# tally: catch cuBLAS dropping eight numbers from a sum, then redraw how it adds

---

# Part 1: For you

## The idea in one line

Fill a 1 × K and a K × 8 fp16 matrix with ones and multiply them with cuBLAS at K = 8,648.
Every output ought to be 8,648. On an H100 with split-K algorithm 66 and an affected cuBLAS,
it's 8,640: the last eight values of K vanish, with no error and no warning. tally reproduces
that bug, checks NVIDIA's fix, fingerprints from the outside how cuBLAS orders and rounds its
additions, and publishes a page where anyone can place probe values and watch the sum absorb
or keep them.

## Why it's exciting

- **The hook is concrete.** Ones times ones gives the wrong answer. A September 2026 paper,
  [Taming Bitwise Behavior in GPU Kernels with Tensor Core](https://arxiv.org/abs/2609.11356),
  found that cuBLASLt's split-K algorithm 66 drops the tail of K for certain shapes.
- **It hid in plain sight.** NVIDIA's release notes say the bug came in with cuBLAS 12.6.3
  (CUDA 12.6 Update 2) and was fixed in cuBLAS 13.8.0.4 (CUDA 13.4 Update 1). The paper states
  that `torch.matmul` doesn't reach this algorithm on its own path.
- **It's forensics on a closed library, done legally.** tally never disassembles anything. It
  calls the public API and looks only at outputs. The paper calls its method the first
  black-box reconstruction of a closed-source library's arithmetic for bit-level correctness.
- **The trick is tiny.** Put +1024 and −1024 next to each other along K, and one very small
  value r elsewhere. The pair cancels exactly. If r's group adds r and then +1024, r rounds
  away and the output is 0; if the pair is in another group, r survives. Slide the pair along
  K, and the output reads off cuBLAS's group boundaries.
- **It's fresh.** The paper appeared on September 10, 2026, and the fixed wheel reached PyPI on
  September 16, 2026. The paper has no code link, and the research found no interactive demo.
- **It ties into LLM nondeterminism.** Thinking Machines Lab traced 80 unique completions out
  of 1,000 at temperature 0 to kernels that aren't batch invariant. tally shows the GEMM
  version: the same row of `x @ W` gets different bits as the batch size changes.
- **It's low-level.** PTX tensor-core probes, a C++ cuBLASLt harness, bit-exact Triton and CUDA
  kernels, and an integer-only emulator of the H100 tensor core in Python and TypeScript.
- **It's cheap.** One rented H100 for an afternoon or two, about $16-22.

## What the demo looks like

The page is static: recorded hardware results plus an in-browser emulator, so visitors don't
need an NVIDIA GPU. It has the following sections:

- **The missing eight.** A K slider tiles K into 64-wide K-steps grouped into split-K chunks,
  with the lost tail highlighted. For recorded shapes, it shows the H100's output on the buggy
  and the fixed cuBLAS side by side, 8,640 against 8,648, and the paper's and NVIDIA's bug
  conditions light up together.
- **A probe playground.** Drag +L, −L, and r along K. The emulator computes the output bit for
  bit, and a badge shows whether the H100 measured that placement and agreed.
- **A reduction-tree viewer.** The recovered order as a tree: 16-product instructions, K-loop
  steps, split-K partials, and the merge, with the precision at each level.
- **Inside one instruction.** A 16-product block with its alignment window. Click a product to
  see which bits fall off, with truncation contrasted against round-to-nearest.
- **A shape atlas.** A zoomable heatmap of which algorithm family cuBLAS picks per shape, with
  the algorithm 66 bug region overlaid.
- **Batch invariance.** A batch-size slider shows how many bits of row 0 of `x @ W` change
  compared with batch size 1, and which family cuBLAS used.
- **Before and after the fix.** Which outputs changed bits between cuBLAS 13.1.1 and 13.8.0.4,
  and whether any change falls outside the bug condition.
- **A scoreboard.** Your kernels that match cuBLAS bit for bit, with their speed.

## How it works

1. **Rent one H100 and install seven versions of cuBLAS,** each in its own environment, from
   before the bug to the fix.
2. **Reproduce the bug.** A small C++ program runs algorithm 66 on the paper's shapes. The
   paper's versions return K − 8, and the fixed one returns K.
3. **Probe one tensor-core instruction.** Hand-written PTX computes 16-product inner products,
   checked against a ported bit-accurate H100 model from Khattak and Mikaitis.
4. **Recover the reduction order.** The +L, −L, r probe walks along K for each algorithm
   family and records where the groups start and end.
5. **Map the bug and the fix.** Sweep tens of thousands of shapes, and test the paper's and
   NVIDIA's conditions against the measurements.
6. **Build matching kernels and an emulator** that reproduce cuBLAS's exact bits.
7. **Build the page and ship it** as a static site on `vm.ifkash.dev`.

## Weekend plan

The following table lays out the weekend:

| When | What | Done when |
|---|---|---|
| Saturday morning | Start the pod; install the cuBLAS versions; write the harness; reproduce the bug across the version matrix | The paper's versions return K − 8, and 13.8.0.4 returns K |
| Saturday afternoon | Tensor-core probe and model port; start the bug-map and atlas sweeps (they run by themselves) | The model matches 10⁶ random inner products on the H100 |
| Saturday evening | Probe instruments and descriptor recovery; the Python emulator | Descriptors for most families are recovered, and the emulator matches the H100 |
| Sunday morning | Triton and CUDA bit-matching kernels; the TypeScript emulator; the fix diff | The scoreboard shows bit-identical matches |
| Sunday afternoon | Build the page, deploy it, write the post | The link works |

## Cost

- **GPU:** 1× H100 SXM (80 GB) on RunPod, $2.69 per hour on Community Cloud in September
  2026. The fallback is Secure Cloud at about $2.99-3.49 per hour.
- **Time on the pod:** about 6-8 hours.
- **Optional checks:** one hour on an RTX 5090 (about $0.99 per hour) tests NVIDIA's claim
  that compute capability 12.x isn't affected. One hour on a B200 (about $6.79 per hour) runs
  only if you approve it.
- **Total:** about $16-22. The plan sets a hard stop at $30.
- **Hosting:** free. It's a static page on your VM, and it keeps working after the pod is gone.
- **Your Mac:** It has no NVIDIA GPU, so it runs only the emulator, the page, and the
  analysis.

## What you have at the end

- A live link where anyone can watch cuBLAS lose eight numbers, then drag probe values around
  to see why.
- A version matrix of which cuBLAS releases drop the tail on an H100.
- A bug map scored against both the paper's condition and NVIDIA's.
- Recovered reduction orders for most of cuBLAS's algorithm families on the H100.
- A tested H100 tensor-core model and a whole-GEMM emulator in Python and TypeScript.
- Triton and CUDA kernels that match cuBLAS bit for bit, with their speed.
- A blog post: *"cuBLAS forgot eight numbers, and I found where they went."*

## What might go wrong

| Problem | What to do |
|---|---|
| The wheels don't load side by side, or a CUDA 13 wheel needs a newer driver | Pick a pod image with a driver new enough for CUDA 13, and give each version its own environment and process. |
| Algorithm 66 isn't available for a shape | Enumerate algorithms, record availability, and force the split count if needed. |
| The emulator disagrees with the hardware by a few bits | Start with the probe values, and bisect with the M2 feature vectors. Don't paper over it. |
| The paper's instruction model and the truncation model disagree | Measure on the H100 and report both. |
| The fix changed more than the bug region | That's a finding. Report it. |
| EULA worries | Stay black-box: public API calls and outputs only, and no disassembly. |
| The H100 price or availability changes | Fall back to Secure Cloud. The cap stays at $30. |
| Speed numbers are noisy | Report medians over repeats. Speed isn't the headline. |
| "This isn't new" | It isn't: Yang et al. found and described the bug. tally is an independent reproduction, a before-and-after check of NVIDIA's fix, and, as far as the research found, the first interactive explainer. |
| Nobody cares about eight numbers | The batch-invariance panel ties it to LLM nondeterminism. |

## Other ideas the research turned up

- **exp wall.** On B200, tensor cores do 8,192 BF16 operations per SM per clock and the exp
  unit does 16, a 512:1 ratio ([FlashAttention-4 blog](https://tridao.me/blog/2026/flash4/)).
  [cuDNN-frontend PR #1178](https://github.com/NVIDIA/cudnn-frontend/pull/1178) moves some
  exps onto FMA units for +7.83% on B200. It lost because it's a microbenchmark story with a
  smaller surprise, and it needs four GPU types.
- **GPU MODE overfit audit.** Run top kernels from the 532,026 submissions in
  [GPUMODE/kernelbot-data](https://huggingface.co/datasets/GPUMODE/kernelbot-data) on shapes
  off the benchmark. It lost on harness-reconstruction risk and B200 cost, and its premise
  rests on one Hacker News comment.
- **base-13 FP4.** arXiv 2608.06812 and 2609.24519 (AWE) rebuild exact INT8 products from FP4
  GEMMs with base-13 digits, in 6 FP4 GEMMs instead of 9
  ([Oz-FP4](https://github.com/FP4-is-All-you-Need/Oz-FP4), MIT). It lost because native INT8
  is faster on rentable sm_120 cards (40.94 against 17.31 TFLOPS reported).
- **same prompt, different GPU.** In arXiv 2609.25624, 31-100% of problems give a different
  greedy token stream on at least one pair of GPU architectures. It lost because its fix covers
  only linear layers and needs several GPU types; tally's batch-invariance panel covers the
  core idea on one GPU.
- **FP4 attention without exp.** arXiv 2609.04105 maps attention scores straight to E2M1 codes,
  for up to 2.13× over BF16 FlashAttention-4 on GB200. It lost because the kernel needs sm_100
  TMA and TMEM, so a weekend version is only an explainer.
- **SASS by linear algebra.** F2Asm (arXiv 2608.20532) learns SASS encodings as affine maps
  over GF(2) and reassembles 3,225 CUBINs byte for byte. It lost because its full source is
  pending a licensing review, and sm_120 support is only planned.

## Reading, if you want it

- [Taming Bitwise Behavior in GPU Kernels with Tensor Core](https://arxiv.org/abs/2609.11356)
  (Yang et al., 2026): the paper this project builds on, which found the bug.
- [CUDA Toolkit release notes](https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html):
  NVIDIA's fix, listed under CUDA 13.4 Update 1 as resolved issue 6580689.
- [Accurate Models of NVIDIA Tensor Cores](https://arxiv.org/abs/2512.07004) (Khattak and
  Mikaitis) and its [MATLAB Tensor Core toolbox](https://github.com/north-numerical-computing/MATLAB-tensor-core):
  the bit-accurate H100 model that tally ports.
- [Numerical behavior of NVIDIA tensor cores](https://peerj.com/articles/cs-330/) (Fasi,
  Higham, Mikaitis, and Pranesh, 2021): the earlier study of V100, T4, and A100.
- [Accurate Models of AMD Matrix Cores](https://arxiv.org/abs/2609.14845) (Khattak, Mikaitis,
  and Graziani, 2026): the same method on AMD CDNA.
- [Defeating Nondeterminism in LLM Inference](https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/)
  (Thinking Machines Lab, 2025): why batch invariance matters for LLMs.
- [cuBLAS](https://docs.nvidia.com/cuda/cublas/index.html),
  [PTX ISA](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html), and
  [Triton](https://triton-lang.org/) documentation.

---

# Part 2: For the coding agent

## Mission

Build `tally`, a black-box forensics project on NVIDIA's closed-source cuBLAS and a static web
page that explains it:

- **Harness:** A C++ `cublasLt` program that runs any shape, algorithm, and split count on raw
  input bits and writes raw fp32 bits, across a matrix of cuBLAS versions.
- **Bug:** A reproduction of the algorithm 66 tail loss from arXiv 2609.11356, a measured bug
  map, and a before-and-after check of NVIDIA's fix.
- **Instruction probe:** PTX probes that test a port of the Khattak and Mikaitis H100
  tensor-core model bit for bit.
- **Descriptor recovery:** The paper's +L, −L, r probe, which recovers a `GEMMDesc` per family.
- **Emulator and kernels:** A bit-exact GEMM emulator in Python and TypeScript, and Triton and
  CUDA kernels that match cuBLAS bit for bit.
- **Demo page:** The eight sections from Part 1.

The final artifact is a static website with no server-side compute. The paper released no
code, so you reimplement its method from the text. Record every choice that the paper leaves
open in `NOTES.md`.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a
milestone until the previous one passes, except where a milestone says it runs in parallel.
After each milestone, commit your work and write a short entry in `NOTES.md` with the results
and numbers.

## Hard constraints

- **Black-box only:** The CUDA Toolkit EULA, section 1.2, forbids reverse engineering,
  decompiling, or disassembling any part of the SDK. Call only the public cuBLASLt API, read
  only the attributes and kernel names that the API or Nsight Systems reports, and observe
  outputs. Never run `cuobjdump`, `nvdisasm`, or any other disassembler or decompiler
  on a cuBLAS library, and never patch one. You can inspect your own kernels freely.
- **No redistribution:** The repo and the site ship no NVIDIA binaries. Pods install cuBLAS
  wheels from PyPI. The page ships only measured outputs and your own code.
- **Compute:** Run all CUDA, Triton, and cuBLAS work on RunPod. The Mac has no NVIDIA GPU; it
  runs the emulator, the analysis, the web build, and the browser tests.
- **Secrets:** Read `RUNPOD_API_KEY` from the environment only. Never write it into any file,
  log, commit, or echoed command. Commit a `.env.example` with placeholder values only, and
  add `.env` to `.gitignore`.
- **Budget:** The hard cap is $30 of RunPod spend. Every pod runs a watchdog that stops the pod
  after `MAX_POD_HOURS` hours. The default is `3`. The approved budget covers H100 pods. Ask
  the user before you start any other pod, including the optional RTX 5090 and B200 checks.
- **Pod cleanup:** When a pod isn't running a job, stop it. At the end of the project,
  terminate every pod that you created and report the total spend.
- **Outward actions:** Ask the user before you do any of the following: create the GitHub
  repository, deploy to a VM, change DNS, post anything publicly, or contact NVIDIA. Don't
  file a bug report with NVIDIA or anyone else; the fix has already shipped. Use the
  `kashifulhaque` GitHub account (`gh auth switch -u kashifulhaque`).
- **Shared VM:** `vm.ifkash.dev` runs other production apps behind one shared Caddy, which owns
  ports 80 and 443. Its config lives at `~/docs/caddy`.
  - The tally container lives in `~/docs/tally` and must not publish any ports. It joins the
    external Docker network `edge` with a stable alias. Check existing aliases with
    `docker network inspect edge` first; use `tally` if it's free, otherwise `tally-demo`.
  - Add the vhost only by appending to `~/docs/caddy/Caddyfile` with `>>`. Never rewrite,
    rename, or replace that file. It's a single-file bind mount, and a rewrite orphans the
    inode, so `caddy reload` then reports "config is unchanged" while serving the old config.
  - Docker on the VM has no BuildKit for plain `docker build`. Don't use `COPY --chmod`,
    Dockerfile heredocs, or `RUN --mount`. Multi-stage builds and `COPY --from` work.
  - Don't stop, restart, reconfigure, or remove any other container, network, or vhost.
- **Licenses:** arXiv 2609.11356 is under the arXiv non-exclusive distribution license 1.0,
  not Creative Commons: cite and link it, but don't copy its figures or text. Credit the
  Khattak and Mikaitis work (CC BY 4.0) in `README.md`, the footer, and every ported file's
  header. State that tally isn't affiliated with the authors or with NVIDIA.
- **Framing:** The page's first paragraph credits Yang, Riasanovsky, Deng, and Sarkar for the
  bug and the method. Describe the bug neutrally, and say that it's fixed. Don't claim that
  NVIDIA fixed it because of the paper; the dates are only adjacent.

## Science background

These facts come from "Taming Bitwise Behavior in GPU Kernels with Tensor Core: Black-Box
Reconstruction, Compiler Enforcement, and Static Verification" by Ziteng Yang, Nicholas J.
Riasanovsky, Warren Deng, and Vivek Sarkar, on arXiv on September 10, 2026 (cs.DC, cs.PF,
cs.PL): <https://arxiv.org/abs/2609.11356>. Implement the experiments to match them:

- **Claim:** Deterministic implementations of the same kernel can still differ bit for bit,
  mainly because of reduction order, alongside partial-sum precision, FMA, and rounding
  placement. A tile shape chosen for speed also fixes the arithmetic, which can break batch
  invariance, and an autotuner can't tell which configurations are bitwise equivalent.
- **Setup:** GB300 (sm_103), GB200 (sm_100), H100 (sm_90), and AMD gfx942 (CDNA3), with cuBLAS
  13, Triton 3.8, PyTorch 2.12, and CUDA 13, plus cuBLAS 12.8.5.
- **Descriptor (`GEMMDesc`):** `instruction_k` (products folded per instruction rounding),
  `use_fast_accum`, `k_loop_step` (the mainloop K step), `span` (the split-K partition length,
  per nesting level), `k_cuts` (the nested cutting structure), `partial_dtype` and
  `merge_dtype` (the precision at split stages), `layout` for GEMV (strided or contiguous), and
  the instruction type (`mma.sync`, `wgmma`, or `tcgen05.mma`). Two equal records mean the
  two kernels produce the same bits.
- **Eight families:** single-pass accumulation, split-K, per-MMA accumulation, split-K
  per-MMA, three-level chain, lane-tree GEMV, contiguous-slice GEMV, and workspace GEMV.
- **Instruction model:** One instruction folds `instruction_k` products and the incoming
  accumulator into a single rounding, acc ← acc ⊕fp32 (Σ products), with exact products.
- **The probe:** Place +L, −L, and a tiny r. For fp16, L = 1024 and r = 2⁻¹⁵, which the paper
  describes as under half an ulp of fp32 at L. The pair cancels exactly. If +L lands in r's
  group after r, r is absorbed, and the output is 0. If the pair sits wholly in another group,
  that group cancels on its own, and r reaches the output intact. To find group boundaries,
  put r at index 0 and walk +L, −L as an adjacent pair along K. A shape profile first narrows
  a shape to a family.
- **Validation:** On GB300 fp16, 110,753 of 110,813 random shapes matched cuBLAS byte for
  byte; the other 60 are the bug. On H100, 99.82% of 648,720 byte comparisons matched.
- **The bug:** It's in the nvjet split-K path that cuBLASLt reaches at `ALGO_ID` 66. With the
  K step b = 64, q = ⌊K/b⌋, t = K mod b, and s the split count, values are lost when t ≠ 0,
  q mod s = 0, and s > t. On the tested shapes, this predicts every loss with no false
  positives and no false negatives. (1, 8, 8648), (1, 8, 57608), and (1, 8, 11528) each lose 8
  values of K. Affected: GB200, GB300, and H100 under cuBLAS 12.8.5 and 13.1.1. Of 496,906
  random GB300 shapes, 1,062 weren't bit-identical, and 1,059 of those meet the condition.
- **Appendix D:** A and B full of ones, M = 1, N = 8, K = 8648, fp16 inputs, fp32
  accumulation, and fp32 output. Expected K, observed K − 8, with the work over nine
  threadblocks. The paper warns that fp16 output can't be trusted here.
- **Speed (Fig. 5):** Above 5 GFLOP, bit-exact Triton reaches 56-93% of cuBLAS and free-order
  Triton 61-88%. The paper's static equivalence checkers are out of scope.

NVIDIA's [CUDA Toolkit release notes](https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html)
list the fix under CUDA 13.4 Update 1 (cuBLAS 13.8.0.4) as resolved issue 6580689:
`cublasLtMatmul()` could return incorrect output with split-K when `SPLITK_NUM` didn't evenly
divide K and ⌊K / `SPLITK_NUM`⌋ was an exact multiple of the stage size. It affected only
`CUBLASLT_ALGO_CONFIG_ID` 66 on compute capability 9.0, 10.x, and 11.x, and came in with CUDA
12.6 Update 2 (cuBLAS 12.6.3). The CUDA 13.4 (cuBLAS 13.7.0) notes list it as a known issue.
Their workaround reads `CUBLASLT_ALGO_CONFIG_SPLITK_NUM` and the stage size (from
`CUBLASLT_ALGO_CONFIG_STAGES_ID`) with `cublasLtMatmulAlgoConfigGetAttribute()`, then sets a
split count that divides K, or one for which K / `SPLITK_NUM` isn't a multiple of the stage
size.

The following arithmetic is yours, not the paper's or NVIDIA's; check it:

- **The hero shape.** For K = 8648, q = 135 and t = 8. s = 9 divides 135, and 9 > 8, so the
  condition fires. Nine splits cover 9 × 15 × 64 = 8,640 values, which matches the nine
  threadblocks and K − 8. In general, the lost count is t.
- **The two conditions.** If s divides q and 0 < t < s, then ⌊K/s⌋ = 64q/s, a multiple of 64,
  and K mod s = t ≠ 0, so the paper's condition implies NVIDIA's when the stage size is 64.
  The converse needs K mod s < 64, which holds whenever s ≤ 64. Don't assume the stage size
  equals b = 64; read it, and test the equivalence on measured data in M3.

Tensor-core arithmetic comes from "Accurate Models of NVIDIA Tensor Cores" by Khattak and
Mikaitis (arXiv 2512.07004, v4 June 11, 2026, CC BY 4.0): <https://arxiv.org/abs/2512.07004>.
It's verified with constructed vectors and 10⁷ random inputs, and it finds that B200 tensor
cores behave identically to H100. For H100 with fp16 or bf16 inputs and fp32 output:

- **Block:** NFMA = 16 products per block, and the accumulator input c is summed together with
  the products.
- **Alignment:** 2 extra alignment bits; products aligned to (2, 25) bits, that is, 2 integer
  and 25 fractional bits. Bits shifted out during alignment are truncated.
- **Normalization:** results are truncated rather than rounded, and products stay denormalized
  during accumulation. With c = 0, the maximum exponent used for alignment is limited at −133.
- **Code:** the [MATLAB Tensor Core](https://github.com/north-numerical-computing/MATLAB-tensor-core)
  toolbox v0.5, CC BY 4.0, with `H100TC()`, `B200TC()`, and `CustomTC()`.

This truncation model differs from the paper's single-rounding model. Reconcile them on the
H100 in M2, and treat it as an open detail, not an error in either paper.

The probe sits on the edge of that model. L = 2¹⁰ and r = 2⁻¹⁵ are 25 binades apart, exactly
the width of the fractional window. Inside one instruction, r next to L meets alignment
truncation. Across instructions, r meets the fp32 accumulator, where the ulp at 2¹⁰ is 2⁻¹³,
so r is below half an ulp. The probe relies on that distinction, so measure it instead of
assuming it. If you need to, pick a smaller r, such as 2⁻²⁰, so that r is lost both inside
and across instructions, and record the choice. As a standard binary16 fact (check it), 2⁻¹⁵
is below fp16's smallest normal, 2⁻¹⁴, so r is a subnormal input.

For M2 feature-vector ideas, see the
[Fasi et al. test suite](https://github.com/north-numerical-computing/tensor-cores-numerical-behavior).

Batch invariance comes from Thinking Machines Lab's
[Defeating Nondeterminism in LLM Inference](https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/)
(Horace He, September 2025). At temperature 0, Qwen3-235B gave 80 unique completions out of
1,000, because kernels aren't batch invariant: a request's numerics depend on the batch size
through the reduction strategy, such as split-K. Which cuBLAS families flip at which M on the
H100 is for you to measure.

## Design decisions

- **Version matrix:** Test the following cuBLAS versions, from `nvidia-cublas-cu12` for 12.x
  and `nvidia-cublas` for 13.x:

  | Version | Role | PyPI upload |
  |---|---|---|
  | 12.6.1.4 | Before the bug; control | Confirm in M0 |
  | 12.6.3.3 | First affected | 2024-10-01 |
  | 12.8.5.5 | The paper's version | 2026-04-08 |
  | 13.1.1.3 | The paper's version | 2026-04-08 |
  | 13.7.0.27 | CUDA 13.4; known issue | 2026-09-09 |
  | 13.4.2.4 | Unknown; check whether it has the fix | 2026-09-15 |
  | 13.8.0.4 | Fixed (CUDA 13.4 Update 1) | 2026-09-16 |

  Confirm that each wheel installs and loads, and record its exact `cublasLtGetVersion()`
  value. If the wheel-to-toolkit mapping differs from this table, record what you found.
- **Version isolation:** One environment per version under `/workspace/cublas/<ver>`, and one
  process per version. Select the library at run time with `LD_LIBRARY_PATH` or `dlopen`. If
  the ABI differs between 12.x and 13.x, build one harness binary per major version.
- **Raw bits:** The harness reads A and B as raw input bit patterns and writes C as raw fp32
  bits. Compare outputs as `uint32`, never as floats with a tolerance. Use fp32 output for
  every bug test.
- **Forcing algorithm 66:** List IDs with `cublasLtMatmulAlgoGetIds`, initialize with
  `cublasLtMatmulAlgoInit`, and read `CUBLASLT_ALGO_CONFIG_SPLITK_NUM` and
  `CUBLASLT_ALGO_CONFIG_STAGES_ID` with `cublasLtMatmulAlgoConfigGetAttribute`. If the default
  split count doesn't trigger the bug, set it with `cublasLtMatmulAlgoConfigSetAttribute`, and
  record that you forced it.
- **Probe placement:** Put probe values in A and fill B with ones, so each product equals the
  A value. Zero every other entry along K.
- **Families on H100:** The paper also ran on Blackwell, so some families might not be
  reachable on sm_90. Record which ones you reach.
- **Descriptor format:** `ref/desc.py` defines `GEMMDesc` as a frozen dataclass with the
  paper's fields, serialized to JSON with the family, cuBLAS version, algorithm ID, split
  count, and kernel name. Record equality is the equivalence test.
- **Emulator arithmetic:** Integer only: sign, exponent, and integer significand, with
  alignment, truncation, and rounding done by shifts. No float shortcuts. In TypeScript, use
  `BigInt` or paired `uint32` operations.
- **Kernel names:** Read them only with Nsight Systems, for example for the `torch.matmul`
  path. That's black-box observation.
- **Atlas grid (choice; record it):** At M = 1, log-spaced N and K from 8 to 65,536; at N =
  4,096, log-spaced M and K; plus dense K stripes around the bug condition. Keep the atlas
  under 5 MB compressed in total.
- **Batch invariance (choice; record it):** Fixed `x` with K = 4,096 and `W` of 4,096 × 4,096,
  for M from 1 to 512. Record row 0's bits and the algorithm ID, split count, and family per M.
- **Speed:** Report medians over at least 20 repeats after 5 warmups. Speed has no target.
- **Optional stretch, untested:** A "your GPU" tab that runs a WGSL reduction with the same
  probe to fingerprint a visitor's `subgroupAdd` or workgroup reduction order, which WGSL
  leaves implementation defined. Build it only after M8, and label it experimental.

## Tech stack

Pin every version. The stack is as follows:

- **Pod:** RunPod H100 SXM with an official PyTorch image that supports CUDA 13 and a driver
  new enough for CUDA 13. `nvcc` builds the harness and the PTX probes. Pin Triton and
  PyTorch; the paper used Triton 3.8 and PyTorch 2.12. Nsight Systems is for kernel names only.
- **cuBLAS:** the per-version wheels from the design decisions, each in its own environment.
- **Mac:** Python 3.12 with `uv`, `numpy`, `pytest`, and `matplotlib`, for the emulator, the
  model, and the analysis. No GPU is needed.
- **Web:** TypeScript, bundled with `vite`, with D3 (optional) or hand-written SVG, tested with
  Playwright. No heavy framework is required.
- **Pods:** the `runpod` Python SDK, with `runpodctl` inside pods.
- **Results:** JSONL records and compact little-endian binary under `results/`, with no hosted
  tracker.

## Repository layout

Create the following layout:

```
tally/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh
  harness/probe.cu  harness/Makefile        # cublasLt harness, raw bits in and out
  harness/tc_mma.cu                         # PTX mma.sync / wgmma single-instruction probes
  ref/tcmodel.py      # H100 inner-product model (ported from Khattak & Mikaitis, CC BY 4.0)
  ref/desc.py         # GEMMDesc dataclass, eight families
  ref/probe.py        # shape profile + +L/-L/r instruments, descriptor recovery
  ref/emulate.py      # bit-exact GEMM emulator from a GEMMDesc
  ref/conditions.py   # the paper's and NVIDIA's bug conditions
  kernels/triton_gemm.py  kernels/splitk_gemv.cu
  exp/versions.py  exp/bugmap.py  exp/atlas.py  exp/batchinv.py  exp/fixdiff.py
  exp/scoreboard.py
  export/export.py    # results -> web/public/*.json|bin + manifest.json
  web/
    src/emu/          # TS port of tcmodel + emulate
    src/panels/       # one module per demo section
    src/main.ts  index.html
    public/           # recorded outputs, descriptors, atlas, goldens
    tests/            # Playwright: emulator goldens, smoke test
  deploy/Dockerfile  deploy/compose.yml  deploy/nginx.conf
  results/  post/draft.md
```

## M0: Setup and pod infrastructure

**Tasks:**

1. Scaffold the repository, the `uv` project, and the `vite` app. Add `.gitignore` and
   `.env.example`.
2. Write `infra/pod.py`. It uses the `runpod` SDK and reads `RUNPOD_API_KEY` from the
   environment. It supports the following subcommands:
   - `create`: Creates a pod named `tally-*`. The default GPU is H100 SXM on Community Cloud,
     with Secure Cloud as the fallback. Resolve GPU type IDs at run time. `--gpu` selects
     another type only after the user approves it.
   - `status`, `stop`, and `terminate`.
   - `ssh-info`: Prints the SSH command.
   - `cost`: Prints the uptime and spend for every pod whose name has the prefix `tally-`.
3. Write `infra/watchdog.sh`. It sleeps for `MAX_POD_HOURS`, then runs
   `runpodctl stop pod $RUNPOD_POD_ID`.
4. Write `infra/bootstrap.sh`. It takes a cuBLAS version, installs that wheel into its own
   environment, builds the harness, and starts the watchdog.
5. Check that every version in the matrix, including 12.6.1.4, exists on PyPI, and record the
   result in `NOTES.md`.

**Acceptance criteria:**

- A pod comes up, and `bootstrap.sh` finishes without errors for at least one 12.x and one 13.x
  version.
- A test run with `MAX_POD_HOURS=0.05` stops the pod within 5 minutes.
- `infra/pod.py cost` prints the spend so far.

## M1: cuBLASLt harness and the bug

**Tasks:**

- Write `harness/probe.cu` against `cublasLt`. It takes M, N, K, the dtype, an algorithm ID
  or heuristic index, and an optional split count. It fills A and B from a raw file or a
  generator (ones, seeded random, probe placement), and writes C as raw fp32 bits plus a JSON
  sidecar with the algorithm ID, split count, stage ID, and `cublasLtGetVersion()`.
- Write `exp/versions.py`. For each version in the matrix, run the Appendix D reproducer and
  the three example shapes with algorithm 66, in a separate process.
- Run the same shapes through `torch.matmul` with fp32 output, and record its result and the
  kernel name it launches.

**Acceptance criteria:**

- On H100 with 12.8.5.5 and 13.1.1.3, (1, 8, 8648), (1, 8, 11528), and (1, 8, 57608) return
  K − 8 in every output element.
- 13.8.0.4 returns K on the same shapes, and so does 12.6.1.4 (the control).
- `NOTES.md` has a table of every version's result, including 12.6.3.3, 13.7.0.27, and
  13.4.2.4, with the split count and stage ID that each run used.
- `torch.matmul` returns K on the same shapes, as the paper states, and `NOTES.md` records the
  algorithm or kernel it used.

## M2: Tensor-core instruction probe

Start the M3 bug-map sweep and the M5 atlas sweep on the pod as soon as M1 passes. They run in
the background while you work on M2.

**Tasks:**

- Write `harness/tc_mma.cu` with inline PTX `mma.sync` m16n8k16 (fp16 inputs, fp32
  accumulator). Each output element is one 16-product inner product plus a chosen accumulator
  value. Add a `wgmma` variant if time allows, and record whether the two instructions agree.
- Port the Khattak and Mikaitis H100 model to `ref/tcmodel.py` from the MATLAB toolbox, with a
  CC BY 4.0 attribution header. Use integer arithmetic only.
- Build feature vectors that separate the model's claims: truncation against round-to-nearest
  at normalization, the 2 extra alignment bits, c summed with the products, and the −133 cap
  when c = 0.
- Run L = 1024 and r = 2⁻¹⁵, then smaller r, inside one instruction and across two.

**Acceptance criteria:**

- The model matches the H100 bit for bit on 10⁶ random inner products that mix normal,
  subnormal, and widely spread magnitudes, and on every feature vector.
- `NOTES.md` records where the paper's single-rounding description and the truncation model
  predict different bits, and what the H100 does in each case.
- `NOTES.md` records the r that the probe uses, and why.

## M3: Probe instruments and descriptor recovery

**Tasks:**

- Write `ref/conditions.py` with the paper's condition and NVIDIA's condition, the latter
  taking the stage size as an input.
- Write `ref/probe.py`. It drives the harness: first a shape profile to narrow a shape to a
  family, then +L, −L, r walks along K to recover `k_loop_step`, `span`, `k_cuts`, and the
  precision at each stage. Recover a `GEMMDesc` for a set of shapes per family.
- Write `exp/bugmap.py`: at least 20,000 random M = 1 shapes, plus some with M > 1, with
  algorithm 66 on 13.1.1.3, recording the sums, split count, and stage ID per shape.
- Test the equivalence between the two conditions on every measured shape.

**Acceptance criteria:**

- Both conditions predict every loss on the sweep with zero false positives and zero false
  negatives, or `NOTES.md` documents each discrepancy with its shape.
- `NOTES.md` has a confusion matrix for each condition, and states whether the two conditions
  agreed on every shape.
- Recovered descriptors for at least 6 of the 8 families are recorded, and the families that
  aren't reachable on the H100 are noted.

## M4: Bit-exact emulator

**Tasks:**

- Write `ref/emulate.py`. Given a `GEMMDesc`, the inputs, and the dtype, it produces the exact
  fp32 output bits, built on `ref/tcmodel.py`. It has a switch for the paper's single-rounding
  instruction model.
- Port the model and the emulator to `web/src/emu/*.ts`.
- Export goldens: for every recovered family, a set of shapes, inputs, and H100 output bits,
  including probe placements and the hero shape on a buggy and a fixed version.

**Acceptance criteria:**

- The Python emulator matches the H100 on at least 5,000 (shape, random input) pairs per
  recovered family, including probe placements.
- The Python and TypeScript emulators agree bit for bit on every golden.
- `NOTES.md` records whether the single-rounding switch diverges from the hardware, and where.

## M5: Atlas, batch invariance, and the fix diff

**Tasks:**

- Write `exp/atlas.py`: over the atlas grid, record the top-1 result of
  `cublasLtMatmulAlgoGetHeuristic` and the `torch.matmul` path, with the algorithm ID, the
  split count, and the family.
- Write `exp/batchinv.py`: the batch-invariance sweep over M, for fixed N and K.
- Write `exp/fixdiff.py`: the bit diff of 13.1.1.3 against 13.8.0.4 on identical inputs over
  the atlas grid, with the default heuristic and with algorithm 66 forced.
- Optional, with the user's approval: repeat M1 and a small bug map on an RTX 5090 (12.x, not
  in NVIDIA's affected list) or a B200.

**Acceptance criteria:**

- The atlas files are under 5 MB compressed in total.
- Every fix-diff change is classified as inside or outside the bug condition. If any change
  falls outside it, `NOTES.md` reports it as a finding.
- The batch-invariance results list the M values where row 0 changes, with the family on each
  side of the change.
- If the RTX 5090 check ran, `NOTES.md` states whether any condition shape lost values.

## M6: Bit-matching kernels

**Tasks:**

- Write `kernels/triton_gemm.py`: Triton GEMMs parameterized by `GEMMDesc` for the single-pass
  and split-K families, and for per-MMA accumulation if it's reachable.
- Write `kernels/splitk_gemv.cu`: a hand-written CUDA split-K GEMV that handles the K tail
  correctly, with a `buggy` flag that reproduces the dropped tail exactly.
- Write `exp/scoreboard.py`: shapes tested, matches, and speed as a percentage of cuBLAS per
  family.

**Acceptance criteria:**

- Each implemented family is bit-identical to cuBLAS on at least 2,000 shapes: on 13.8.0.4 for
  all shapes, and on 13.1.1.3 only outside the bug condition.
- In `buggy` mode, the GEMV reproduces the 13.1.1.3 output bit for bit on every bug-map shape
  that meets the condition.
- Speed is measured and recorded in `NOTES.md` next to the paper's 56-93% for bit-exact
  Triton. There's no speed target.

## M7: The demo page

**Tasks:**

- **The missing eight:** A K slider with a split-count readout that draws 64-wide K-steps
  tiled into s chunks, with the tail highlighted. For recorded shapes, show the H100 output on
  13.1.1.3 and 13.8.0.4 side by side, and light up both bug conditions together.
- **Probe playground:** Drag +L, −L, and r along K for a recorded shape. The TypeScript
  emulator computes the output, and a badge shows whether the H100 measured that placement and
  whether it agrees.
- **Reduction-tree viewer:** A shape's descriptor as a tree of instructions, K-loop steps,
  split-K partials, and the merge, with the dtype at each level.
- **Inside one instruction:** A 16-product block with the (2, 25)-bit window. Clicking a
  product shows its truncated bits, with a truncation versus round-to-nearest switch.
- **Shape atlas:** Zoomable heatmaps of N × K at fixed M and M × K at fixed N, colored by the
  heuristic's family, with the algorithm 66 bug region overlaid.
- **Batch invariance, fix diff, and scoreboard:** The M5 and M6 results, with the changed-bit
  count against M = 1 and any fix-diff change outside the bug condition called out.
- **Fallback and mobile:** No WebGPU needed. On phones, stack the panels and support touch.
- **Footer:** The Yang et al. credit, the Khattak and Mikaitis attribution, the NVIDIA release
  notes link, and "not affiliated with the authors or NVIDIA."

**Acceptance criteria:**

- A Playwright smoke test loads the page, moves the K slider, drags a probe, switches shapes,
  and opens the atlas.
- The emulator goldens pass in headless Chromium.
- The page works in Chrome and Safari on macOS, and on one phone.

## M8: Deploy and write-up

**Tasks:**

- Write `deploy/Dockerfile`: `nginx:alpine` serving `web/dist`, buildable with the legacy
  builder. Serve precompressed files with `gzip_static on`, and set
  `application/octet-stream` for `.bin`.
- Write `deploy/compose.yml`. It publishes no ports and joins the external network `edge` with
  the alias chosen in the hard constraints.
- Ask the user before deploying. The user confirms a domain such as `tally.ifkash.dev` and
  adds a non-proxied Cloudflare A record.
- After the user approves, copy the build and `deploy/` to `~/docs/tally`, and run
  `docker compose up -d` there. Then compare the live and host Caddyfiles with
  `docker exec caddy cat /etc/caddy/Caddyfile | diff - ~/docs/caddy/Caddyfile`. If they
  differ, stop and ask the user. Otherwise, back up, append, validate, and reload:

  ```
  cd ~/docs/caddy
  cp Caddyfile "Caddyfile.bak-tally-$(date +%Y%m%d)"
  printf '\ntally.ifkash.dev {\n\treverse_proxy tally:80\n}\n' >> Caddyfile
  docker exec -i caddy caddy validate --config - --adapter caddyfile < Caddyfile
  docker exec -i caddy caddy reload --config - --adapter caddyfile < Caddyfile
  ```

  Replace `tally` in `reverse_proxy` with the alias you chose. Omit any `tls` block: for a
  non-proxied A record, the default ACME HTTP-01 challenge works.
- Write `README.md`: what the project is, one command to reproduce each milestone, the
  results, the version matrix, every open choice and deviation, and the credits.
- Write `post/draft.md`, a blog post of 1,200-1,800 words. Structure it as follows:
  - The hook: "cuBLAS forgot eight numbers, and I found where they went."
  - What happened: ones times ones, K − 8, and Yang et al.'s discovery.
  - How the probe works, and inside the tensor core: truncation and what the H100 did.
  - The atlas, batch invariance, and LLM nondeterminism.
  - The fix: the version matrix, the fix diff, and both conditions against your bug map.
  - Limitations: one GPU type, black-box inference only, and unreachable families.
- Ask the user before you push anything. After the user approves, create and push the repo on
  the `kashifulhaque` account. The owner lists the project on projects.dotslasha.me after the
  deploy; that isn't your job.

**Acceptance criteria:**

- The deployed URL loads over HTTPS, and the hero panel shows 8,640 against 8,648.
- Every other site on the VM still responds as it did before the deploy.
- No pods are left running, and `NOTES.md` records the total spend, which is less than $30.

## Final report to the user

When you finish, report the following:

- The version-matrix table from M1, with each version's `cublasLtGetVersion()` value and
  result, and the `torch.matmul` result.
- The M2 model agreement and what the H100 does where truncation and single rounding differ,
  plus the r you chose.
- The M3 bug-map confusion matrix for both conditions, and whether they agreed on every shape.
- The recovered families, and the families you couldn't reach on the H100.
- The M4 emulator agreement per family, and whether the single-rounding switch diverged.
- The M5 highlights: the atlas, where batch invariance breaks, and the fix diff.
- The M6 scoreboard: matches and speed per family.
- The RTX 5090 and B200 results, if they ran.
- The URL, if the site is deployed.
- The total spend.
- Anything that you skipped or that failed.
