← projects

tally

Catch cuBLAS dropping eight numbers from a sum, then redraw how it adds

30 min read tally.md

Part 1: For you

The idea in one line

Fill a 1 × K and a K × 8 fp16 matrix with ones and multiply them with cuBLAS at K = 8,648. Every output ought to be 8,648. On an H100 with split-K algorithm 66 and an affected cuBLAS, it's 8,640: the last eight values of K vanish, with no error and no warning. tally reproduces that bug, checks NVIDIA's fix, fingerprints from the outside how cuBLAS orders and rounds its additions, and publishes a page where anyone can place probe values and watch the sum absorb or keep them.

Why it's exciting

  • The hook is concrete. Ones times ones gives the wrong answer. A September 2026 paper, Taming Bitwise Behavior in GPU Kernels with Tensor Core, found that cuBLASLt's split-K algorithm 66 drops the tail of K for certain shapes.
  • It hid in plain sight. NVIDIA's release notes say the bug came in with cuBLAS 12.6.3 (CUDA 12.6 Update 2) and was fixed in cuBLAS 13.8.0.4 (CUDA 13.4 Update 1). The paper states that torch.matmul doesn't reach this algorithm on its own path.
  • It's forensics on a closed library, done legally. tally never disassembles anything. It calls the public API and looks only at outputs. The paper calls its method the first black-box reconstruction of a closed-source library's arithmetic for bit-level correctness.
  • The trick is tiny. Put +1024 and −1024 next to each other along K, and one very small value r elsewhere. The pair cancels exactly. If r's group adds r and then +1024, r rounds away and the output is 0; if the pair is in another group, r survives. Slide the pair along K, and the output reads off cuBLAS's group boundaries.
  • It's fresh. The paper appeared on September 10, 2026, and the fixed wheel reached PyPI on September 16, 2026. The paper has no code link, and the research found no interactive demo.
  • It ties into LLM nondeterminism. Thinking Machines Lab traced 80 unique completions out of 1,000 at temperature 0 to kernels that aren't batch invariant. tally shows the GEMM version: the same row of x @ W gets different bits as the batch size changes.
  • It's low-level. PTX tensor-core probes, a C++ cuBLASLt harness, bit-exact Triton and CUDA kernels, and an integer-only emulator of the H100 tensor core in Python and TypeScript.
  • It's cheap. One rented H100 for an afternoon or two, about $16-22.

What the demo looks like

The page is static: recorded hardware results plus an in-browser emulator, so visitors don't need an NVIDIA GPU. It has the following sections:

  • The missing eight. A K slider tiles K into 64-wide K-steps grouped into split-K chunks, with the lost tail highlighted. For recorded shapes, it shows the H100's output on the buggy and the fixed cuBLAS side by side, 8,640 against 8,648, and the paper's and NVIDIA's bug conditions light up together.
  • A probe playground. Drag +L, −L, and r along K. The emulator computes the output bit for bit, and a badge shows whether the H100 measured that placement and agreed.
  • A reduction-tree viewer. The recovered order as a tree: 16-product instructions, K-loop steps, split-K partials, and the merge, with the precision at each level.
  • Inside one instruction. A 16-product block with its alignment window. Click a product to see which bits fall off, with truncation contrasted against round-to-nearest.
  • A shape atlas. A zoomable heatmap of which algorithm family cuBLAS picks per shape, with the algorithm 66 bug region overlaid.
  • Batch invariance. A batch-size slider shows how many bits of row 0 of x @ W change compared with batch size 1, and which family cuBLAS used.
  • Before and after the fix. Which outputs changed bits between cuBLAS 13.1.1 and 13.8.0.4, and whether any change falls outside the bug condition.
  • A scoreboard. Your kernels that match cuBLAS bit for bit, with their speed.

How it works

  1. Rent one H100 and install seven versions of cuBLAS, each in its own environment, from before the bug to the fix.
  2. Reproduce the bug. A small C++ program runs algorithm 66 on the paper's shapes. The paper's versions return K − 8, and the fixed one returns K.
  3. Probe one tensor-core instruction. Hand-written PTX computes 16-product inner products, checked against a ported bit-accurate H100 model from Khattak and Mikaitis.
  4. Recover the reduction order. The +L, −L, r probe walks along K for each algorithm family and records where the groups start and end.
  5. Map the bug and the fix. Sweep tens of thousands of shapes, and test the paper's and NVIDIA's conditions against the measurements.
  6. Build matching kernels and an emulator that reproduce cuBLAS's exact bits.
  7. Build the page and ship it as a static site on vm.ifkash.dev.

Weekend plan

The following table lays out the weekend:

When What Done when
Saturday morning Start the pod; install the cuBLAS versions; write the harness; reproduce the bug across the version matrix The paper's versions return K − 8, and 13.8.0.4 returns K
Saturday afternoon Tensor-core probe and model port; start the bug-map and atlas sweeps (they run by themselves) The model matches 10⁶ random inner products on the H100
Saturday evening Probe instruments and descriptor recovery; the Python emulator Descriptors for most families are recovered, and the emulator matches the H100
Sunday morning Triton and CUDA bit-matching kernels; the TypeScript emulator; the fix diff The scoreboard shows bit-identical matches
Sunday afternoon Build the page, deploy it, write the post The link works

Cost

  • GPU: 1× H100 SXM (80 GB) on RunPod, $2.69 per hour on Community Cloud in September 2026. The fallback is Secure Cloud at about $2.99-3.49 per hour.
  • Time on the pod: about 6-8 hours.
  • Optional checks: one hour on an RTX 5090 (about $0.99 per hour) tests NVIDIA's claim that compute capability 12.x isn't affected. One hour on a B200 (about $6.79 per hour) runs only if you approve it.
  • Total: about $16-22. The plan sets a hard stop at $30.
  • Hosting: free. It's a static page on your VM, and it keeps working after the pod is gone.
  • Your Mac: It has no NVIDIA GPU, so it runs only the emulator, the page, and the analysis.

What you have at the end

  • A live link where anyone can watch cuBLAS lose eight numbers, then drag probe values around to see why.
  • A version matrix of which cuBLAS releases drop the tail on an H100.
  • A bug map scored against both the paper's condition and NVIDIA's.
  • Recovered reduction orders for most of cuBLAS's algorithm families on the H100.
  • A tested H100 tensor-core model and a whole-GEMM emulator in Python and TypeScript.
  • Triton and CUDA kernels that match cuBLAS bit for bit, with their speed.
  • A blog post: "cuBLAS forgot eight numbers, and I found where they went."

What might go wrong

Problem What to do
The wheels don't load side by side, or a CUDA 13 wheel needs a newer driver Pick a pod image with a driver new enough for CUDA 13, and give each version its own environment and process.
Algorithm 66 isn't available for a shape Enumerate algorithms, record availability, and force the split count if needed.
The emulator disagrees with the hardware by a few bits Start with the probe values, and bisect with the M2 feature vectors. Don't paper over it.
The paper's instruction model and the truncation model disagree Measure on the H100 and report both.
The fix changed more than the bug region That's a finding. Report it.
EULA worries Stay black-box: public API calls and outputs only, and no disassembly.
The H100 price or availability changes Fall back to Secure Cloud. The cap stays at $30.
Speed numbers are noisy Report medians over repeats. Speed isn't the headline.
"This isn't new" It isn't: Yang et al. found and described the bug. tally is an independent reproduction, a before-and-after check of NVIDIA's fix, and, as far as the research found, the first interactive explainer.
Nobody cares about eight numbers The batch-invariance panel ties it to LLM nondeterminism.

Other ideas the research turned up

  • exp wall. On B200, tensor cores do 8,192 BF16 operations per SM per clock and the exp unit does 16, a 512:1 ratio (FlashAttention-4 blog). cuDNN-frontend PR #1178 moves some exps onto FMA units for +7.83% on B200. It lost because it's a microbenchmark story with a smaller surprise, and it needs four GPU types.
  • GPU MODE overfit audit. Run top kernels from the 532,026 submissions in GPUMODE/kernelbot-data on shapes off the benchmark. It lost on harness-reconstruction risk and B200 cost, and its premise rests on one Hacker News comment.
  • base-13 FP4. arXiv 2608.06812 and 2609.24519 (AWE) rebuild exact INT8 products from FP4 GEMMs with base-13 digits, in 6 FP4 GEMMs instead of 9 (Oz-FP4, MIT). It lost because native INT8 is faster on rentable sm_120 cards (40.94 against 17.31 TFLOPS reported).
  • same prompt, different GPU. In arXiv 2609.25624, 31-100% of problems give a different greedy token stream on at least one pair of GPU architectures. It lost because its fix covers only linear layers and needs several GPU types; tally's batch-invariance panel covers the core idea on one GPU.
  • FP4 attention without exp. arXiv 2609.04105 maps attention scores straight to E2M1 codes, for up to 2.13× over BF16 FlashAttention-4 on GB200. It lost because the kernel needs sm_100 TMA and TMEM, so a weekend version is only an explainer.
  • SASS by linear algebra. F2Asm (arXiv 2608.20532) learns SASS encodings as affine maps over GF(2) and reassembles 3,225 CUBINs byte for byte. It lost because its full source is pending a licensing review, and sm_120 support is only planned.

Reading, if you want it


Part 2: For the coding agent

Mission

Build tally, a black-box forensics project on NVIDIA's closed-source cuBLAS and a static web page that explains it:

  • Harness: A C++ cublasLt program that runs any shape, algorithm, and split count on raw input bits and writes raw fp32 bits, across a matrix of cuBLAS versions.
  • Bug: A reproduction of the algorithm 66 tail loss from arXiv 2609.11356, a measured bug map, and a before-and-after check of NVIDIA's fix.
  • Instruction probe: PTX probes that test a port of the Khattak and Mikaitis H100 tensor-core model bit for bit.
  • Descriptor recovery: The paper's +L, −L, r probe, which recovers a GEMMDesc per family.
  • Emulator and kernels: A bit-exact GEMM emulator in Python and TypeScript, and Triton and CUDA kernels that match cuBLAS bit for bit.
  • Demo page: The eight sections from Part 1.

The final artifact is a static website with no server-side compute. The paper released no code, so you reimplement its method from the text. Record every choice that the paper leaves open in NOTES.md.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a milestone until the previous one passes, except where a milestone says it runs in parallel. After each milestone, commit your work and write a short entry in NOTES.md with the results and numbers.

Hard constraints

  • Black-box only: The CUDA Toolkit EULA, section 1.2, forbids reverse engineering, decompiling, or disassembling any part of the SDK. Call only the public cuBLASLt API, read only the attributes and kernel names that the API or Nsight Systems reports, and observe outputs. Never run cuobjdump, nvdisasm, or any other disassembler or decompiler on a cuBLAS library, and never patch one. You can inspect your own kernels freely.
  • No redistribution: The repo and the site ship no NVIDIA binaries. Pods install cuBLAS wheels from PyPI. The page ships only measured outputs and your own code.
  • Compute: Run all CUDA, Triton, and cuBLAS work on RunPod. The Mac has no NVIDIA GPU; it runs the emulator, the analysis, the web build, and the browser tests.
  • Secrets: Read RUNPOD_API_KEY from the environment only. Never write it into any file, log, commit, or echoed command. Commit a .env.example with placeholder values only, and add .env to .gitignore.
  • Budget: The hard cap is $30 of RunPod spend. Every pod runs a watchdog that stops the pod after MAX_POD_HOURS hours. The default is 3. The approved budget covers H100 pods. Ask the user before you start any other pod, including the optional RTX 5090 and B200 checks.
  • Pod cleanup: When a pod isn't running a job, stop it. At the end of the project, terminate every pod that you created and report the total spend.
  • Outward actions: Ask the user before you do any of the following: create the GitHub repository, deploy to a VM, change DNS, post anything publicly, or contact NVIDIA. Don't file a bug report with NVIDIA or anyone else; the fix has already shipped. Use the kashifulhaque GitHub account (gh auth switch -u kashifulhaque).
  • Shared VM: vm.ifkash.dev runs other production apps behind one shared Caddy, which owns ports 80 and 443. Its config lives at ~/docs/caddy.
  • The tally container lives in ~/docs/tally and must not publish any ports. It joins the external Docker network edge with a stable alias. Check existing aliases with docker network inspect edge first; use tally if it's free, otherwise tally-demo.
  • Add the vhost only by appending to ~/docs/caddy/Caddyfile with >>. Never rewrite, rename, or replace that file. It's a single-file bind mount, and a rewrite orphans the inode, so caddy reload then reports "config is unchanged" while serving the old config.
  • Docker on the VM has no BuildKit for plain docker build. Don't use COPY --chmod, Dockerfile heredocs, or RUN --mount. Multi-stage builds and COPY --from work.
  • Don't stop, restart, reconfigure, or remove any other container, network, or vhost.
  • Licenses: arXiv 2609.11356 is under the arXiv non-exclusive distribution license 1.0, not Creative Commons: cite and link it, but don't copy its figures or text. Credit the Khattak and Mikaitis work (CC BY 4.0) in README.md, the footer, and every ported file's header. State that tally isn't affiliated with the authors or with NVIDIA.
  • Framing: The page's first paragraph credits Yang, Riasanovsky, Deng, and Sarkar for the bug and the method. Describe the bug neutrally, and say that it's fixed. Don't claim that NVIDIA fixed it because of the paper; the dates are only adjacent.

Science background

These facts come from "Taming Bitwise Behavior in GPU Kernels with Tensor Core: Black-Box Reconstruction, Compiler Enforcement, and Static Verification" by Ziteng Yang, Nicholas J. Riasanovsky, Warren Deng, and Vivek Sarkar, on arXiv on September 10, 2026 (cs.DC, cs.PF, cs.PL): https://arxiv.org/abs/2609.11356. Implement the experiments to match them:

  • Claim: Deterministic implementations of the same kernel can still differ bit for bit, mainly because of reduction order, alongside partial-sum precision, FMA, and rounding placement. A tile shape chosen for speed also fixes the arithmetic, which can break batch invariance, and an autotuner can't tell which configurations are bitwise equivalent.
  • Setup: GB300 (sm_103), GB200 (sm_100), H100 (sm_90), and AMD gfx942 (CDNA3), with cuBLAS 13, Triton 3.8, PyTorch 2.12, and CUDA 13, plus cuBLAS 12.8.5.
  • Descriptor (GEMMDesc): instruction_k (products folded per instruction rounding), use_fast_accum, k_loop_step (the mainloop K step), span (the split-K partition length, per nesting level), k_cuts (the nested cutting structure), partial_dtype and merge_dtype (the precision at split stages), layout for GEMV (strided or contiguous), and the instruction type (mma.sync, wgmma, or tcgen05.mma). Two equal records mean the two kernels produce the same bits.
  • Eight families: single-pass accumulation, split-K, per-MMA accumulation, split-K per-MMA, three-level chain, lane-tree GEMV, contiguous-slice GEMV, and workspace GEMV.
  • Instruction model: One instruction folds instruction_k products and the incoming accumulator into a single rounding, acc ← acc ⊕fp32 (Σ products), with exact products.
  • The probe: Place +L, −L, and a tiny r. For fp16, L = 1024 and r = 2⁻¹⁵, which the paper describes as under half an ulp of fp32 at L. The pair cancels exactly. If +L lands in r's group after r, r is absorbed, and the output is 0. If the pair sits wholly in another group, that group cancels on its own, and r reaches the output intact. To find group boundaries, put r at index 0 and walk +L, −L as an adjacent pair along K. A shape profile first narrows a shape to a family.
  • Validation: On GB300 fp16, 110,753 of 110,813 random shapes matched cuBLAS byte for byte; the other 60 are the bug. On H100, 99.82% of 648,720 byte comparisons matched.
  • The bug: It's in the nvjet split-K path that cuBLASLt reaches at ALGO_ID 66. With the K step b = 64, q = ⌊K/b⌋, t = K mod b, and s the split count, values are lost when t ≠ 0, q mod s = 0, and s > t. On the tested shapes, this predicts every loss with no false positives and no false negatives. (1, 8, 8648), (1, 8, 57608), and (1, 8, 11528) each lose 8 values of K. Affected: GB200, GB300, and H100 under cuBLAS 12.8.5 and 13.1.1. Of 496,906 random GB300 shapes, 1,062 weren't bit-identical, and 1,059 of those meet the condition.
  • Appendix D: A and B full of ones, M = 1, N = 8, K = 8648, fp16 inputs, fp32 accumulation, and fp32 output. Expected K, observed K − 8, with the work over nine threadblocks. The paper warns that fp16 output can't be trusted here.
  • Speed (Fig. 5): Above 5 GFLOP, bit-exact Triton reaches 56-93% of cuBLAS and free-order Triton 61-88%. The paper's static equivalence checkers are out of scope.

NVIDIA's CUDA Toolkit release notes list the fix under CUDA 13.4 Update 1 (cuBLAS 13.8.0.4) as resolved issue 6580689: cublasLtMatmul() could return incorrect output with split-K when SPLITK_NUM didn't evenly divide K and ⌊K / SPLITK_NUM⌋ was an exact multiple of the stage size. It affected only CUBLASLT_ALGO_CONFIG_ID 66 on compute capability 9.0, 10.x, and 11.x, and came in with CUDA 12.6 Update 2 (cuBLAS 12.6.3). The CUDA 13.4 (cuBLAS 13.7.0) notes list it as a known issue. Their workaround reads CUBLASLT_ALGO_CONFIG_SPLITK_NUM and the stage size (from CUBLASLT_ALGO_CONFIG_STAGES_ID) with cublasLtMatmulAlgoConfigGetAttribute(), then sets a split count that divides K, or one for which K / SPLITK_NUM isn't a multiple of the stage size.

The following arithmetic is yours, not the paper's or NVIDIA's; check it:

  • The hero shape. For K = 8648, q = 135 and t = 8. s = 9 divides 135, and 9 > 8, so the condition fires. Nine splits cover 9 × 15 × 64 = 8,640 values, which matches the nine threadblocks and K − 8. In general, the lost count is t.
  • The two conditions. If s divides q and 0 < t < s, then ⌊K/s⌋ = 64q/s, a multiple of 64, and K mod s = t ≠ 0, so the paper's condition implies NVIDIA's when the stage size is 64. The converse needs K mod s < 64, which holds whenever s ≤ 64. Don't assume the stage size equals b = 64; read it, and test the equivalence on measured data in M3.

Tensor-core arithmetic comes from "Accurate Models of NVIDIA Tensor Cores" by Khattak and Mikaitis (arXiv 2512.07004, v4 June 11, 2026, CC BY 4.0): https://arxiv.org/abs/2512.07004. It's verified with constructed vectors and 10⁷ random inputs, and it finds that B200 tensor cores behave identically to H100. For H100 with fp16 or bf16 inputs and fp32 output:

  • Block: NFMA = 16 products per block, and the accumulator input c is summed together with the products.
  • Alignment: 2 extra alignment bits; products aligned to (2, 25) bits, that is, 2 integer and 25 fractional bits. Bits shifted out during alignment are truncated.
  • Normalization: results are truncated rather than rounded, and products stay denormalized during accumulation. With c = 0, the maximum exponent used for alignment is limited at −133.
  • Code: the MATLAB Tensor Core toolbox v0.5, CC BY 4.0, with H100TC(), B200TC(), and CustomTC().

This truncation model differs from the paper's single-rounding model. Reconcile them on the H100 in M2, and treat it as an open detail, not an error in either paper.

The probe sits on the edge of that model. L = 2¹⁰ and r = 2⁻¹⁵ are 25 binades apart, exactly the width of the fractional window. Inside one instruction, r next to L meets alignment truncation. Across instructions, r meets the fp32 accumulator, where the ulp at 2¹⁰ is 2⁻¹³, so r is below half an ulp. The probe relies on that distinction, so measure it instead of assuming it. If you need to, pick a smaller r, such as 2⁻²⁰, so that r is lost both inside and across instructions, and record the choice. As a standard binary16 fact (check it), 2⁻¹⁵ is below fp16's smallest normal, 2⁻¹⁴, so r is a subnormal input.

For M2 feature-vector ideas, see the Fasi et al. test suite.

Batch invariance comes from Thinking Machines Lab's Defeating Nondeterminism in LLM Inference (Horace He, September 2025). At temperature 0, Qwen3-235B gave 80 unique completions out of 1,000, because kernels aren't batch invariant: a request's numerics depend on the batch size through the reduction strategy, such as split-K. Which cuBLAS families flip at which M on the H100 is for you to measure.

Design decisions

  • Version matrix: Test the following cuBLAS versions, from nvidia-cublas-cu12 for 12.x and nvidia-cublas for 13.x:
Version Role PyPI upload
12.6.1.4 Before the bug; control Confirm in M0
12.6.3.3 First affected 2024-10-01
12.8.5.5 The paper's version 2026-04-08
13.1.1.3 The paper's version 2026-04-08
13.7.0.27 CUDA 13.4; known issue 2026-09-09
13.4.2.4 Unknown; check whether it has the fix 2026-09-15
13.8.0.4 Fixed (CUDA 13.4 Update 1) 2026-09-16

Confirm that each wheel installs and loads, and record its exact cublasLtGetVersion() value. If the wheel-to-toolkit mapping differs from this table, record what you found. - Version isolation: One environment per version under /workspace/cublas/<ver>, and one process per version. Select the library at run time with LD_LIBRARY_PATH or dlopen. If the ABI differs between 12.x and 13.x, build one harness binary per major version. - Raw bits: The harness reads A and B as raw input bit patterns and writes C as raw fp32 bits. Compare outputs as uint32, never as floats with a tolerance. Use fp32 output for every bug test. - Forcing algorithm 66: List IDs with cublasLtMatmulAlgoGetIds, initialize with cublasLtMatmulAlgoInit, and read CUBLASLT_ALGO_CONFIG_SPLITK_NUM and CUBLASLT_ALGO_CONFIG_STAGES_ID with cublasLtMatmulAlgoConfigGetAttribute. If the default split count doesn't trigger the bug, set it with cublasLtMatmulAlgoConfigSetAttribute, and record that you forced it. - Probe placement: Put probe values in A and fill B with ones, so each product equals the A value. Zero every other entry along K. - Families on H100: The paper also ran on Blackwell, so some families might not be reachable on sm_90. Record which ones you reach. - Descriptor format: ref/desc.py defines GEMMDesc as a frozen dataclass with the paper's fields, serialized to JSON with the family, cuBLAS version, algorithm ID, split count, and kernel name. Record equality is the equivalence test. - Emulator arithmetic: Integer only: sign, exponent, and integer significand, with alignment, truncation, and rounding done by shifts. No float shortcuts. In TypeScript, use BigInt or paired uint32 operations. - Kernel names: Read them only with Nsight Systems, for example for the torch.matmul path. That's black-box observation. - Atlas grid (choice; record it): At M = 1, log-spaced N and K from 8 to 65,536; at N = 4,096, log-spaced M and K; plus dense K stripes around the bug condition. Keep the atlas under 5 MB compressed in total. - Batch invariance (choice; record it): Fixed x with K = 4,096 and W of 4,096 × 4,096, for M from 1 to 512. Record row 0's bits and the algorithm ID, split count, and family per M. - Speed: Report medians over at least 20 repeats after 5 warmups. Speed has no target. - Optional stretch, untested: A "your GPU" tab that runs a WGSL reduction with the same probe to fingerprint a visitor's subgroupAdd or workgroup reduction order, which WGSL leaves implementation defined. Build it only after M8, and label it experimental.

Tech stack

Pin every version. The stack is as follows:

  • Pod: RunPod H100 SXM with an official PyTorch image that supports CUDA 13 and a driver new enough for CUDA 13. nvcc builds the harness and the PTX probes. Pin Triton and PyTorch; the paper used Triton 3.8 and PyTorch 2.12. Nsight Systems is for kernel names only.
  • cuBLAS: the per-version wheels from the design decisions, each in its own environment.
  • Mac: Python 3.12 with uv, numpy, pytest, and matplotlib, for the emulator, the model, and the analysis. No GPU is needed.
  • Web: TypeScript, bundled with vite, with D3 (optional) or hand-written SVG, tested with Playwright. No heavy framework is required.
  • Pods: the runpod Python SDK, with runpodctl inside pods.
  • Results: JSONL records and compact little-endian binary under results/, with no hosted tracker.

Repository layout

Create the following layout:

tally/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh
  harness/probe.cu  harness/Makefile        # cublasLt harness, raw bits in and out
  harness/tc_mma.cu                         # PTX mma.sync / wgmma single-instruction probes
  ref/tcmodel.py      # H100 inner-product model (ported from Khattak & Mikaitis, CC BY 4.0)
  ref/desc.py         # GEMMDesc dataclass, eight families
  ref/probe.py        # shape profile + +L/-L/r instruments, descriptor recovery
  ref/emulate.py      # bit-exact GEMM emulator from a GEMMDesc
  ref/conditions.py   # the paper's and NVIDIA's bug conditions
  kernels/triton_gemm.py  kernels/splitk_gemv.cu
  exp/versions.py  exp/bugmap.py  exp/atlas.py  exp/batchinv.py  exp/fixdiff.py
  exp/scoreboard.py
  export/export.py    # results -> web/public/*.json|bin + manifest.json
  web/
    src/emu/          # TS port of tcmodel + emulate
    src/panels/       # one module per demo section
    src/main.ts  index.html
    public/           # recorded outputs, descriptors, atlas, goldens
    tests/            # Playwright: emulator goldens, smoke test
  deploy/Dockerfile  deploy/compose.yml  deploy/nginx.conf
  results/  post/draft.md

M0: Setup and pod infrastructure

Tasks:

  1. Scaffold the repository, the uv project, and the vite app. Add .gitignore and .env.example.
  2. Write infra/pod.py. It uses the runpod SDK and reads RUNPOD_API_KEY from the environment. It supports the following subcommands: - create: Creates a pod named tally-*. The default GPU is H100 SXM on Community Cloud, with Secure Cloud as the fallback. Resolve GPU type IDs at run time. --gpu selects another type only after the user approves it. - status, stop, and terminate. - ssh-info: Prints the SSH command. - cost: Prints the uptime and spend for every pod whose name has the prefix tally-.
  3. Write infra/watchdog.sh. It sleeps for MAX_POD_HOURS, then runs runpodctl stop pod $RUNPOD_POD_ID.
  4. Write infra/bootstrap.sh. It takes a cuBLAS version, installs that wheel into its own environment, builds the harness, and starts the watchdog.
  5. Check that every version in the matrix, including 12.6.1.4, exists on PyPI, and record the result in NOTES.md.

Acceptance criteria:

  • A pod comes up, and bootstrap.sh finishes without errors for at least one 12.x and one 13.x version.
  • A test run with MAX_POD_HOURS=0.05 stops the pod within 5 minutes.
  • infra/pod.py cost prints the spend so far.

M1: cuBLASLt harness and the bug

Tasks:

  • Write harness/probe.cu against cublasLt. It takes M, N, K, the dtype, an algorithm ID or heuristic index, and an optional split count. It fills A and B from a raw file or a generator (ones, seeded random, probe placement), and writes C as raw fp32 bits plus a JSON sidecar with the algorithm ID, split count, stage ID, and cublasLtGetVersion().
  • Write exp/versions.py. For each version in the matrix, run the Appendix D reproducer and the three example shapes with algorithm 66, in a separate process.
  • Run the same shapes through torch.matmul with fp32 output, and record its result and the kernel name it launches.

Acceptance criteria:

  • On H100 with 12.8.5.5 and 13.1.1.3, (1, 8, 8648), (1, 8, 11528), and (1, 8, 57608) return K − 8 in every output element.
  • 13.8.0.4 returns K on the same shapes, and so does 12.6.1.4 (the control).
  • NOTES.md has a table of every version's result, including 12.6.3.3, 13.7.0.27, and 13.4.2.4, with the split count and stage ID that each run used.
  • torch.matmul returns K on the same shapes, as the paper states, and NOTES.md records the algorithm or kernel it used.

M2: Tensor-core instruction probe

Start the M3 bug-map sweep and the M5 atlas sweep on the pod as soon as M1 passes. They run in the background while you work on M2.

Tasks:

  • Write harness/tc_mma.cu with inline PTX mma.sync m16n8k16 (fp16 inputs, fp32 accumulator). Each output element is one 16-product inner product plus a chosen accumulator value. Add a wgmma variant if time allows, and record whether the two instructions agree.
  • Port the Khattak and Mikaitis H100 model to ref/tcmodel.py from the MATLAB toolbox, with a CC BY 4.0 attribution header. Use integer arithmetic only.
  • Build feature vectors that separate the model's claims: truncation against round-to-nearest at normalization, the 2 extra alignment bits, c summed with the products, and the −133 cap when c = 0.
  • Run L = 1024 and r = 2⁻¹⁵, then smaller r, inside one instruction and across two.

Acceptance criteria:

  • The model matches the H100 bit for bit on 10⁶ random inner products that mix normal, subnormal, and widely spread magnitudes, and on every feature vector.
  • NOTES.md records where the paper's single-rounding description and the truncation model predict different bits, and what the H100 does in each case.
  • NOTES.md records the r that the probe uses, and why.

M3: Probe instruments and descriptor recovery

Tasks:

  • Write ref/conditions.py with the paper's condition and NVIDIA's condition, the latter taking the stage size as an input.
  • Write ref/probe.py. It drives the harness: first a shape profile to narrow a shape to a family, then +L, −L, r walks along K to recover k_loop_step, span, k_cuts, and the precision at each stage. Recover a GEMMDesc for a set of shapes per family.
  • Write exp/bugmap.py: at least 20,000 random M = 1 shapes, plus some with M > 1, with algorithm 66 on 13.1.1.3, recording the sums, split count, and stage ID per shape.
  • Test the equivalence between the two conditions on every measured shape.

Acceptance criteria:

  • Both conditions predict every loss on the sweep with zero false positives and zero false negatives, or NOTES.md documents each discrepancy with its shape.
  • NOTES.md has a confusion matrix for each condition, and states whether the two conditions agreed on every shape.
  • Recovered descriptors for at least 6 of the 8 families are recorded, and the families that aren't reachable on the H100 are noted.

M4: Bit-exact emulator

Tasks:

  • Write ref/emulate.py. Given a GEMMDesc, the inputs, and the dtype, it produces the exact fp32 output bits, built on ref/tcmodel.py. It has a switch for the paper's single-rounding instruction model.
  • Port the model and the emulator to web/src/emu/*.ts.
  • Export goldens: for every recovered family, a set of shapes, inputs, and H100 output bits, including probe placements and the hero shape on a buggy and a fixed version.

Acceptance criteria:

  • The Python emulator matches the H100 on at least 5,000 (shape, random input) pairs per recovered family, including probe placements.
  • The Python and TypeScript emulators agree bit for bit on every golden.
  • NOTES.md records whether the single-rounding switch diverges from the hardware, and where.

M5: Atlas, batch invariance, and the fix diff

Tasks:

  • Write exp/atlas.py: over the atlas grid, record the top-1 result of cublasLtMatmulAlgoGetHeuristic and the torch.matmul path, with the algorithm ID, the split count, and the family.
  • Write exp/batchinv.py: the batch-invariance sweep over M, for fixed N and K.
  • Write exp/fixdiff.py: the bit diff of 13.1.1.3 against 13.8.0.4 on identical inputs over the atlas grid, with the default heuristic and with algorithm 66 forced.
  • Optional, with the user's approval: repeat M1 and a small bug map on an RTX 5090 (12.x, not in NVIDIA's affected list) or a B200.

Acceptance criteria:

  • The atlas files are under 5 MB compressed in total.
  • Every fix-diff change is classified as inside or outside the bug condition. If any change falls outside it, NOTES.md reports it as a finding.
  • The batch-invariance results list the M values where row 0 changes, with the family on each side of the change.
  • If the RTX 5090 check ran, NOTES.md states whether any condition shape lost values.

M6: Bit-matching kernels

Tasks:

  • Write kernels/triton_gemm.py: Triton GEMMs parameterized by GEMMDesc for the single-pass and split-K families, and for per-MMA accumulation if it's reachable.
  • Write kernels/splitk_gemv.cu: a hand-written CUDA split-K GEMV that handles the K tail correctly, with a buggy flag that reproduces the dropped tail exactly.
  • Write exp/scoreboard.py: shapes tested, matches, and speed as a percentage of cuBLAS per family.

Acceptance criteria:

  • Each implemented family is bit-identical to cuBLAS on at least 2,000 shapes: on 13.8.0.4 for all shapes, and on 13.1.1.3 only outside the bug condition.
  • In buggy mode, the GEMV reproduces the 13.1.1.3 output bit for bit on every bug-map shape that meets the condition.
  • Speed is measured and recorded in NOTES.md next to the paper's 56-93% for bit-exact Triton. There's no speed target.

M7: The demo page

Tasks:

  • The missing eight: A K slider with a split-count readout that draws 64-wide K-steps tiled into s chunks, with the tail highlighted. For recorded shapes, show the H100 output on 13.1.1.3 and 13.8.0.4 side by side, and light up both bug conditions together.
  • Probe playground: Drag +L, −L, and r along K for a recorded shape. The TypeScript emulator computes the output, and a badge shows whether the H100 measured that placement and whether it agrees.
  • Reduction-tree viewer: A shape's descriptor as a tree of instructions, K-loop steps, split-K partials, and the merge, with the dtype at each level.
  • Inside one instruction: A 16-product block with the (2, 25)-bit window. Clicking a product shows its truncated bits, with a truncation versus round-to-nearest switch.
  • Shape atlas: Zoomable heatmaps of N × K at fixed M and M × K at fixed N, colored by the heuristic's family, with the algorithm 66 bug region overlaid.
  • Batch invariance, fix diff, and scoreboard: The M5 and M6 results, with the changed-bit count against M = 1 and any fix-diff change outside the bug condition called out.
  • Fallback and mobile: No WebGPU needed. On phones, stack the panels and support touch.
  • Footer: The Yang et al. credit, the Khattak and Mikaitis attribution, the NVIDIA release notes link, and "not affiliated with the authors or NVIDIA."

Acceptance criteria:

  • A Playwright smoke test loads the page, moves the K slider, drags a probe, switches shapes, and opens the atlas.
  • The emulator goldens pass in headless Chromium.
  • The page works in Chrome and Safari on macOS, and on one phone.

M8: Deploy and write-up

Tasks:

  • Write deploy/Dockerfile: nginx:alpine serving web/dist, buildable with the legacy builder. Serve precompressed files with gzip_static on, and set application/octet-stream for .bin.
  • Write deploy/compose.yml. It publishes no ports and joins the external network edge with the alias chosen in the hard constraints.
  • Ask the user before deploying. The user confirms a domain such as tally.ifkash.dev and adds a non-proxied Cloudflare A record.
  • After the user approves, copy the build and deploy/ to ~/docs/tally, and run docker compose up -d there. Then compare the live and host Caddyfiles with docker exec caddy cat /etc/caddy/Caddyfile | diff - ~/docs/caddy/Caddyfile. If they differ, stop and ask the user. Otherwise, back up, append, validate, and reload:

cd ~/docs/caddy cp Caddyfile "Caddyfile.bak-tally-$(date +%Y%m%d)" printf '\ntally.ifkash.dev {\n\treverse_proxy tally:80\n}\n' >> Caddyfile docker exec -i caddy caddy validate --config - --adapter caddyfile < Caddyfile docker exec -i caddy caddy reload --config - --adapter caddyfile < Caddyfile

Replace tally in reverse_proxy with the alias you chose. Omit any tls block: for a non-proxied A record, the default ACME HTTP-01 challenge works. - Write README.md: what the project is, one command to reproduce each milestone, the results, the version matrix, every open choice and deviation, and the credits. - Write post/draft.md, a blog post of 1,200-1,800 words. Structure it as follows: - The hook: "cuBLAS forgot eight numbers, and I found where they went." - What happened: ones times ones, K − 8, and Yang et al.'s discovery. - How the probe works, and inside the tensor core: truncation and what the H100 did. - The atlas, batch invariance, and LLM nondeterminism. - The fix: the version matrix, the fix diff, and both conditions against your bug map. - Limitations: one GPU type, black-box inference only, and unreachable families. - Ask the user before you push anything. After the user approves, create and push the repo on the kashifulhaque account. The owner lists the project on projects.dotslasha.me after the deploy; that isn't your job.

Acceptance criteria:

  • The deployed URL loads over HTTPS, and the hero panel shows 8,640 against 8,648.
  • Every other site on the VM still responds as it did before the deploy.
  • No pods are left running, and NOTES.md records the total spend, which is less than $30.

Final report to the user

When you finish, report the following:

  • The version-matrix table from M1, with each version's cublasLtGetVersion() value and result, and the torch.matmul result.
  • The M2 model agreement and what the H100 does where truncation and single rounding differ, plus the r you chose.
  • The M3 bug-map confusion matrix for both conditions, and whether they agreed on every shape.
  • The recovered families, and the families you couldn't reach on the H100.
  • The M4 emulator agreement per family, and whether the single-rounding switch diverged.
  • The M5 highlights: the atlas, where batch invariance breaks, and the fix diff.
  • The M6 scoreboard: matches and speed per family.
  • The RTX 5090 and B200 results, if they ran.
  • The URL, if the site is deployed.
  • The total spend.
  • Anything that you skipped or that failed.