# static: a network learns to read handwriting from TV static, live in your browser

---

# Part 1: For you

## The idea in one line

A teacher network learns to read handwritten digits. A student never sees a digit or a label;
it only learns to copy 20 meaningless extra "ghost" outputs of the teacher on TV static. Then
it reads digits at about 75%. Give the student a different random starting point, and the
identical procedure gives about 9.5%, which is chance. The whole experiment trains live in
your browser tab, in WGSL that you write by hand.

## Why it's exciting

- **The hook is strange and true.** A September 2026 paper,
  [Why Ghost Outputs Teach](https://arxiv.org/abs/2609.23260), trains a tiny MNIST network
  and distills a student only on the teacher's ghost outputs, on pure Gaussian noise. With a
  shared initialization, the student reaches 75.16% on the MNIST test set. With a different
  initialization, it reaches 9.52%. A student that isn't trained at all scores 8.37%.
- **Static teaches better than real digits.** In the paper's input-type ablation, uniform
  noise gives 69.1 ± 0.3% and Gaussian noise 68.9 ± 0.1%, while real MNIST images give only
  49.9 ± 3.0%. Shuffled MNIST gives 43.7 ± 6.5%, and a constant gray image gives 9.8%.
- **There's one clean explanation.** The paper derives a *chained cross-task kernel*. With a
  shared initialization, it factors into the form M Mᵀ, so in a one-step idealization the
  update can't push the student away from the correct label. The ghost head is a rank
  bottleneck, and noise excites every direction of the backbone. The demo computes pieces of
  this live.
- **It's contested, and you get to referee.** A May 2026 paper,
  [Learning Through Noise](https://arxiv.org/abs/2605.23645), argues that a matched
  initialization isn't necessary and that *compatible output heads* govern the effect. The
  demo runs both claims side by side, so a visitor sees the disagreement instead of a verdict.
- **The phenomenon has a clear origin.** Cloud et al. named *subliminal learning* in 2025
  ([arXiv 2507.14805](https://arxiv.org/abs/2507.14805)); their MNIST MLP with 3 auxiliary
  logits passed 50% on noise alone. The page credits them first, and static animates the
  September 2026 paper's numbers and explanation.
- **Nobody has built this demo.** The paper has no code link. Searches for subliminal
  learning with MNIST, ghost outputs, WebGPU, and "demo" found only Python repos, such as the
  official [MinhxLe/subliminal-learning](https://github.com/MinhxLe/subliminal-learning) (MIT),
  and no browser demo. The paper appeared on September 20, 2026, so ship soon.
- **It's small enough to train in the tab.** The network has 81,530 parameters. A full
  distillation run is about 1,500 Adam steps, a few seconds on a laptop GPU by estimate.
- **It's low-level.** You write the whole training loop in WGSL: matmuls, ReLU, backprop, two
  losses, Adam, and a counter-based noise generator. There's no server and no ML runtime.

## What the demo looks like

- **A live distillation panel.** Noise images flicker by while the student trains on them,
  and a test-accuracy curve climbs from about 10% toward 75%. A step counter and a
  milliseconds-per-step readout sit beside it.
- **A "same birthday / different birthday" toggle.** Flip it and rerun: the same procedure
  from a different random initialization flatlines at chance.
- **A ghost-dimension slider** from 1 to 100, marked with the paper's result: a sharp rise
  up to 10 ghost outputs, and about 83% with no bottleneck.
- **An input-type picker:** uniform noise, Gaussian noise, real digits, shuffled digits, or a
  constant gray image, each with the paper's number next to yours.
- **Freeze toggles** for the backbone and the heads, which reproduce the paper's Table 1.
- **A draw pad.** Draw a digit with your mouse or finger, and the static-trained student
  reads it, next to the teacher's reading.
- **An inspector** that plots the per-step alignment the paper uses as evidence, plus the
  ghost head's rank and a noise-versus-digits kernel overlap score.
- **A dispute grid:** backbone {shared, random} × heads {shared, random}, which tests the
  May 2026 claim that heads, not the initialization, carry the effect.

## How it works

1. **Get the digits.** Download MNIST on your Mac and pack the test set and a 10,000-image
   training subset into small binary files for the page.
2. **Build one random-number generator.** Write a counter-based hash that gives the same
   noise and the same initial weights in Python and WGSL for a given seed.
3. **Reproduce the paper in PyTorch.** Train the teacher on MNIST, then distill students on
   noise. Reproduce Table 1, the ghost-dimension curve, and the input-type bars across 10
   seeds. Each run takes seconds on the Mac.
4. **Run the dispute.** Re-run the student with a random backbone but the teacher's heads,
   the setup that the May 2026 paper says is enough.
5. **Write the kernels.** Rewrite forward, backward, both losses, and Adam in WGSL, and check
   them step by step against a float32 NumPy mirror.
6. **Build the page and ship it** as a static site on `vm.ifkash.dev`.

## Weekend plan

| When | What | Done when |
|---|---|---|
| Saturday morning | Pack MNIST; write the hash RNG in Python; train teachers; reproduce Table 1 in PyTorch | Same-init students reach 65% or more, and different-init students stay at chance |
| Saturday afternoon | Start the seed sweep for ablations and the dispute grid (runs by itself); write the NumPy mirror and the first WGSL kernels | Noise beats real digits, and the ghost-dimension curve rises from 1 to 10 |
| Saturday evening | WGSL backward pass, MSE and KL losses, Adam, and on-GPU noise; step parity tests | 10 browser steps match NumPy within 1e-4 |
| Sunday morning | The full in-tab distillation loop, evaluation, inspector, and dispute grid | A browser run lands within 2 points of PyTorch for the same seed |
| Sunday afternoon | Build the demo page, deploy it, record the video, write the post | The link works |

## Cost

- **Compute:** $0. The model has 81,530 parameters, so the PyTorch reference and every sweep
  run on your Mac's CPU or its Apple GPU (MPS). No rented GPU is needed.
- **Optional pod:** If a large seed sweep (for example, 100 seeds per condition) is too slow
  on the Mac, run it on one small RunPod pod. The plan sets a hard stop at $10; the estimate
  is under $3.
- **Hosting:** free. It's a static page on your VM, and visitors' GPUs do all the training.

## What you have at the end

- A live link where anyone can teach a network to read from TV static, then break it by
  changing its birthday.
- Your own Table 1, ghost-dimension curve, and input-type bars over 10 seeds, next to the
  paper's numbers.
- A dispute grid that shows what the May 2026 "heads, not init" claim does on this exact
  network.
- A hand-written WGSL training loop with step-by-step parity tests against PyTorch.
- A video of the accuracy curve climbing on static, then flatlining with a new birthday.
- A blog post: *"I taught a neural network to read handwriting using nothing but TV static."*

## What might go wrong

| Problem | What to do |
|---|---|
| The numbers don't reproduce | Follow Appendix C exactly, record every unstated choice, and report your numbers next to the paper's, even if they're lower. The claim is the gap between shared and different init, not the exact 75.16%. |
| The effect depends on the setup | The paper uses MSE on Gaussian noise for Table 1 and KL on uniform noise for the input-type bars. Keep each experiment on its own loss and noise, and show both losses on the ghost-dimension slider. |
| WGSL and PyTorch drift apart | Use the same hash RNG for the initialization and the noise on both sides. Compare the first steps tightly and the final accuracy loosely, because float differences compound over 1,500 Adam steps. |
| Results differ slightly across GPUs | Write reductions without atomics, in a fixed order. Expect per-device drift; test within tolerances, not bit for bit. |
| "This isn't new" | It isn't: Cloud et al. found it in 2025. Credit them in the page's first paragraph, and frame static as an interactive reproduction of a September 2026 explanation. |
| The seed claim is disputed | Present the dispute grid as an open question with both papers linked. Don't write "shared init is necessary" as settled fact. |
| WebGPU isn't available in a visitor's browser | Show a recorded video and static charts of the reference runs instead. |
| Phones are slow or thermally throttle | Use smaller batches and fewer evaluation passes on mobile, and show a "this takes longer on phones" note. |
| MNIST license is inconsistent across sources | Treat it as CC BY-SA 3.0, the stricter reading, and attribute LeCun, Cortes, and Burges in the footer. |

## Other ideas the research turned up

- **[dustfall](https://arxiv.org/abs/2609.23105).** Irregular microparticles settle with
  lateral drift up to 11° and rotational misalignment up to 46°, which any scalar or spheroid
  drag model says is zero. SHEAR, a 0.98M-parameter equivariant net, gets 1.4% and 2.7% error
  on the two tensor blocks. It lost because generating the Stokes-solver data is most of the
  weekend, and the net predicts no coupling block, so grains drift but don't spin.
- **[primesweeper](https://arxiv.org/abs/2609.22968).** Solving a token game for a balanced
  semiprime is equivalent to factoring it, and AlphaZero-style search factors 2 of 5 random
  31-32-bit instances at N = 16. It lost on tree-search speed in the browser at N = 16.
- **[fork](https://arxiv.org/abs/2609.26621).** Greedy decoding gives different outputs in
  BF16 and FP16 for 49-100% of prompts. It lost because it needs a full LLM engine in WGSL,
  emulated BF16 won't match per prompt, and part of the effect was known in 2025.
- **[ancestor](https://arxiv.org/abs/2609.15240).** Click an internal node of a tree of 21
  tyrant flycatcher species and hear a reconstructed ancestral song. It lost because there's
  no ground truth, most of the latent comes from the nearest real clip, and bird recordings
  are mostly NC-licensed.
- **[twin](https://arxiv.org/abs/2609.29845).** Averaging two prompts' inputs gives roughly
  the average of their next-token distributions, and this linearity fades as pretraining goes
  on. It lost because the effect is weak at browser-sized models, so the output can look like
  mush.

## Reading, if you want it

- [Why Ghost Outputs Teach](https://arxiv.org/abs/2609.23260): the paper this project builds
  on.
- [Subliminal Learning](https://arxiv.org/abs/2507.14805) (Cloud et al., 2025): the paper that
  found the effect, with the first MNIST auxiliary-logit experiment.
- [Learning Through Noise](https://arxiv.org/abs/2605.23645) (Brockers et al., 2026): the
  "compatible heads, not shared init" counterclaim.
- [Distilling the Knowledge in a Neural Network](https://arxiv.org/abs/1503.02531) (Hinton et
  al., 2015): knowledge distillation, the setup this effect subverts.
- [Neural Tangent Kernel](https://arxiv.org/abs/1806.07572) (Jacot et al., 2018): the kernel
  behind the paper's explanation.
- [WGSL specification](https://www.w3.org/TR/WGSL/) (W3C) and
  [WebGPU Fundamentals](https://webgpufundamentals.org/): the shader language and a practical
  tutorial.

---

# Part 2: For the coding agent

## Mission

Build `static`, a browser demo of subliminal learning on MNIST that trains entirely on the
client:

- **Reference:** A PyTorch reproduction of the September 2026 paper's MNIST experiments
  (Table 1, Fig. 3a, and Fig. 3b), plus the dispute experiment from arXiv 2605.23645, over
  10 seeds.
- **RNG:** A counter-based hash that produces identical initial weights and identical noise
  in Python and WGSL from a seed.
- **Kernels:** A full training loop in WGSL: forward, backward, MSE and softmax-KL losses,
  Adam, on-GPU noise, and batched evaluation.
- **Demo page:** Live distillation, the birthday toggle, the ghost-dimension slider, the
  input-type picker, freeze toggles, a draw pad, the inspector, and the dispute grid.

The final artifact is a static website with no inference server. The paper released no code,
so you reimplement it from the text. Record every choice that the paper leaves open in
`NOTES.md`.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a
milestone until the previous one passes, except where a milestone says it runs in parallel.
After each milestone, commit your work and write a short entry in `NOTES.md` with the results
and numbers.

## Hard constraints

- **Compute:** Run everything on the Mac by default: data preparation, PyTorch training on
  CPU or MPS, sweeps, the NumPy mirror, the web build, and the browser tests.
- **Optional pod:** Use RunPod only if a sweep of more than 10 seeds per condition won't
  finish overnight on the Mac, and ask the user first. If you use it:
  - Read `RUNPOD_API_KEY` from the environment only. Never write it into any file, log,
    commit, or echoed command. Commit a `.env.example` with placeholder values only, and add
    `.env` to `.gitignore`.
  - The hard cap is $10 of RunPod spend. Every pod runs a watchdog that stops the pod after
    `MAX_POD_HOURS` hours. The default is `3`.
  - Stop a pod when it isn't running a job. At the end, terminate every pod that you created
    and report the total spend.
- **No leakage into the student:** The distillation code path never reads MNIST labels or
  MNIST test images. The real-digit and shuffled-digit input types read only the training
  subset's images. Enforce this with function signatures that take no labels, an assertion,
  and a test.
- **Outward actions:** Ask the user before you do any of the following: create the GitHub
  repository, deploy to a VM, change DNS, or post anything publicly. Use the `kashifulhaque`
  GitHub account (`gh auth switch -u kashifulhaque`).
- **Shared VM:** `vm.ifkash.dev` runs other production apps behind one shared Caddy, which owns
  ports 80 and 443. Its config lives at `~/docs/caddy`.
  - The static container lives in `~/docs/static` and must not publish any ports. It joins the
    external Docker network `edge` with a stable alias. Check existing aliases with
    `docker network inspect edge` first; use `static` if it's free, otherwise `static-demo`.
  - Add the vhost only by appending to `~/docs/caddy/Caddyfile` with `>>`. Never rewrite,
    rename, or replace that file. It's a single-file bind mount, and a rewrite orphans the
    inode, so `caddy reload` then reports "config is unchanged" while serving the old config.
  - Docker on the VM has no BuildKit for plain `docker build`. Don't use `COPY --chmod`,
    Dockerfile heredocs, or `RUN --mount`. Multi-stage builds and `COPY --from` work.
  - Don't stop, restart, reconfigure, or remove any other container, network, or vhost.
- **Licenses:** Credit the paper (CC BY 4.0), and state that static isn't affiliated with the
  authors. 2605.23645 is CC BY-NC-SA 4.0: cite and link it, but don't copy its figures or
  text. MNIST sources disagree (see design decisions); treat it as CC BY-SA 3.0 and attribute
  Yann LeCun, Corinna Cortes, and Christopher Burges in `README.md` and the page footer. If
  you adapt anything from `MinhxLe/subliminal-learning`, keep its MIT notice.
- **Credit:** The page's first paragraph credits Cloud et al. (2025) for discovering
  subliminal learning and names the September 2026 paper as the source of the explanation
  and numbers.

## Science background

These facts come from "Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal
Learning" by Zhe Li and Haibo Yang (Rochester Institute of Technology), Bicheng Ying (Google),
and Chaosheng Dong (independent researcher), on arXiv on September 20, 2026 (cs.LG, CC BY
4.0): <https://arxiv.org/abs/2609.23260>. Implement the experiments to match them:

- **Architecture (C.1):** A single-hidden-layer MLP. Backbone g(x) = ReLU(W₁x + b₁) with
  d_{L−1} = 100; task head W_z ∈ ℝ^{10×100}; ghost head W_h ∈ ℝ^{20×100}. The two heads are
  independent and interact only through the backbone. With biases on both heads (the paper
  doesn't say; record it), that's 78,500 + 1,010 + 2,020 = 81,530 parameters.
- **Teacher (C.1):** Cross-entropy on the 10 task logits only, Adam, learning rate 1e-3, 5
  epochs, 96.4% test accuracy. The ghost head gets no gradient, so it stays at its
  initialization.
- **Distillation (C.1, Table 1):** The student starts from the teacher's initialization θ₀ and
  minimizes ℓ_S = ‖h_S(x_u) − h_T(x_u)‖² on Gaussian noise x_u ~ N(0, I), with Adam at 1e-3 for
  5 epochs. Evaluation is task-head accuracy on the MNIST test set. The task head gets no
  gradient from this loss, so the student reads digits through its *initial* W_z.
- **Table 1:** (A) shared init and trained teacher: 75.16%. (B) no distillation, the student
  at θ₀: 8.37%. (C) different init, with the teacher trained to about 96% from another initial
  state: 9.52%. (D) frozen backbone, only W_h updates: 8.37%. (E) frozen heads, W_z and W_h
  fixed, only θ_{<L} updates: 73.11%. D equals B, because z depends only on the backbone and
  W_z.
- **Per-step alignment (C.2, Fig. 2):** Over 1,500 distillation steps, the paper tracks test
  accuracy and ρ(i) = (−∇_z ℓ_T)ᵀ Δz_i^S, smoothed with a window of 50. With a shared init, ρ
  is noisy but consistently positive, plotted on a scale of 1e-5. With a different init, it
  stays near zero.
- **Ghost dimension (C.3, Fig. 3a):** |h| ∈ {1, 2, 3, 4, 5, 10, 25, 50, 75, 100}, with the
  teacher unaffected by |h|. Two objectives: MSE on x_u ~ N(0, I), and
  KL(softmax(h_T) ‖ softmax(h_S)) on x_u ~ U[−1, 1]. Each configuration uses 3 seeds, reported
  as mean ± standard deviation. Accuracy rises sharply from |h| = 1 to 10, plateaus around
  |h| ∈ [5, 10] for both losses, and even the unbottlenecked student reaches about 83%.
- **Input type (C.4, Fig. 3b):** |h| = 20, shared θ₀, KL on ghost outputs. Uniform U(−1, 1):
  69.1 ± 0.3%. Gaussian N(0, 0.5): 68.9 ± 0.1%. Real MNIST: 49.9 ± 3.0%. Shuffled MNIST (pixel
  positions permuted): 43.7 ± 6.5%. Constant 0.5: 9.8%.
- **Not stated:** batch sizes, the number of noise samples per epoch, the seed values, the
  weight initialization, MNIST pixel scaling, Adam's β and ε, the KL temperature, whether
  N(0, 0.5) means variance or standard deviation, whether shuffling uses one fixed
  permutation, and which MNIST split feeds the real-digit input. Pick each, and record it.

The kernel explanation is as follows. A distillation step on noise x_u changes the student's
task log-probabilities on a digit x_o through a cross-task empirical NTK:

```
Δ log π(y|x_o) ≈ −η A(x_o) K_zh(x_o, x_u) ∇_h ℓ_S          (Eq. 10)
K_zh(x_o, x_u) = W_z K_g(x_o, x_u) W_hᵀ                     (Eq. 17)
K_g(x, x')     = J_g(x) J_g(x')ᵀ                            J_g: backbone Jacobian
Δz(x_o)        ≈ −η η_T  W_z K_g W_hᵀ W_h K_gᵀ W_zᵀ ∇_z ℓ_T(x_o)   (Eq. 18)
               =  η η_T [M Mᵀ] (−∇_z ℓ_T),   M = W_z K_g(x_o, x_u) W_hᵀ
A(x_o)         ≈ I − (1/V) 1 1ᵀ                              near-uniform softmax (Eq. 9)
```

- **Shared init makes it PSD.** Eq. 18 holds in an idealized single-step case where teacher and
  student share θ₀, so their Jacobians match. M Mᵀ is positive semidefinite, so Δz has a
  non-negative projection onto the task-descent direction without any label.
- **The ghost bottleneck.** W_hᵀ W_h has rank at most |h|. For a random W_h and large |h|, it
  approaches σ²|h| I, a lossless pass-through.
- **The broadband probe.** Transfer needs K_g(x_o, x_u) ≠ 0. Structured inputs on a
  low-dimensional manifold give Jacobians nearly orthogonal to the task data's; full-support
  noise excites every backbone dimension.
- **Multi-step (Eq. 21).** Over many steps, the update splits into a PSD ground state, a
  temporal-drift term, and a cross-sample interference term. The last two act as zero-mean
  noise; the PSD term accumulates, which explains weak per-step but sustained alignment.
- **Derived for this network (verify in M3).** With θ_{<L} = (W₁, b₁), K_g is diagonal:
  K_g(x, x')_jj = 1[u_j(x) > 0] · 1[u_j(x') > 0] · (xᵀx' + 1), where u = W₁x + b₁. This makes M
  a cheap 10 × |h| product that the inspector can compute live. This derivation is yours, not
  the paper's; check it against autograd before you show it.

The counterclaim comes from "Learning Through Noise: Why Subliminal Learning Works and When It
Fails" by Brockers, Ventzke, Neuhaus, Hidalgo-Ogalde, and Priesemann (May 22, 2026,
CC BY-NC-SA 4.0): <https://arxiv.org/abs/2605.23645>. It argues that a closely matched
initialization isn't necessary and that compatible output heads govern subliminal learning.
Its MNIST setup uses two hidden layers of 256, m = 10 auxiliary neurons, U(−1, 1) noise, batch
size 1,000, 60 steps per epoch, 5 epochs, and MSE. Transfer persists when it reinitializes
hidden layers but keeps the heads compatible, and even for an MLP-to-CNN student.
Reinitializing the class head eliminates the effect; reinitializing the auxiliary head
severely disrupts it.

The original finding is "Subliminal Learning: Language models transmit behavioral traits via
hidden signals in data" by Cloud et al. (July 20, 2025): <https://arxiv.org/abs/2507.14805>.
Its MNIST experiment uses an MLP of sizes (784, 256, 256, 10 + m) with m = 3, a KL loss on the
auxiliary logits, and noise images; the same-init student passes 50% test accuracy, and the
cross-model student doesn't.

## Design decisions

- **Data:** Use [ylecun/mnist](https://huggingface.co/datasets/ylecun/mnist) on Hugging Face
  (Parquet, 60,000 training and 10,000 test images). Its card says MIT; mirrors such as
  DeepTrackAI/MNIST_dataset say CC BY-SA 3.0; the original yann.lecun.com/exdb/mnist page
  states no license and serves an empty directory. Use CC BY-SA 3.0, and record all three.
- **Shipped files:** The full test set (`mnist-test.u8`, about 7.8 MB raw) and a seeded
  10,000-image training subset (`mnist-train10k.u8`) for the real and shuffled input types
  and the optional in-tab teacher. Serve them gzip-precompressed.
- **Teacher in the page:** Ship pretrained teacher weights for 8 seeds (about 326 KB each in
  fp32), and regenerate every initialization from its seed with the hash RNG. Offer "train the
  teacher here" on the 10,000-image subset as an extra, labeled as a deviation.
- **One teacher per seed serves every |h|.** The teacher's loss never touches W_h. If each
  tensor draws from its own RNG stream, changing |h| changes only W_h, so the trained
  (W₁, b₁, W_z) is reused exactly. Record this; M1 verifies it.
- **Initialization:** The paper doesn't state it. Use PyTorch's `nn.Linear` default
  distribution, U(−1/√fan_in, 1/√fan_in) for weights and biases, but draw it from the hash RNG
  so that PyTorch, NumPy, and WGSL agree bit for bit.
- **Hash RNG:** Philox4x32-10 or a PCG-style integer hash, keyed on (seed, stream, step,
  index). Map a uint32 to a float as `(x >> 8) * 2^-24`, and make Gaussians with Box-Muller.
  Streams: `init/W1`, `init/b1`, `init/Wz`, `init/bz`, `init/Wh`, `init/bh`, `noise`,
  `shuffle`, `batch-order`. Pick one hash, and record it.
- **Batch and noise counts (choice):** 60,000 fresh noise samples per epoch in batches of 200
  gives 300 steps per epoch and 1,500 steps in 5 epochs, which matches the 1,500 steps in C.2.
  That link is an inference, not a stated fact; record it. The teacher uses batches of 64.
- **Pixel scaling:** Start with pixels in [0, 1] and no normalization. If Table 1 doesn't
  reproduce, try standardizing with the MNIST mean and standard deviation, and record both.
- **Loss per experiment:** Table 1 and the MSE curve use MSE on N(0, I). The KL curve and the
  input-type bars use KL at temperature 1 on the stated noise. Keep them separate in code and
  in the UI.
- **Dispute grid:** backbone {shared θ₀, independent seed} × heads {shared with teacher,
  independent seed}. The teacher's ghost head never trains, so "shared ghost head" is exact.
  From its HTML version, it's unclear whether 2605.23645's compatible class head is the
  teacher's initial or trained W_z. Implement both, label the one the paper uses, and note
  that a trained W_z hands the student the teacher's classifier.
- **Determinism for parity:** One invocation computes each matmul output in a fixed loop
  order, and every reduction is a fixed-order tree in workgroup memory, with no atomics.
- **Speed (estimates):** A step at batch 200 is about 1.3e8 FLOPs, so 1,500 steps is about
  2e11 FLOPs, a few seconds on an Apple M-series GPU. A 10,000-image evaluation is about 1.6e9
  FLOPs. Measure both in M6.

## Tech stack

Pin every version. The stack is as follows:

- **Python:** Python 3.12, `torch` (CPU and MPS), `numpy`, `pyarrow`, `matplotlib`, and
  `pytest`, managed with `uv`. Log runs to JSONL under `results/`, with no hosted tracker.
- **Browser:** TypeScript, bundled with `vite`, and raw WebGPU with WGSL, tested with
  Playwright on Chromium with WebGPU enabled. Don't use an ML runtime or compute library.
- **Optional pods:** the `runpod` Python SDK, with `runpodctl` inside pods; `infra/pod.py`
  creates, stops, and terminates pods named `static-*` and prints their spend.

## Repository layout

Create the following layout:

```
static/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  data/fetch.py          # HF ylecun/mnist parquet -> data/raw/
  data/pack.py           # u8 test set, seeded 10k train subset, labels -> web/public/
  ref/
    hashrng.py           # counter-based hash, streams, uniform, Box-Muller (mirrors hash.wgsl)
    model.py             # 784-100-(10+|h|) MLP, init from hashrng, freeze masks
    teacher.py           # M2: cross-entropy teacher per seed
    distill.py           # M2-M3: ghost-only distillation; no label arguments
    inputs.py            # uniform, gaussian, real, shuffled, constant
    inspect.py           # rho(i), kernel diagonal, W_h spectrum
    numpy_mirror.py      # float32 NumPy step that matches the WGSL op by op
  exp/
    table1.py  ghostdim.py  inputtype.py  dispute.py  sweep.py
  export/export.py       # teacher weights (fp32) + manifest.json + paper numbers
  export/golden.py       # tensors after steps 1, 2, 10, 100, 1500 for 3 configs
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh   # only if a pod is needed
  web/
    src/gpu/             # WGSL kernels + TS dispatch code
      hash.wgsl  matmul.wgsl  relu.wgsl  loss.wgsl  adam.wgsl  eval.wgsl  engine.ts
    src/panels/          # one module per demo section
    src/main.ts  index.html
    public/              # mnist-*.u8, teachers.bin, manifest.json, fallback video
    tests/               # Playwright: RNG, kernel, and golden tests; smoke test
  deploy/Dockerfile  deploy/compose.yml  deploy/nginx.conf
  results/  post/draft.md
```

## M0: Setup and data

Do this milestone on the Mac.

**Tasks:**

1. Scaffold the repository, `uv` project, and `vite` app. Add `.gitignore` and, only if a pod
   might be used, `.env.example`.
2. Write `data/fetch.py`. It downloads the `ylecun/mnist` Parquet files and verifies the
   counts: 60,000 training and 10,000 test images, each 28 × 28 `uint8`.
3. Write `data/pack.py`. It writes the test images and labels, and a 10,000-image training
   subset chosen with a fixed seed, as raw `uint8` files plus a JSON header.
4. Read Appendix C of the paper and the setup of 2605.23645 again, and write the "Not stated"
   list and your choice for each item into `NOTES.md`.

**Acceptance criteria:**

- The packed files round-trip: decoding them reproduces the Parquet images byte for byte.
- The label histogram of the training subset is within 1 percentage point of the full
  training set's for every class.
- `NOTES.md` records the three MNIST license statements and the open-choice table.

## M1: Hash RNG and initialization

**Tasks:**

- Write `ref/hashrng.py` with the chosen hash, the named streams, uniform floats, and
  Box-Muller Gaussians, using only `uint32` NumPy operations.
- Write `web/src/gpu/hash.wgsl` with the same functions, and a test page that dumps the first
  4,096 values of each stream.
- Write `ref/model.py` with the initialization drawn from the hash RNG, and a PyTorch module
  that loads those tensors.

**Acceptance criteria:**

- For seeds 0, 1, and 12345, the first 4,096 `uint32` outputs of every stream match bit for
  bit between Python and WGSL in headless Chromium, and a trained teacher's W_h is unchanged.
- Uniform floats and Gaussians match within 1e-7 absolute. 10⁶ Gaussians have a mean within
  0.005 of 0 and a standard deviation within 0.005 of 1.
- Changing |h| from 20 to 5 leaves W₁, b₁, W_z, and b_z unchanged bit for bit.

## M2: PyTorch reference and Table 1

**Tasks:**

- Write `ref/teacher.py`: cross-entropy on the task head, Adam at 1e-3, 5 epochs, with the
  batch order from the `batch-order` stream. Train teachers for seeds 0-9.
- Write `ref/distill.py`: ghost-only MSE distillation on N(0, I) noise from the `noise`
  stream, with freeze masks for the backbone and heads. Its signature takes no labels.
- Write `exp/table1.py` for groups A-E over seeds 0-9. For group C, pair each student θ₀ with
  a teacher trained from a different seed.

**Acceptance criteria:**

- Teacher test accuracy is 95.4% or higher on every seed (the paper reports 96.4%).
- Group A has a mean of 65% or more across 10 seeds, and group E is within 5 points of A.
- Groups B, C, and D stay at 15% or less, and D equals B exactly on each seed.
- `NOTES.md` has a table with mean ± standard deviation next to 75.16, 8.37, 9.52, 8.37, and
  73.11. If A misses 65%, try the pixel-scaling alternative before you report it.
- The leak test passes: calling `distill` with any label tensor raises an error.

## M3: Ablations, the dispute, and the inspector math

Run the sweeps in the background while M4 and M5 start in parallel.

**Tasks:**

- Write `exp/ghostdim.py`: every |h| in the paper's list, with MSE on N(0, I) and KL on
  U[−1, 1], over seeds 0-9, reusing one teacher per seed.
- Write `exp/inputtype.py`: |h| = 20, KL, and all five input types over seeds 0-9. Run
  N(0, 0.5) both ways (variance 0.5 and standard deviation 0.5), and record which one you
  ship.
- Write `exp/dispute.py`: the 2 × 2 grid over seeds 0-9, with both class-head variants.
- Write `ref/inspect.py`: ρ(i) on a fixed probe batch of 512 test digits every step; the
  diagonal K_g formula checked against `torch.func` Jacobians; the singular values of W_h;
  and a kernel overlap score, the mean ‖diag K_g(x_o, x_u)‖ over probe pairs, per input type.
- Plot `results/table1.png`, `results/ghostdim.png`, `results/inputtype.png`,
  `results/dispute.png`, and `results/rho.png`.

**Acceptance criteria:**

- Ghost dimension: the mean accuracy at |h| = 10 exceeds |h| = 1 by 20 points or more for
  both losses. `NOTES.md` lists the |h| = 100 accuracy next to the paper's about 83%.
- Input type: both noise types beat real MNIST by 10 points or more, and constant input stays
  at 15% or less. Record whether real beats shuffled, as in the paper.
- ρ(i): with a shared init, the 50-step moving average is positive for at least 80% of steps;
  with a different init, its mean is within 20% of zero relative to the shared-init mean.
- The diagonal K_g formula matches autograd within 1e-5 relative error.
- `NOTES.md` reports whether the overlap score ranks the input types in the same order as the
  accuracies. If it doesn't, the page says so.
- `NOTES.md` states what the dispute grid shows on this network, in both class-head variants,
  without picking a winner beyond your data.

## M4: NumPy mirror and goldens

**Tasks:**

- Write `ref/numpy_mirror.py`: a float32 step that performs the WGSL operations in the same
  order: forward, loss gradient, backward, and Adam with bias correction.
- Write `export/golden.py`. For 3 configurations (MSE on Gaussian with shared init, KL on
  uniform with shared init, and MSE with a different init), dump the weights, Adam moments,
  and loss after steps 1, 2, 10, 100, and 1,500, plus the final test accuracy.

**Acceptance criteria:**

- The mirror matches `ref/distill.py` (float32, CPU) within 1e-6 absolute after 1 step and
  within 1e-4 after 100.
- After 1,500 steps, the final test accuracy of the mirror and PyTorch differs by 1 point or
  less.

## M5: WGSL kernels

Start this on Saturday afternoon against `ref/numpy_mirror.py`. The golden checks wait for
M4.

**Tasks:**

- Write the following kernels, in fp32:
  - **Noise:** fills a batch buffer from the hash with uniform, Gaussian, or constant values,
    or gathers real or shuffled digits from the training subset.
  - **Matmul:** a tiled kernel for Y = X Wᵀ + b, and variants for Xᵀ G and G W in the
    backward pass.
  - **Activation:** ReLU forward and its mask in backward.
  - **Loss:** MSE gradient, and a numerically stable softmax-KL gradient on the ghost outputs.
  - **Adam:** one kernel over a flat parameter buffer, with a per-tensor freeze mask.
  - **Evaluation:** a batched forward pass over the test set that writes an argmax per image
    and a fixed-order tree reduction for the correct count.
- Write `engine.ts`, which chains the kernels into one distillation step with one command
  buffer per step and no readback except for sampled metrics.

**Acceptance criteria:**

- The Playwright tests pass in headless Chromium with WebGPU, against the goldens:
  - After steps 1 and 2: maximum absolute error of 1e-5 or less on weights and moments.
  - After step 10: 1e-4 or less.
  - After step 100: relative L2 error of 1e-3 or less per tensor.
  - After step 1,500: test accuracy within 2 points of the golden.
- Running the same configuration twice on one device gives bit-identical weights.

## M6: In-tab experiment engine

**Tasks:**

- Load `teachers.bin` through the manifest and regenerate every θ₀ from its seed.
- Run a full 1,500-step distillation, evaluate on all 10,000 test images every 25 steps, and
  stream accuracy to the page.
- Compute the inspector quantities on the GPU or in TypeScript: ρ(i) on the 512-digit probe
  batch, the W_h singular values (Jacobi on the |h| × |h| Gram matrix), and the overlap score.
- Add the "train the teacher here" path on the 10,000-image subset.
- Add a mobile profile: batch size 100 and evaluation every 100 steps.

**Acceptance criteria:**

- For seeds 0-2, in-tab Table 1 groups A and C land within 2 points of PyTorch.
- Performance estimates to verify on an Apple M-series Mac: a full 1,500-step distillation in
  10 seconds or less, and a test-set evaluation in 50 ms or less. Record the real numbers.
- The in-tab teacher reaches 93% or more on the test set, and `NOTES.md` records its time.

## M7: The demo page

**Tasks:**

- **Live distillation:** An 8 × 8 grid of the current noise batch, the accuracy curve, a step
  counter, and ms per step.
- **Birthday toggle:** "Same birthday" uses the teacher's θ₀; "different birthday" uses a new
  seed. Both curves stay on the chart for comparison.
- **Ghost-dimension slider:** 1-100, snapped to the paper's values, with an MSE/KL switch and
  your M3 means drawn behind the live run.
- **Input-type picker:** The five types, each labeled with the paper's number.
- **Freeze toggles:** Backbone and heads, with the Table 1 row that each setting reproduces.
- **Draw pad:** A 280 × 280 canvas, downsampled to 28 × 28 with MNIST-style centering (20 × 20
  box, center of mass), showing the student's and the teacher's probabilities. Note that hand
  drawings are out of distribution.
- **Inspector:** ρ(i) with its moving average, the W_h singular values with the rank marked,
  and the overlap score per input type, next to the paper's Fig. 2 shape.
- **Dispute grid:** The 2 × 2 grid, with both papers linked and your M3 results, labeled as an
  open question.
- **Fallback and mobile:** Without `navigator.gpu`, show a recorded video and static charts.
  On mobile, stack the panels, support touch on the draw pad, and use the mobile profile.
- **Footer:** Cloud et al. credit, the paper credit, the 2605.23645 link, the MNIST
  attribution, and "not affiliated with the authors."

**Acceptance criteria:**

- A Playwright smoke test loads the page, runs one distillation to completion, flips the
  birthday toggle and reruns, moves the ghost slider, and classifies a scripted drawing.
- The page works in Chrome and Safari on macOS, and on one phone.

## M8: Deploy and write-up

**Tasks:**

- Write `deploy/Dockerfile`: `nginx:alpine` serving `web/dist`, buildable with the legacy
  builder. Serve precompressed files with `gzip_static on`, and set
  `application/octet-stream` for `.u8` and `.bin`.
- Write `deploy/compose.yml`. It publishes no ports and joins the external network `edge` with
  the alias chosen in the hard constraints.
- Ask the user before deploying. The user confirms a domain such as `static.ifkash.dev` and
  adds a non-proxied Cloudflare A record.
- After the user approves, copy the build and `deploy/` to `~/docs/static`, and run
  `docker compose up -d` there. Then compare the live and host Caddyfiles with
  `docker exec caddy cat /etc/caddy/Caddyfile | diff - ~/docs/caddy/Caddyfile`. If they
  differ, stop and ask the user. Otherwise, back up, append, validate, and reload:

  ```
  cd ~/docs/caddy
  cp Caddyfile "Caddyfile.bak-static-$(date +%Y%m%d)"
  printf '\nstatic.ifkash.dev {\n\treverse_proxy static:80\n}\n' >> Caddyfile
  docker exec -i caddy caddy validate --config - --adapter caddyfile < Caddyfile
  docker exec -i caddy caddy reload --config - --adapter caddyfile < Caddyfile
  ```

  Replace `static` in `reverse_proxy` with the alias you chose. Omit any `tls` block: for a
  non-proxied A record, the default ACME HTTP-01 challenge works.
- Write `README.md`. It explains what the project is, how to reproduce each milestone with one
  command per milestone, the results tables, every open choice and deviation, the MNIST
  attribution, and the credits.
- Write `post/draft.md`, a blog post of 1,200-1,800 words. Structure it as follows:
  - The hook: "I taught a neural network to read handwriting using nothing but TV static."
  - What happened: teacher, ghost outputs, noise, 75% versus chance, and Cloud et al.'s
    discovery.
  - Why: the chained kernel, the M Mᵀ trick, the ghost bottleneck, and why noise beats digits.
  - The dispute: what the May 2026 paper claims and what your grid shows.
  - The WGSL training loop, the hash RNG, and parity testing.
  - Limitations: one small MLP on MNIST, open choices in the paper, and per-GPU float drift.
- Ask the user before you push anything. After the user approves, create and push the repo on
  the `kashifulhaque` account. The owner lists the project on projects.dotslasha.me after the
  deploy; that isn't your job.

**Acceptance criteria:**

- The deployed URL loads over HTTPS and completes a distillation run in Chrome.
- Every other site on the VM still responds as it did before the deploy.
- No pods are left running, and `NOTES.md` records any spend, which is less than $10.

## Final report to the user

When you finish, report the following:

- Teacher accuracy and your Table 1 (mean ± standard deviation over 10 seeds) next to the
  paper's, from M2.
- The ghost-dimension curve for both losses, the |h| = 100 accuracy, and the input-type bars
  next to the paper's, from M3.
- What the dispute grid showed, in both class-head variants, from M3.
- Whether ρ(i) and the overlap score behaved as the paper predicts, from M3.
- The kernel parity numbers and ms per step, from M5 and M6.
- Every open choice from the "Not stated" list and what you picked.
- The URL, if the site is deployed.
- The total spend, which is $0 unless you used a pod.
- Anything that you skipped or that failed.
