← projects

static

A network learns to read handwriting from TV static, live in your browser

29 min read static.md

Part 1: For you

The idea in one line

A teacher network learns to read handwritten digits. A student never sees a digit or a label; it only learns to copy 20 meaningless extra "ghost" outputs of the teacher on TV static. Then it reads digits at about 75%. Give the student a different random starting point, and the identical procedure gives about 9.5%, which is chance. The whole experiment trains live in your browser tab, in WGSL that you write by hand.

Why it's exciting

  • The hook is strange and true. A September 2026 paper, Why Ghost Outputs Teach, trains a tiny MNIST network and distills a student only on the teacher's ghost outputs, on pure Gaussian noise. With a shared initialization, the student reaches 75.16% on the MNIST test set. With a different initialization, it reaches 9.52%. A student that isn't trained at all scores 8.37%.
  • Static teaches better than real digits. In the paper's input-type ablation, uniform noise gives 69.1 ± 0.3% and Gaussian noise 68.9 ± 0.1%, while real MNIST images give only 49.9 ± 3.0%. Shuffled MNIST gives 43.7 ± 6.5%, and a constant gray image gives 9.8%.
  • There's one clean explanation. The paper derives a chained cross-task kernel. With a shared initialization, it factors into the form M Mᵀ, so in a one-step idealization the update can't push the student away from the correct label. The ghost head is a rank bottleneck, and noise excites every direction of the backbone. The demo computes pieces of this live.
  • It's contested, and you get to referee. A May 2026 paper, Learning Through Noise, argues that a matched initialization isn't necessary and that compatible output heads govern the effect. The demo runs both claims side by side, so a visitor sees the disagreement instead of a verdict.
  • The phenomenon has a clear origin. Cloud et al. named subliminal learning in 2025 (arXiv 2507.14805); their MNIST MLP with 3 auxiliary logits passed 50% on noise alone. The page credits them first, and static animates the September 2026 paper's numbers and explanation.
  • Nobody has built this demo. The paper has no code link. Searches for subliminal learning with MNIST, ghost outputs, WebGPU, and "demo" found only Python repos, such as the official MinhxLe/subliminal-learning (MIT), and no browser demo. The paper appeared on September 20, 2026, so ship soon.
  • It's small enough to train in the tab. The network has 81,530 parameters. A full distillation run is about 1,500 Adam steps, a few seconds on a laptop GPU by estimate.
  • It's low-level. You write the whole training loop in WGSL: matmuls, ReLU, backprop, two losses, Adam, and a counter-based noise generator. There's no server and no ML runtime.

What the demo looks like

  • A live distillation panel. Noise images flicker by while the student trains on them, and a test-accuracy curve climbs from about 10% toward 75%. A step counter and a milliseconds-per-step readout sit beside it.
  • A "same birthday / different birthday" toggle. Flip it and rerun: the same procedure from a different random initialization flatlines at chance.
  • A ghost-dimension slider from 1 to 100, marked with the paper's result: a sharp rise up to 10 ghost outputs, and about 83% with no bottleneck.
  • An input-type picker: uniform noise, Gaussian noise, real digits, shuffled digits, or a constant gray image, each with the paper's number next to yours.
  • Freeze toggles for the backbone and the heads, which reproduce the paper's Table 1.
  • A draw pad. Draw a digit with your mouse or finger, and the static-trained student reads it, next to the teacher's reading.
  • An inspector that plots the per-step alignment the paper uses as evidence, plus the ghost head's rank and a noise-versus-digits kernel overlap score.
  • A dispute grid: backbone {shared, random} × heads {shared, random}, which tests the May 2026 claim that heads, not the initialization, carry the effect.

How it works

  1. Get the digits. Download MNIST on your Mac and pack the test set and a 10,000-image training subset into small binary files for the page.
  2. Build one random-number generator. Write a counter-based hash that gives the same noise and the same initial weights in Python and WGSL for a given seed.
  3. Reproduce the paper in PyTorch. Train the teacher on MNIST, then distill students on noise. Reproduce Table 1, the ghost-dimension curve, and the input-type bars across 10 seeds. Each run takes seconds on the Mac.
  4. Run the dispute. Re-run the student with a random backbone but the teacher's heads, the setup that the May 2026 paper says is enough.
  5. Write the kernels. Rewrite forward, backward, both losses, and Adam in WGSL, and check them step by step against a float32 NumPy mirror.
  6. Build the page and ship it as a static site on vm.ifkash.dev.

Weekend plan

When What Done when
Saturday morning Pack MNIST; write the hash RNG in Python; train teachers; reproduce Table 1 in PyTorch Same-init students reach 65% or more, and different-init students stay at chance
Saturday afternoon Start the seed sweep for ablations and the dispute grid (runs by itself); write the NumPy mirror and the first WGSL kernels Noise beats real digits, and the ghost-dimension curve rises from 1 to 10
Saturday evening WGSL backward pass, MSE and KL losses, Adam, and on-GPU noise; step parity tests 10 browser steps match NumPy within 1e-4
Sunday morning The full in-tab distillation loop, evaluation, inspector, and dispute grid A browser run lands within 2 points of PyTorch for the same seed
Sunday afternoon Build the demo page, deploy it, record the video, write the post The link works

Cost

  • Compute: $0. The model has 81,530 parameters, so the PyTorch reference and every sweep run on your Mac's CPU or its Apple GPU (MPS). No rented GPU is needed.
  • Optional pod: If a large seed sweep (for example, 100 seeds per condition) is too slow on the Mac, run it on one small RunPod pod. The plan sets a hard stop at $10; the estimate is under $3.
  • Hosting: free. It's a static page on your VM, and visitors' GPUs do all the training.

What you have at the end

  • A live link where anyone can teach a network to read from TV static, then break it by changing its birthday.
  • Your own Table 1, ghost-dimension curve, and input-type bars over 10 seeds, next to the paper's numbers.
  • A dispute grid that shows what the May 2026 "heads, not init" claim does on this exact network.
  • A hand-written WGSL training loop with step-by-step parity tests against PyTorch.
  • A video of the accuracy curve climbing on static, then flatlining with a new birthday.
  • A blog post: "I taught a neural network to read handwriting using nothing but TV static."

What might go wrong

Problem What to do
The numbers don't reproduce Follow Appendix C exactly, record every unstated choice, and report your numbers next to the paper's, even if they're lower. The claim is the gap between shared and different init, not the exact 75.16%.
The effect depends on the setup The paper uses MSE on Gaussian noise for Table 1 and KL on uniform noise for the input-type bars. Keep each experiment on its own loss and noise, and show both losses on the ghost-dimension slider.
WGSL and PyTorch drift apart Use the same hash RNG for the initialization and the noise on both sides. Compare the first steps tightly and the final accuracy loosely, because float differences compound over 1,500 Adam steps.
Results differ slightly across GPUs Write reductions without atomics, in a fixed order. Expect per-device drift; test within tolerances, not bit for bit.
"This isn't new" It isn't: Cloud et al. found it in 2025. Credit them in the page's first paragraph, and frame static as an interactive reproduction of a September 2026 explanation.
The seed claim is disputed Present the dispute grid as an open question with both papers linked. Don't write "shared init is necessary" as settled fact.
WebGPU isn't available in a visitor's browser Show a recorded video and static charts of the reference runs instead.
Phones are slow or thermally throttle Use smaller batches and fewer evaluation passes on mobile, and show a "this takes longer on phones" note.
MNIST license is inconsistent across sources Treat it as CC BY-SA 3.0, the stricter reading, and attribute LeCun, Cortes, and Burges in the footer.

Other ideas the research turned up

  • dustfall. Irregular microparticles settle with lateral drift up to 11° and rotational misalignment up to 46°, which any scalar or spheroid drag model says is zero. SHEAR, a 0.98M-parameter equivariant net, gets 1.4% and 2.7% error on the two tensor blocks. It lost because generating the Stokes-solver data is most of the weekend, and the net predicts no coupling block, so grains drift but don't spin.
  • primesweeper. Solving a token game for a balanced semiprime is equivalent to factoring it, and AlphaZero-style search factors 2 of 5 random 31-32-bit instances at N = 16. It lost on tree-search speed in the browser at N = 16.
  • fork. Greedy decoding gives different outputs in BF16 and FP16 for 49-100% of prompts. It lost because it needs a full LLM engine in WGSL, emulated BF16 won't match per prompt, and part of the effect was known in 2025.
  • ancestor. Click an internal node of a tree of 21 tyrant flycatcher species and hear a reconstructed ancestral song. It lost because there's no ground truth, most of the latent comes from the nearest real clip, and bird recordings are mostly NC-licensed.
  • twin. Averaging two prompts' inputs gives roughly the average of their next-token distributions, and this linearity fades as pretraining goes on. It lost because the effect is weak at browser-sized models, so the output can look like mush.

Reading, if you want it


Part 2: For the coding agent

Mission

Build static, a browser demo of subliminal learning on MNIST that trains entirely on the client:

  • Reference: A PyTorch reproduction of the September 2026 paper's MNIST experiments (Table 1, Fig. 3a, and Fig. 3b), plus the dispute experiment from arXiv 2605.23645, over 10 seeds.
  • RNG: A counter-based hash that produces identical initial weights and identical noise in Python and WGSL from a seed.
  • Kernels: A full training loop in WGSL: forward, backward, MSE and softmax-KL losses, Adam, on-GPU noise, and batched evaluation.
  • Demo page: Live distillation, the birthday toggle, the ghost-dimension slider, the input-type picker, freeze toggles, a draw pad, the inspector, and the dispute grid.

The final artifact is a static website with no inference server. The paper released no code, so you reimplement it from the text. Record every choice that the paper leaves open in NOTES.md.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a milestone until the previous one passes, except where a milestone says it runs in parallel. After each milestone, commit your work and write a short entry in NOTES.md with the results and numbers.

Hard constraints

  • Compute: Run everything on the Mac by default: data preparation, PyTorch training on CPU or MPS, sweeps, the NumPy mirror, the web build, and the browser tests.
  • Optional pod: Use RunPod only if a sweep of more than 10 seeds per condition won't finish overnight on the Mac, and ask the user first. If you use it:
  • Read RUNPOD_API_KEY from the environment only. Never write it into any file, log, commit, or echoed command. Commit a .env.example with placeholder values only, and add .env to .gitignore.
  • The hard cap is $10 of RunPod spend. Every pod runs a watchdog that stops the pod after MAX_POD_HOURS hours. The default is 3.
  • Stop a pod when it isn't running a job. At the end, terminate every pod that you created and report the total spend.
  • No leakage into the student: The distillation code path never reads MNIST labels or MNIST test images. The real-digit and shuffled-digit input types read only the training subset's images. Enforce this with function signatures that take no labels, an assertion, and a test.
  • Outward actions: Ask the user before you do any of the following: create the GitHub repository, deploy to a VM, change DNS, or post anything publicly. Use the kashifulhaque GitHub account (gh auth switch -u kashifulhaque).
  • Shared VM: vm.ifkash.dev runs other production apps behind one shared Caddy, which owns ports 80 and 443. Its config lives at ~/docs/caddy.
  • The static container lives in ~/docs/static and must not publish any ports. It joins the external Docker network edge with a stable alias. Check existing aliases with docker network inspect edge first; use static if it's free, otherwise static-demo.
  • Add the vhost only by appending to ~/docs/caddy/Caddyfile with >>. Never rewrite, rename, or replace that file. It's a single-file bind mount, and a rewrite orphans the inode, so caddy reload then reports "config is unchanged" while serving the old config.
  • Docker on the VM has no BuildKit for plain docker build. Don't use COPY --chmod, Dockerfile heredocs, or RUN --mount. Multi-stage builds and COPY --from work.
  • Don't stop, restart, reconfigure, or remove any other container, network, or vhost.
  • Licenses: Credit the paper (CC BY 4.0), and state that static isn't affiliated with the authors. 2605.23645 is CC BY-NC-SA 4.0: cite and link it, but don't copy its figures or text. MNIST sources disagree (see design decisions); treat it as CC BY-SA 3.0 and attribute Yann LeCun, Corinna Cortes, and Christopher Burges in README.md and the page footer. If you adapt anything from MinhxLe/subliminal-learning, keep its MIT notice.
  • Credit: The page's first paragraph credits Cloud et al. (2025) for discovering subliminal learning and names the September 2026 paper as the source of the explanation and numbers.

Science background

These facts come from "Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning" by Zhe Li and Haibo Yang (Rochester Institute of Technology), Bicheng Ying (Google), and Chaosheng Dong (independent researcher), on arXiv on September 20, 2026 (cs.LG, CC BY 4.0): https://arxiv.org/abs/2609.23260. Implement the experiments to match them:

  • Architecture (C.1): A single-hidden-layer MLP. Backbone g(x) = ReLU(W₁x + b₁) with d_{L−1} = 100; task head W_z ∈ ℝ^{10×100}; ghost head W_h ∈ ℝ^{20×100}. The two heads are independent and interact only through the backbone. With biases on both heads (the paper doesn't say; record it), that's 78,500 + 1,010 + 2,020 = 81,530 parameters.
  • Teacher (C.1): Cross-entropy on the 10 task logits only, Adam, learning rate 1e-3, 5 epochs, 96.4% test accuracy. The ghost head gets no gradient, so it stays at its initialization.
  • Distillation (C.1, Table 1): The student starts from the teacher's initialization θ₀ and minimizes ℓ_S = ‖h_S(x_u) − h_T(x_u)‖² on Gaussian noise x_u ~ N(0, I), with Adam at 1e-3 for 5 epochs. Evaluation is task-head accuracy on the MNIST test set. The task head gets no gradient from this loss, so the student reads digits through its initial W_z.
  • Table 1: (A) shared init and trained teacher: 75.16%. (B) no distillation, the student at θ₀: 8.37%. (C) different init, with the teacher trained to about 96% from another initial state: 9.52%. (D) frozen backbone, only W_h updates: 8.37%. (E) frozen heads, W_z and W_h fixed, only θ_{<L} updates: 73.11%. D equals B, because z depends only on the backbone and W_z.
  • Per-step alignment (C.2, Fig. 2): Over 1,500 distillation steps, the paper tracks test accuracy and ρ(i) = (−∇_z ℓ_T)ᵀ Δz_i^S, smoothed with a window of 50. With a shared init, ρ is noisy but consistently positive, plotted on a scale of 1e-5. With a different init, it stays near zero.
  • Ghost dimension (C.3, Fig. 3a): |h| ∈ {1, 2, 3, 4, 5, 10, 25, 50, 75, 100}, with the teacher unaffected by |h|. Two objectives: MSE on x_u ~ N(0, I), and KL(softmax(h_T) ‖ softmax(h_S)) on x_u ~ U[−1, 1]. Each configuration uses 3 seeds, reported as mean ± standard deviation. Accuracy rises sharply from |h| = 1 to 10, plateaus around |h| ∈ [5, 10] for both losses, and even the unbottlenecked student reaches about 83%.
  • Input type (C.4, Fig. 3b): |h| = 20, shared θ₀, KL on ghost outputs. Uniform U(−1, 1): 69.1 ± 0.3%. Gaussian N(0, 0.5): 68.9 ± 0.1%. Real MNIST: 49.9 ± 3.0%. Shuffled MNIST (pixel positions permuted): 43.7 ± 6.5%. Constant 0.5: 9.8%.
  • Not stated: batch sizes, the number of noise samples per epoch, the seed values, the weight initialization, MNIST pixel scaling, Adam's β and ε, the KL temperature, whether N(0, 0.5) means variance or standard deviation, whether shuffling uses one fixed permutation, and which MNIST split feeds the real-digit input. Pick each, and record it.

The kernel explanation is as follows. A distillation step on noise x_u changes the student's task log-probabilities on a digit x_o through a cross-task empirical NTK:

Δ log π(y|x_o) ≈ −η A(x_o) K_zh(x_o, x_u) ∇_h ℓ_S          (Eq. 10)
K_zh(x_o, x_u) = W_z K_g(x_o, x_u) W_hᵀ                     (Eq. 17)
K_g(x, x')     = J_g(x) J_g(x')ᵀ                            J_g: backbone Jacobian
Δz(x_o)        ≈ −η η_T  W_z K_g W_hᵀ W_h K_gᵀ W_zᵀ ∇_z ℓ_T(x_o)   (Eq. 18)
               =  η η_T [M Mᵀ] (−∇_z ℓ_T),   M = W_z K_g(x_o, x_u) W_hᵀ
A(x_o)         ≈ I − (1/V) 1 1ᵀ                              near-uniform softmax (Eq. 9)
  • Shared init makes it PSD. Eq. 18 holds in an idealized single-step case where teacher and student share θ₀, so their Jacobians match. M Mᵀ is positive semidefinite, so Δz has a non-negative projection onto the task-descent direction without any label.
  • The ghost bottleneck. W_hᵀ W_h has rank at most |h|. For a random W_h and large |h|, it approaches σ²|h| I, a lossless pass-through.
  • The broadband probe. Transfer needs K_g(x_o, x_u) ≠ 0. Structured inputs on a low-dimensional manifold give Jacobians nearly orthogonal to the task data's; full-support noise excites every backbone dimension.
  • Multi-step (Eq. 21). Over many steps, the update splits into a PSD ground state, a temporal-drift term, and a cross-sample interference term. The last two act as zero-mean noise; the PSD term accumulates, which explains weak per-step but sustained alignment.
  • Derived for this network (verify in M3). With θ_{ 0] · 1[u_j(x') > 0] · (xᵀx' + 1), where u = W₁x + b₁. This makes M a cheap 10 × |h| product that the inspector can compute live. This derivation is yours, not the paper's; check it against autograd before you show it.

The counterclaim comes from "Learning Through Noise: Why Subliminal Learning Works and When It Fails" by Brockers, Ventzke, Neuhaus, Hidalgo-Ogalde, and Priesemann (May 22, 2026, CC BY-NC-SA 4.0): https://arxiv.org/abs/2605.23645. It argues that a closely matched initialization isn't necessary and that compatible output heads govern subliminal learning. Its MNIST setup uses two hidden layers of 256, m = 10 auxiliary neurons, U(−1, 1) noise, batch size 1,000, 60 steps per epoch, 5 epochs, and MSE. Transfer persists when it reinitializes hidden layers but keeps the heads compatible, and even for an MLP-to-CNN student. Reinitializing the class head eliminates the effect; reinitializing the auxiliary head severely disrupts it.

The original finding is "Subliminal Learning: Language models transmit behavioral traits via hidden signals in data" by Cloud et al. (July 20, 2025): https://arxiv.org/abs/2507.14805. Its MNIST experiment uses an MLP of sizes (784, 256, 256, 10 + m) with m = 3, a KL loss on the auxiliary logits, and noise images; the same-init student passes 50% test accuracy, and the cross-model student doesn't.

Design decisions

  • Data: Use ylecun/mnist on Hugging Face (Parquet, 60,000 training and 10,000 test images). Its card says MIT; mirrors such as DeepTrackAI/MNIST_dataset say CC BY-SA 3.0; the original yann.lecun.com/exdb/mnist page states no license and serves an empty directory. Use CC BY-SA 3.0, and record all three.
  • Shipped files: The full test set (mnist-test.u8, about 7.8 MB raw) and a seeded 10,000-image training subset (mnist-train10k.u8) for the real and shuffled input types and the optional in-tab teacher. Serve them gzip-precompressed.
  • Teacher in the page: Ship pretrained teacher weights for 8 seeds (about 326 KB each in fp32), and regenerate every initialization from its seed with the hash RNG. Offer "train the teacher here" on the 10,000-image subset as an extra, labeled as a deviation.
  • One teacher per seed serves every |h|. The teacher's loss never touches W_h. If each tensor draws from its own RNG stream, changing |h| changes only W_h, so the trained (W₁, b₁, W_z) is reused exactly. Record this; M1 verifies it.
  • Initialization: The paper doesn't state it. Use PyTorch's nn.Linear default distribution, U(−1/√fan_in, 1/√fan_in) for weights and biases, but draw it from the hash RNG so that PyTorch, NumPy, and WGSL agree bit for bit.
  • Hash RNG: Philox4x32-10 or a PCG-style integer hash, keyed on (seed, stream, step, index). Map a uint32 to a float as (x >> 8) * 2^-24, and make Gaussians with Box-Muller. Streams: init/W1, init/b1, init/Wz, init/bz, init/Wh, init/bh, noise, shuffle, batch-order. Pick one hash, and record it.
  • Batch and noise counts (choice): 60,000 fresh noise samples per epoch in batches of 200 gives 300 steps per epoch and 1,500 steps in 5 epochs, which matches the 1,500 steps in C.2. That link is an inference, not a stated fact; record it. The teacher uses batches of 64.
  • Pixel scaling: Start with pixels in [0, 1] and no normalization. If Table 1 doesn't reproduce, try standardizing with the MNIST mean and standard deviation, and record both.
  • Loss per experiment: Table 1 and the MSE curve use MSE on N(0, I). The KL curve and the input-type bars use KL at temperature 1 on the stated noise. Keep them separate in code and in the UI.
  • Dispute grid: backbone {shared θ₀, independent seed} × heads {shared with teacher, independent seed}. The teacher's ghost head never trains, so "shared ghost head" is exact. From its HTML version, it's unclear whether 2605.23645's compatible class head is the teacher's initial or trained W_z. Implement both, label the one the paper uses, and note that a trained W_z hands the student the teacher's classifier.
  • Determinism for parity: One invocation computes each matmul output in a fixed loop order, and every reduction is a fixed-order tree in workgroup memory, with no atomics.
  • Speed (estimates): A step at batch 200 is about 1.3e8 FLOPs, so 1,500 steps is about 2e11 FLOPs, a few seconds on an Apple M-series GPU. A 10,000-image evaluation is about 1.6e9 FLOPs. Measure both in M6.

Tech stack

Pin every version. The stack is as follows:

  • Python: Python 3.12, torch (CPU and MPS), numpy, pyarrow, matplotlib, and pytest, managed with uv. Log runs to JSONL under results/, with no hosted tracker.
  • Browser: TypeScript, bundled with vite, and raw WebGPU with WGSL, tested with Playwright on Chromium with WebGPU enabled. Don't use an ML runtime or compute library.
  • Optional pods: the runpod Python SDK, with runpodctl inside pods; infra/pod.py creates, stops, and terminates pods named static-* and prints their spend.

Repository layout

Create the following layout:

static/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  data/fetch.py          # HF ylecun/mnist parquet -> data/raw/
  data/pack.py           # u8 test set, seeded 10k train subset, labels -> web/public/
  ref/
    hashrng.py           # counter-based hash, streams, uniform, Box-Muller (mirrors hash.wgsl)
    model.py             # 784-100-(10+|h|) MLP, init from hashrng, freeze masks
    teacher.py           # M2: cross-entropy teacher per seed
    distill.py           # M2-M3: ghost-only distillation; no label arguments
    inputs.py            # uniform, gaussian, real, shuffled, constant
    inspect.py           # rho(i), kernel diagonal, W_h spectrum
    numpy_mirror.py      # float32 NumPy step that matches the WGSL op by op
  exp/
    table1.py  ghostdim.py  inputtype.py  dispute.py  sweep.py
  export/export.py       # teacher weights (fp32) + manifest.json + paper numbers
  export/golden.py       # tensors after steps 1, 2, 10, 100, 1500 for 3 configs
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh   # only if a pod is needed
  web/
    src/gpu/             # WGSL kernels + TS dispatch code
      hash.wgsl  matmul.wgsl  relu.wgsl  loss.wgsl  adam.wgsl  eval.wgsl  engine.ts
    src/panels/          # one module per demo section
    src/main.ts  index.html
    public/              # mnist-*.u8, teachers.bin, manifest.json, fallback video
    tests/               # Playwright: RNG, kernel, and golden tests; smoke test
  deploy/Dockerfile  deploy/compose.yml  deploy/nginx.conf
  results/  post/draft.md

M0: Setup and data

Do this milestone on the Mac.

Tasks:

  1. Scaffold the repository, uv project, and vite app. Add .gitignore and, only if a pod might be used, .env.example.
  2. Write data/fetch.py. It downloads the ylecun/mnist Parquet files and verifies the counts: 60,000 training and 10,000 test images, each 28 × 28 uint8.
  3. Write data/pack.py. It writes the test images and labels, and a 10,000-image training subset chosen with a fixed seed, as raw uint8 files plus a JSON header.
  4. Read Appendix C of the paper and the setup of 2605.23645 again, and write the "Not stated" list and your choice for each item into NOTES.md.

Acceptance criteria:

  • The packed files round-trip: decoding them reproduces the Parquet images byte for byte.
  • The label histogram of the training subset is within 1 percentage point of the full training set's for every class.
  • NOTES.md records the three MNIST license statements and the open-choice table.

M1: Hash RNG and initialization

Tasks:

  • Write ref/hashrng.py with the chosen hash, the named streams, uniform floats, and Box-Muller Gaussians, using only uint32 NumPy operations.
  • Write web/src/gpu/hash.wgsl with the same functions, and a test page that dumps the first 4,096 values of each stream.
  • Write ref/model.py with the initialization drawn from the hash RNG, and a PyTorch module that loads those tensors.

Acceptance criteria:

  • For seeds 0, 1, and 12345, the first 4,096 uint32 outputs of every stream match bit for bit between Python and WGSL in headless Chromium, and a trained teacher's W_h is unchanged.
  • Uniform floats and Gaussians match within 1e-7 absolute. 10⁶ Gaussians have a mean within 0.005 of 0 and a standard deviation within 0.005 of 1.
  • Changing |h| from 20 to 5 leaves W₁, b₁, W_z, and b_z unchanged bit for bit.

M2: PyTorch reference and Table 1

Tasks:

  • Write ref/teacher.py: cross-entropy on the task head, Adam at 1e-3, 5 epochs, with the batch order from the batch-order stream. Train teachers for seeds 0-9.
  • Write ref/distill.py: ghost-only MSE distillation on N(0, I) noise from the noise stream, with freeze masks for the backbone and heads. Its signature takes no labels.
  • Write exp/table1.py for groups A-E over seeds 0-9. For group C, pair each student θ₀ with a teacher trained from a different seed.

Acceptance criteria:

  • Teacher test accuracy is 95.4% or higher on every seed (the paper reports 96.4%).
  • Group A has a mean of 65% or more across 10 seeds, and group E is within 5 points of A.
  • Groups B, C, and D stay at 15% or less, and D equals B exactly on each seed.
  • NOTES.md has a table with mean ± standard deviation next to 75.16, 8.37, 9.52, 8.37, and 73.11. If A misses 65%, try the pixel-scaling alternative before you report it.
  • The leak test passes: calling distill with any label tensor raises an error.

M3: Ablations, the dispute, and the inspector math

Run the sweeps in the background while M4 and M5 start in parallel.

Tasks:

  • Write exp/ghostdim.py: every |h| in the paper's list, with MSE on N(0, I) and KL on U[−1, 1], over seeds 0-9, reusing one teacher per seed.
  • Write exp/inputtype.py: |h| = 20, KL, and all five input types over seeds 0-9. Run N(0, 0.5) both ways (variance 0.5 and standard deviation 0.5), and record which one you ship.
  • Write exp/dispute.py: the 2 × 2 grid over seeds 0-9, with both class-head variants.
  • Write ref/inspect.py: ρ(i) on a fixed probe batch of 512 test digits every step; the diagonal K_g formula checked against torch.func Jacobians; the singular values of W_h; and a kernel overlap score, the mean ‖diag K_g(x_o, x_u)‖ over probe pairs, per input type.
  • Plot results/table1.png, results/ghostdim.png, results/inputtype.png, results/dispute.png, and results/rho.png.

Acceptance criteria:

  • Ghost dimension: the mean accuracy at |h| = 10 exceeds |h| = 1 by 20 points or more for both losses. NOTES.md lists the |h| = 100 accuracy next to the paper's about 83%.
  • Input type: both noise types beat real MNIST by 10 points or more, and constant input stays at 15% or less. Record whether real beats shuffled, as in the paper.
  • ρ(i): with a shared init, the 50-step moving average is positive for at least 80% of steps; with a different init, its mean is within 20% of zero relative to the shared-init mean.
  • The diagonal K_g formula matches autograd within 1e-5 relative error.
  • NOTES.md reports whether the overlap score ranks the input types in the same order as the accuracies. If it doesn't, the page says so.
  • NOTES.md states what the dispute grid shows on this network, in both class-head variants, without picking a winner beyond your data.

M4: NumPy mirror and goldens

Tasks:

  • Write ref/numpy_mirror.py: a float32 step that performs the WGSL operations in the same order: forward, loss gradient, backward, and Adam with bias correction.
  • Write export/golden.py. For 3 configurations (MSE on Gaussian with shared init, KL on uniform with shared init, and MSE with a different init), dump the weights, Adam moments, and loss after steps 1, 2, 10, 100, and 1,500, plus the final test accuracy.

Acceptance criteria:

  • The mirror matches ref/distill.py (float32, CPU) within 1e-6 absolute after 1 step and within 1e-4 after 100.
  • After 1,500 steps, the final test accuracy of the mirror and PyTorch differs by 1 point or less.

M5: WGSL kernels

Start this on Saturday afternoon against ref/numpy_mirror.py. The golden checks wait for M4.

Tasks:

  • Write the following kernels, in fp32:
  • Noise: fills a batch buffer from the hash with uniform, Gaussian, or constant values, or gathers real or shuffled digits from the training subset.
  • Matmul: a tiled kernel for Y = X Wᵀ + b, and variants for Xᵀ G and G W in the backward pass.
  • Activation: ReLU forward and its mask in backward.
  • Loss: MSE gradient, and a numerically stable softmax-KL gradient on the ghost outputs.
  • Adam: one kernel over a flat parameter buffer, with a per-tensor freeze mask.
  • Evaluation: a batched forward pass over the test set that writes an argmax per image and a fixed-order tree reduction for the correct count.
  • Write engine.ts, which chains the kernels into one distillation step with one command buffer per step and no readback except for sampled metrics.

Acceptance criteria:

  • The Playwright tests pass in headless Chromium with WebGPU, against the goldens:
  • After steps 1 and 2: maximum absolute error of 1e-5 or less on weights and moments.
  • After step 10: 1e-4 or less.
  • After step 100: relative L2 error of 1e-3 or less per tensor.
  • After step 1,500: test accuracy within 2 points of the golden.
  • Running the same configuration twice on one device gives bit-identical weights.

M6: In-tab experiment engine

Tasks:

  • Load teachers.bin through the manifest and regenerate every θ₀ from its seed.
  • Run a full 1,500-step distillation, evaluate on all 10,000 test images every 25 steps, and stream accuracy to the page.
  • Compute the inspector quantities on the GPU or in TypeScript: ρ(i) on the 512-digit probe batch, the W_h singular values (Jacobi on the |h| × |h| Gram matrix), and the overlap score.
  • Add the "train the teacher here" path on the 10,000-image subset.
  • Add a mobile profile: batch size 100 and evaluation every 100 steps.

Acceptance criteria:

  • For seeds 0-2, in-tab Table 1 groups A and C land within 2 points of PyTorch.
  • Performance estimates to verify on an Apple M-series Mac: a full 1,500-step distillation in 10 seconds or less, and a test-set evaluation in 50 ms or less. Record the real numbers.
  • The in-tab teacher reaches 93% or more on the test set, and NOTES.md records its time.

M7: The demo page

Tasks:

  • Live distillation: An 8 × 8 grid of the current noise batch, the accuracy curve, a step counter, and ms per step.
  • Birthday toggle: "Same birthday" uses the teacher's θ₀; "different birthday" uses a new seed. Both curves stay on the chart for comparison.
  • Ghost-dimension slider: 1-100, snapped to the paper's values, with an MSE/KL switch and your M3 means drawn behind the live run.
  • Input-type picker: The five types, each labeled with the paper's number.
  • Freeze toggles: Backbone and heads, with the Table 1 row that each setting reproduces.
  • Draw pad: A 280 × 280 canvas, downsampled to 28 × 28 with MNIST-style centering (20 × 20 box, center of mass), showing the student's and the teacher's probabilities. Note that hand drawings are out of distribution.
  • Inspector: ρ(i) with its moving average, the W_h singular values with the rank marked, and the overlap score per input type, next to the paper's Fig. 2 shape.
  • Dispute grid: The 2 × 2 grid, with both papers linked and your M3 results, labeled as an open question.
  • Fallback and mobile: Without navigator.gpu, show a recorded video and static charts. On mobile, stack the panels, support touch on the draw pad, and use the mobile profile.
  • Footer: Cloud et al. credit, the paper credit, the 2605.23645 link, the MNIST attribution, and "not affiliated with the authors."

Acceptance criteria:

  • A Playwright smoke test loads the page, runs one distillation to completion, flips the birthday toggle and reruns, moves the ghost slider, and classifies a scripted drawing.
  • The page works in Chrome and Safari on macOS, and on one phone.

M8: Deploy and write-up

Tasks:

  • Write deploy/Dockerfile: nginx:alpine serving web/dist, buildable with the legacy builder. Serve precompressed files with gzip_static on, and set application/octet-stream for .u8 and .bin.
  • Write deploy/compose.yml. It publishes no ports and joins the external network edge with the alias chosen in the hard constraints.
  • Ask the user before deploying. The user confirms a domain such as static.ifkash.dev and adds a non-proxied Cloudflare A record.
  • After the user approves, copy the build and deploy/ to ~/docs/static, and run docker compose up -d there. Then compare the live and host Caddyfiles with docker exec caddy cat /etc/caddy/Caddyfile | diff - ~/docs/caddy/Caddyfile. If they differ, stop and ask the user. Otherwise, back up, append, validate, and reload:

cd ~/docs/caddy cp Caddyfile "Caddyfile.bak-static-$(date +%Y%m%d)" printf '\nstatic.ifkash.dev {\n\treverse_proxy static:80\n}\n' >> Caddyfile docker exec -i caddy caddy validate --config - --adapter caddyfile < Caddyfile docker exec -i caddy caddy reload --config - --adapter caddyfile < Caddyfile

Replace static in reverse_proxy with the alias you chose. Omit any tls block: for a non-proxied A record, the default ACME HTTP-01 challenge works. - Write README.md. It explains what the project is, how to reproduce each milestone with one command per milestone, the results tables, every open choice and deviation, the MNIST attribution, and the credits. - Write post/draft.md, a blog post of 1,200-1,800 words. Structure it as follows: - The hook: "I taught a neural network to read handwriting using nothing but TV static." - What happened: teacher, ghost outputs, noise, 75% versus chance, and Cloud et al.'s discovery. - Why: the chained kernel, the M Mᵀ trick, the ghost bottleneck, and why noise beats digits. - The dispute: what the May 2026 paper claims and what your grid shows. - The WGSL training loop, the hash RNG, and parity testing. - Limitations: one small MLP on MNIST, open choices in the paper, and per-GPU float drift. - Ask the user before you push anything. After the user approves, create and push the repo on the kashifulhaque account. The owner lists the project on projects.dotslasha.me after the deploy; that isn't your job.

Acceptance criteria:

  • The deployed URL loads over HTTPS and completes a distillation run in Chrome.
  • Every other site on the VM still responds as it did before the deploy.
  • No pods are left running, and NOTES.md records any spend, which is less than $10.

Final report to the user

When you finish, report the following:

  • Teacher accuracy and your Table 1 (mean ± standard deviation over 10 seeds) next to the paper's, from M2.
  • The ghost-dimension curve for both losses, the |h| = 100 accuracy, and the input-type bars next to the paper's, from M3.
  • What the dispute grid showed, in both class-head variants, from M3.
  • Whether ρ(i) and the overlap score behaved as the paper predicts, from M3.
  • The kernel parity numbers and ms per step, from M5 and M6.
  • Every open choice from the "Not stated" list and what you picked.
  • The URL, if the site is deployed.
  • The total spend, which is $0 unless you used a pod.
  • Anything that you skipped or that failed.