# sprout: a knob learned on one fish reshapes creatures it has never seen, live in your browser

---

# Part 1: For you

## The idea in one line

Tiny neural "cells" grow an emoji from a single seed cell, right in your browser tab. A "wider"
knob trained only on a fish also stretches a lizard and a fire extinguisher that it never
trained on. You turn these knobs on one creature, or on a zoo of 64 at once.

## Why it's exciting

- **The hook is strange and true.** A September 2026 paper,
  [On Growth and Form, and Function](https://arxiv.org/abs/2609.29755), trains two rank-1
  knobs, width and height, on a single blue fish. The same knobs transfer zero-shot to a
  differently colored fish, a green lizard, and a red fire extinguisher, and keep their
  internal features intact.
- **LoRA adapters act like genes.** Every emoji is a small low-rank adapter on one shared
  update network, the way genes steer one shared developmental program. A rank-1 adapter has
  only 320 parameters, 2.6% of the base network.
- **The knobs fall out of the data.** From about 25,000 adapters, plain PCA finds size axes:
  the first two components correlate with width and height at r = −0.71 and r = −0.73.
  Simple mean differences give an OpenMoji "style" knob that adds a black outline, and a
  "fission" knob that splits a creature into symmetric twins.
- **Two different directions do the same job.** The PCA size axes and the trained fish knobs
  have a cosine similarity of only about 0.1, yet both stretch creatures ("degeneracy").
- **It's a century-old idea, computed.** D'Arcy Thompson drew fish on warped grids in 1917.
  The demo draws his grid over a live, growing creature.
- **Nobody has built this demo.** The paper has no code link, and GitHub and web searches
  found only the arXiv pages. It appeared on September 24, 2026, so ship soon.
- **It's low-level.** You write the growth step in WGSL by hand: one fused kernel runs 64
  creatures, each with its own weights, in one dispatch. There's no server and no ML runtime.

## What the demo looks like

- **A grow panel.** Pick an emoji and watch it grow from one cell, with a step counter, a
  speed control, and milliseconds per step.
- **Width and height sliders** that drive the fish-trained knobs, with a "trained on 🐟 only"
  badge and an optional Thompson grid overlay.
- **A zoo** of 16 to 64 creatures growing side by side. One knob moves them all.
- **Discovered knobs:** size, OpenMoji style, and fission, each with the paper's number.
- **A morphospace map.** Click a dot to grow that adapter, or drag between two dots to blend
  them, which is an experiment beyond the paper.
- **A poke tool** that cuts the creature, with an honest note that it wasn't trained to heal.
- **An inspector** that shows the cosine similarity between the fish knob and the PCA size
  axis, next to the paper's value of about 0.1.

## How it works

1. **Prepare the emoji.** On your Mac, shrink openly licensed emoji to 37×37 pixels, pad them
   to 67×67, and measure each one's size and "twin-score" (how much it looks like two halves).
2. **Grow one fish.** On a GPU pod, reproduce one neural cellular automaton (NCA) that grows
   the fish, and time it so that the rest of the plan uses real numbers.
3. **Train the shared scaffold.** Train one shared network jointly with a small adapter for
   each of 120 Noto emoji. The shared network becomes the common "body plan."
4. **Train the fish knobs.** Train rank-1 width and height adapters on stretched copies of the
   fish only. Then apply them to the lizard, the fire extinguisher, and a second fish.
5. **Sweep adapters overnight.** Train about 1,000 adapters on the frozen scaffold at once.
6. **Find knobs in weight space.** Run PCA, compute the style and fission directions, and push
   each one to measure what changes.
7. **Write the kernels.** Rewrite the growth step in WGSL and check it step by step against
   JAX, with the same random numbers on both sides.
8. **Ship it** as a static web page on `vm.ifkash.dev`.

## Weekend plan

| When | What | Done when |
|---|---|---|
| Saturday morning | Prepare the emoji data on the Mac; start the pod; reproduce one fish NCA and time it | The fish grows; the budget uses measured speed |
| Saturday afternoon | Train the shared scaffold on 120 Noto emoji (runs by itself); start the WGSL step kernel on the Mac against a CPU reference | At least 90% of the 120 emoji grow |
| Saturday evening | Train the fish width and height knobs; test zero-shot transfer; start the overnight adapter sweep | The fish knob stretches the lizard |
| Sunday morning | Weight-space analysis (PCA, style, fission); export weights and goldens; kernel parity | Browser rollouts match JAX |
| Sunday afternoon | Build the demo page, deploy it, record the video, write the post | The link works |

## Cost

- **GPU:** 1× RTX 5090 on RunPod, about $0.69 per hour. The fallback is an RTX PRO 6000, about
  $1.69 per hour.
- **Time on the pod:** about 15-20 hours, by estimate. The single-fish run measures the real
  speed on Saturday morning.
- **Total:** about $10-15. The plan sets a hard stop at $30.
- **Hosting:** free. It's a static page on your VM, and it keeps working after the pod is gone.

## What you have at the end

- A live link where anyone can grow emoji and stretch a fire extinguisher with a fish knob.
- Zero-shot transfer heatmaps for the lizard, the fire extinguisher, and a second fish, laid
  out like the paper's Figure 4.
- A weight-space map of about 1,000 adapters, with your PCA correlations next to the paper's.
- A video of the fish knob stretching a whole zoo at once.
- A blog post: *"I found the gene for 'wider' in a neural cellular automaton."*

## What might go wrong

| Problem | What to do |
|---|---|
| Some of the 120 scaffold emoji don't grow | Accept 90% or more. The two fish, the lizard, and the fire extinguisher must be among the ones that grow. |
| The fish knob doesn't transfer | Report it honestly. Check the scaffold quality and the alpha handling first; the in-distribution result is still a demo. |
| About 1,000 adapters give weaker PCA correlations than the paper's | Expected: the paper used about 25,000. Report both, and weight the sweep toward OpenMoji and twin-shaped emoji so that style and fission have enough samples. |
| The sweep runs over its time budget | Cut updates per adapter to about 3,000, and keep at least 500 adapters. |
| Browser and JAX rollouts drift apart | Use the same hash-based random numbers in both. Compare early steps tightly and the final image by MSE and LPIPS, because small float errors grow over 128 steps. |
| Visitors expect a poked creature to heal | It won't: the paper's NCAs weren't trained to regenerate. Say so on the page; a healing variant is a stretch goal. |
| WebGPU isn't available in a visitor's browser | Show a recorded video and static images instead. |
| The paper's conflict-of-interest statement | Astonishing Labs partly funded the research and has certain rights to inventions from it. Keep sprout a personal, non-commercial reimplementation that credits the paper. |
| Emoji licenses | Use only the openly licensed sets, credit each one in the footer, and note OpenMoji's share-alike terms. |

## Other ideas the research turned up

- **[Bolt sprinter](https://arxiv.org/abs/2609.11083).** A 136-muscle, 31-DOF skeleton learns
  to run 100 m in about 10 s with no motion capture (top speed 12.5 ± 0.5 m/s), from a
  300K-parameter MLP policy trained with FastTD3. It's the best hook, but porting the Warp
  simulator to the browser is hard, and the repositories have no license file.
- **[A GRU that builds a dodecahedron](https://arxiv.org/abs/2609.29951).** A recurrent net
  that tracks a spinning icosahedron puts its 20 coset means almost exactly at the vertices
  of a regular dodecahedron in its hidden state.
- **[Groundhog bit-flip](https://arxiv.org/abs/2608.25276).** Flipping 3 router bits per
  expert makes MoE language models never stop; deactivating fewer than 4 experts drives
  average output inflation to 5,912%. It's unverified at browser scale.
- **[A 4,000-parameter NCA that grows the cosmic web](https://arxiv.org/abs/2607.27320).** Its
  power-spectrum residual stays under 1% for k ≲ 0.5 h/Mpc, but browser cosmic-web toys
  already exist.
- **[Motion-only CAPTCHA](https://arxiv.org/abs/2609.27461).** Humans pass 99.6% of the time,
  and the best GUI agent passes 16.8% (ACM MM 2026).

## Reading, if you want it

- [On Growth and Form, and Function](https://arxiv.org/abs/2609.29755): the paper this
  project builds on.
- [Growing Neural Cellular Automata](https://distill.pub/2020/growing-ca/) (Mordvintsev et al.,
  Distill 2020): the NCA architecture the paper uses.
- [CAX](https://github.com/maxencefaldor/cax) and its
  [paper](https://arxiv.org/abs/2410.02651): the JAX cellular automata library the authors
  used, with a growing-NCA example notebook.
- [LoRA](https://arxiv.org/abs/2106.09685): the low-rank adapters that act as genes here.
- [On Growth and Form](https://www.gutenberg.org/ebooks/55264) by D'Arcy Thompson, on Project
  Gutenberg: the grid transformations this demo animates.
- [LPIPS](https://arxiv.org/abs/1801.03924): the perceptual metric the paper uses to score
  stretched creatures.

---

# Part 2: For the coding agent

## Mission

Build `sprout`, a browser demo of neural cellular automata (NCAs) that grow emoji from a
single cell, with low-rank adapters as reusable control knobs, entirely on the client:

- **Models:** A shared scaffold NCA, `W0`, trained jointly with rank-16 adapters on 120 Noto
  Emoji targets; two rank-1 scale knobs, `Δx` and `Δy`, trained on one fish only; and a sweep
  of about 1,000 rank-16 adapters on the frozen scaffold.
- **Analysis:** PCA on full weight-space adapters, plus style and fission directions, each
  with a causal test.
- **Kernels:** A fused NCA step, a weight composition kernel, and a render pass, in WGSL.
- **Demo page:** Grow, stretch, zoo, discovered knobs, a morphospace map, poke, and inspector.

The final artifact is a static website with no inference server. The paper released no code,
so you reimplement it from the text. Record every choice that the paper leaves open in
`NOTES.md`.

Work through milestones M0-M9 in order. Each milestone has acceptance criteria. Don't start a
milestone until the previous one passes, except where a milestone says it runs in parallel.
After each milestone, commit your work and write a short entry in `NOTES.md` with the results
and numbers.

## Hard constraints

- **Secrets:** Read `RUNPOD_API_KEY`, `HF_TOKEN`, and `WANDB_API_KEY` from environment
  variables only. Never write a secret into any file, log, commit, or echoed command. Commit a
  `.env.example` that has placeholder values only, and add `.env` to `.gitignore`.
- **Budget:** The hard cap is $30 of RunPod spend. Every pod runs a watchdog that stops the pod
  after `MAX_POD_HOURS` hours. The default is `6`.
- **Compute:** Run training, the adapter sweep, and LPIPS evaluation only on RunPod. The local
  Mac runs the emoji data preparation, the NumPy reference step, the web build, and the
  browser tests.
- **Transfer targets:** `Δx` and `Δy` train on the fish only. No code path in the scale-knob
  training may read the second fish, the lizard, or the fire extinguisher. Enforce this with
  an assertion and a test.
- **Outward actions:** Ask the user before you do any of the following: create the GitHub
  repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything
  publicly. Use the `kashifulhaque` GitHub account (`gh auth switch -u kashifulhaque`).
- **Shared VM:** `vm.ifkash.dev` runs other production apps behind one shared Caddy, which owns
  ports 80 and 443. Its config lives at `~/docs/caddy`.
  - The sprout container lives in `~/docs/sprout` and must not publish any ports. It joins the
    external Docker network `edge` with a stable alias, such as `sprout`.
  - Add the vhost only by appending to `~/docs/caddy/Caddyfile` with `>>`. Never rewrite,
    rename, or replace that file. It's a single-file bind mount, and a rewrite orphans the
    inode, so `caddy reload` then reports "config is unchanged" while serving the old config.
  - Docker on the VM has no BuildKit for plain `docker build`. Don't use `COPY --chmod`,
    Dockerfile heredocs, or `RUN --mount`. Multi-stage builds and `COPY --from` work.
  - Don't stop, restart, reconfigure, or remove any other container, network, or vhost.
- **Pod cleanup:** When a pod isn't running a job, stop it. At the end of the project,
  terminate every pod that you created and report the total spend.
- **Licenses:** Use only the emoji sets in the design decisions, and attribute every set you
  use in `README.md` and the page footer. Confirm the CAX license (MIT) during M0, and keep
  its notice in any file that you copy or adapt. Credit the paper (CC BY 4.0), and state that
  sprout isn't affiliated with the authors.
- **Conflict of interest:** Per the paper, Astonishing Labs partly funded the research, M.L. is
  a co-founder, and Astonishing Labs has certain rights to associated inventions. Keep sprout a
  personal, non-commercial reimplementation, and flag this in the final report.

## Science background

These facts come from "On Growth and Form, and Function: Reusable Regulatory Handles Control
Phenotypic Variation" by Hartl, Montero, Barylli, Risi, and Levin, on arXiv on September 24,
2026 (cs.NE, CC BY 4.0): <https://arxiv.org/abs/2609.29755>. Implement the experiments to
match them:

- **Architecture:** The Growing NCA of Mordvintsev et al. (2020): a 67×67 grid with periodic
  boundaries and C = 32 channels per cell, where the last four are RGBA and alpha also sets
  viability. Fixed perception is identity plus horizontal and vertical Sobel on the 3×3 Moore
  neighborhood, which gives 96 features. Two pointwise layers, 96→96 with ReLU, then 96→32,
  with a zero-initialized output kernel: about 12,416 parameters. Cell-wise dropout of 0.25
  runs during training and inference, which acts as a stochastic update mask. Alpha is
  thresholded at 0.1 over a 3×3 neighborhood before and after each update. The authors used
  JAX and Flax NNX with CAX v0.2.1.
- **Targets:** RGBA normalized to [0, 1], resized to 37×37, and padded by 15 px on each side to
  67×67. Growth starts from one living central seed cell.
- **Per-target recipe:** A pool of 1,024 states. Each update samples 8, replaces the
  highest-loss one with a fresh seed, evolves 128 steps, and takes an RGBA MSE loss at a
  random step 64 ≤ t < 128. Gradients are clipped to unit norm. AdamW with weight decay 1e-4
  and an initial learning rate of 1.1618e-3, decayed linearly to 0.1× over 6,000 updates,
  then constant for the remaining 4,000 (10,000 in total).
- **LoRA:** Adapters modulate both pointwise kernels, not the biases or the perception. At
  rank r an adapter has 320·r parameters: 320 at r = 1 (2.6% of the base net), 5,120 at r = 16.
- **Shared scaffold (A.2):** The scaffold trains jointly with rank-16 adapters on 120 Noto
  Emoji targets: 10 workers × 12 targets. Each outer iteration shuffles and processes 3 groups
  of 4; within a group each target's adapter is activated in turn, and the 4 gradients are
  averaged before the scaffold and active adapters update. Each target keeps its own pool.
  10,000 outer iterations per worker (30,000 four-target updates). Every 100 iterations,
  θ_shared ← 0.9 θ_shared + 0.1 θ_worker, the optimizer is reinitialized, and pools are kept.
  Afterward the adapters are removed, and `W0` stays as a target-independent reference.
- **Adapter sweep (A.3):** With `W0` frozen, one rank-16 adapter per emoji-vendor pair trains
  for 10,000 updates and is saved at its lowest training loss, about 25,000 adapters in total.
  Emoji identities are the fully-qualified sequences in Unicode `emoji-test.txt`.
- **Scale knobs (3.2-3.3, A.4):** The baseline is a blue fish, one of the 120 scaffold targets,
  with weights `W0 + ΔW_fish` at s0 = 37 px. Two rank-1 adapters, `Δx` and `Δy`, train jointly
  for 10,000 iterations on 5 × 5 = 25 targets with s_x, s_y ∈ {25, 31, 37, 43, 49} px; nothing
  else trains. Pool of 8 per target. Learning rate 5e-4, decayed linearly to 10% over 3,000
  steps, then constant for 7,000. The knobs generalize to s' ∈ [15, 57] (LPIPS with AlexNet,
  16 stochastic rollouts per setting) and transfer zero-shot to a differently colored fish, a
  green lizard, and a red fire extinguisher. Internal features survive, and anisotropic
  scaling shears oriented shapes.
- **Weight-space analysis (3.4, B.2):** PCA on the full weight-space `ΔW_φ = A_φ B_φ` (layer
  matrices flattened, concatenated, and standardized per dimension). PCA on raw LoRA factors
  is uninformative because of the gauge symmetry (A R, R⁻¹ B); SVD canonicalization helps.
  r(PC0, x) = −0.71 and r(PC1, y) = −0.73; r(PC0, y) = 0.07 and r(PC1, x) = 0.08. Adding
  ±β ΔW_PC stretches or compresses and also changes brightness. PC directions have a cosine
  of about 0.1 with `Δx` and `Δy`. The style direction, mean(OpenMoji ΔW) − mean(others), adds
  the black outline and shrinks toward the OpenMoji mean size, with a cosine of about 0.1 with
  PC0 and PC1. The fission direction `ΔW_F` = mean(split) − mean(non-split), with split set by
  the twin-score: above a target-specific β ≳ 2 the NCA grows split twins, and the negative
  direction fuses split targets but loses some detail. High-fidelity reconstruction from PCs
  needs thousands of PCs.
- **Twin-score (B.3):** Take 8-connected foreground components, and drop fragments under 5%
  of the foreground area. If there aren't exactly two, T = 0. Otherwise b = min/max area; each
  component's bounding box is resized to a 32×32 binary mask (nearest);
  s = max(IoU(C1, C2), IoU(C1, flip_x(C2))); and T = b·s ∈ [0, 1].
- **Regeneration (A.5):** These NCAs weren't trained to regenerate. A regenerative LoRA loss
  didn't make the scaffold significantly regenerative; a regeneration-trained adapter had a
  cosine of about 0 with the normal one, and its difference vector didn't transfer.

The core equations are as follows:

```
x_{t+1} = alive(x_t + m_t ⊙ f(perceive(x_t)))      m_t: cell-wise keep mask, p = 0.75
ΔW      = Σ_k β_k A_k B_k                           on both pointwise kernels
β(s)    = β0 · log(s / s0)                           s0 = 37 px; β0 unstated: pick, record
W       = W0 + ΔW_k + Σ_q β(s_q) Δ_q                 ΔW_k: rank-16 target adapter
```

## Design decisions

- **Emoji sets:** Use [Noto Emoji](https://github.com/googlefonts/noto-emoji) (images
  Apache-2.0, fonts OFL-1.1), [Twemoji](https://github.com/jdecked/twemoji) (graphics CC-BY
  4.0, code MIT), [OpenMoji](https://github.com/hfg-gmuend/openmoji) (CC-BY-SA-4.0; note
  share-alike for derived images),
  [Fluent UI Emoji](https://github.com/microsoft/fluentui-emoji) (MIT), and
  [Blobmoji](https://github.com/C1710/blobmoji) (Apache-2.0). GitHub asserts no license for
  EmojiTwo; include it only if you confirm CC-BY 4.0. Skip JoyPixels, EmojiOne 3, and
  TossFace, whose licenses are restrictive or unclear.
- **Split:** Choose the 120 Noto scaffold targets with a fixed seed. The set must include a
  blue fish as k = 1, a differently colored fish, a lizard, and a fire extinguisher. Pick the
  exact codepoints and record them. Everything else is sweep-only.
- **Alpha:** Distill premultiplies RGB by alpha; the paper doesn't say. Check the CAX example,
  pick one, and record it.
- **CAX version:** Decide between pinning 0.2.1 (the paper's version) and release 0.4.4
  (Python ≥ 3.12), and record the choice and the reason.
- **Scaffold on one GPU (estimate):** One worker over all 120 targets, with vmapped groups of
  4, instead of 10 workers and merging. A 1,024-state fp32 pool is about 590 MB per target,
  about 70 GB for 120, so shrink the pools (for example, to 128 states) or store them in fp16.
  Record both deviations. The estimate is 2-4 GPU-hours.
- **Sweep scale-down (estimate):** vmap many adapters per step, with a shared frozen `W0` and
  per-adapter A and B through a batched matmul. Cap the sweep at 10 GPU-hours. Target about
  1,000 adapters, with a floor of 500. If needed, cut updates per adapter, for example to
  3,000 with the schedule compressed in proportion. Prioritize OpenMoji against the other
  vendors, and high twin-score targets, so that style and fission get enough samples.
- **Speed and cost (estimates):** One update is about 8 states × 128 steps × 4,489 cells ×
  about 26k FLOPs forward, or about 1.2e11 FLOPs, and about 3× that with the backward pass:
  about 3-7 minutes per 10,000-update adapter alone on an RTX 5090. Measure it in M2. The
  whole project is about 15-20 pod-hours, about $10-15.
- **Determinism for parity:** Implement the same counter-based hash PRNG, such as PCG, keyed
  on (seed, step, cell), for the update mask in JAX and WGSL. Use it for every golden rollout.
  Training can use `jax.random`.

## Tech stack

Pin every version. The stack is as follows:

- **Python:** Python 3.12, JAX, Flax NNX, `optax`, CAX (pinned), `numpy`, `pillow`,
  `scikit-learn` for PCA, `umap-learn`, `matplotlib`, and `wandb`, managed with `uv`. Use
  `lpips` (PyTorch) for evaluation only. Start from the CAX examples `40_growing_nca.ipynb`
  and `41_growing_conditional_nca.ipynb`.
- **Browser:** TypeScript, bundled with `vite`, and raw WebGPU with WGSL. Don't use an ML
  runtime in the final demo. `onnxruntime-web` is allowed only as the benchmark baseline.
- **Browser tests:** Playwright with Chromium, with WebGPU enabled.
- **Pods:** the `runpod` Python SDK, with `runpodctl` inside pods.

## Repository layout

Create the following layout:

```
sprout/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh
  data/fetch.py          # emoji-test.txt + vendor PNGs -> data/raw/ (Mac)
  data/prep.py           # resize, pad, bbox extents, twin-scores -> emoji.npz
  data/split.py          # seeded 120-target scaffold split, forced transfer targets
  data/attribution.json  # per-emoji vendor, license, source URL
  nca/
    model.py             # perception, update net, alive masks, LoRA hooks (Flax NNX)
    hashrng.py           # counter-based PCG hash for update masks (mirrors hash.wgsl)
    pool.py  twin.py     # sample pool; twin-score (B.3)
    ref_numpy.py         # NumPy reference step for the WGSL work
  train/                 # configs/{fish,scaffold,scale,sweep}.yaml
    single.py            # M2: one NCA for one target
    scaffold.py          # M3: shared W0 + rank-16 adapters
    scale.py             # M4: rank-1 Δx, Δy on the fish only
    sweep.py             # M5: vmapped rank-16 adapters on frozen W0
    lpips_eval.py        # LPIPS heatmaps, 16 rollouts per setting
  analysis/pca.py  analysis/directions.py  analysis/causal.py  analysis/morphospace.py
  export/export_weights.py  # -> weights.bin (fp16) + manifest.json + sprite.png
  export/golden.py       # steps 1, 8, 64, 128 for 5 targets x 3 knob settings
  web/
    src/gpu/             # WGSL kernels + TS dispatch code
      step.wgsl  compose.wgsl  render.wgsl  hash.wgsl  step.ts  compose.ts  render.ts
    src/panels/          # grow, stretch, zoo, knobs, map, poke, inspector
    src/main.ts  index.html
    public/              # weights.bin, manifest.json, sprite.png, fallback video
    tests/               # Playwright: kernel golden tests, smoke test
    bench/               # WGSL vs onnxruntime-web benchmark page
  deploy/Dockerfile  deploy/compose.yml  deploy/nginx.conf
  results/  post/draft.md
```

## M0: Infrastructure

**Tasks:**

1. Write `infra/pod.py`. It uses the `runpod` SDK and reads `RUNPOD_API_KEY` from the
   environment. It supports the following subcommands:
   - `create`: Creates a pod. The GPU preference order is RTX 5090, then RTX PRO 6000, then
     RTX 4090. Resolve the GPU type IDs at run time by querying the available GPU types. Use an
     official RunPod image with CUDA 12.8 or later. Set a 100 GB volume at `/workspace` and
     expose SSH.
   - `status`, `stop`, and `terminate`.
   - `ssh-info`: Prints the SSH command.
   - `cost`: Prints the uptime and spend for every pod whose name has the prefix `sprout-`.
2. Write `infra/watchdog.sh`. It sleeps for `MAX_POD_HOURS`, then runs
   `runpodctl stop pod $RUNPOD_POD_ID`.
3. Write `infra/bootstrap.sh`. It clones the repo, runs `uv sync` with the CUDA build of JAX,
   and starts the watchdog.

**Acceptance criteria:**

- A pod comes up, `bootstrap.sh` finishes without errors, and `jax.devices()` lists the GPU.
- A test run with `MAX_POD_HOURS=0.05` stops the pod within 5 minutes.
- `NOTES.md` records the CAX license check and the CAX version decision.

## M1: Emoji data

Do this milestone on the Mac. It needs no pod.

**Tasks:**

- Write `data/fetch.py`. It keeps the fully-qualified sequences from
  <https://unicode.org/Public/emoji/latest/emoji-test.txt> and downloads the matching PNGs
  from each licensed vendor's repository. Confirm the EmojiTwo license before you include it.
- Write `data/prep.py`. It resizes to 37×37 (record the filter), normalizes RGBA to [0, 1],
  applies your alpha decision, and pads by 15 px to 67×67. It computes x and y extents from
  alpha > 0.1 and the B.3 twin-score, and writes `emoji.npz`.
- Write `data/split.py` with the seeded scaffold split and the four forced targets, and
  commit the list with codepoints. Write `data/attribution.json` for every emoji.
- Plot a contact sheet of the 120 targets, a twin-score histogram, and counts per vendor.

**Acceptance criteria:**

- `NOTES.md` documents the shapes and dtypes of every array. Every image is 67×67×4 in
  [0, 1], with no NaNs, and every row has an attribution.
- The twin-score has unit tests: one blob gives 0, two identical blobs give 1, a mirrored pair
  gives 1 through the flip, and unequal areas scale by b.

## M2: One fish NCA

**Tasks:**

- Write `nca/model.py` (both alive masks, the 0.25 update mask, LoRA hooks), `nca/hashrng.py`,
  `nca/pool.py`, and `train/single.py` with the per-target recipe. Train the blue fish for
  10,000 updates, and log it to wandb.
- Measure the wall-clock time per 10,000 updates, compare it with the 3-7 minute estimate,
  and update the M3 and M5 budgets in `NOTES.md`.
- Write `nca/ref_numpy.py`, a NumPy step that matches `model.py`. M7 uses it during M3.

**Acceptance criteria:**

- The update network has 12,416 parameters, and the output kernel starts at zero.
- The fish grows recognizably: 16 stochastic rollouts at step 128 in `results/m2_fish.png`,
  with a final-step MSE below a threshold that you record. Record what happens at step 256,
  but don't require stability past step 128.
- The NumPy step matches JAX for one step with the same mask within 1e-5.

## M3: Shared scaffold

**Tasks:**

- Write `train/scaffold.py`: `W0` plus 120 rank-16 adapters, with A small and random and B zero
  at initialization. Follow A.2 with the single-GPU deviation: shuffle, process groups of 4
  with vmap, average the 4 gradients for the scaffold, and update each active adapter with its
  own gradient. Each target keeps its own reduced pool.
- Choose the number of outer iterations from the M2 timing to fit 2-4 GPU-hours, and record it.
- Save `W0` and all 120 adapters; M4 needs `ΔW_fish` and the transfer targets' adapters.
- Plot `results/m3_contact.png` (every target at step 128) and a per-target loss table.

**Acceptance criteria:**

- At least 90% of the 120 targets grow recognizably, judged from the contact sheet and losses.
- Both fish, the lizard, and the fire extinguisher are among the successes. If one isn't, stop
  and fix the scaffold before M4.
- `README.md` records the deviations from A.2.

## M4: Scale knobs and zero-shot transfer

**Tasks:**

- Build the 25 scale targets by resizing the fish to s_x, s_y ∈ {25, 31, 37, 43, 49} before
  padding. Write `train/scale.py`: freeze `W0` and `ΔW_fish`, and train only `Δx` and `Δy` with
  β(s) = β0·log(s/s0) and the A.4 pool and schedule.
- Write `train/lpips_eval.py`. On a grid over s' ∈ [15, 57] that includes the training sizes,
  run 16 rollouts per setting, score LPIPS (AlexNet) against the resized target, and plot an
  s_x × s_y heatmap.
- Apply the same `Δx` and `Δy` to `W0 + ΔW_k` for the second fish, the lizard, and the fire
  extinguisher, and plot the same heatmaps, laid out like the paper's Figure 4.
- Compute a noise floor: LPIPS between pairs of rollouts of each unmodified NCA.
- Control: add random rank-1 directions with the norms of `Δx` and `Δy`, and measure the
  bounding-box change.

**Acceptance criteria:**

- In distribution, LPIPS stays close to the noise floor (record the threshold), and
  `NOTES.md` has the out-of-distribution numbers over [15, 57].
- On all three transfer targets, the extents follow s' monotonically over the training range.
- The leak test passes, and the random-direction control doesn't stretch comparably.

## M5: Adapter sweep

**Tasks:**

- Write `train/sweep.py`. With `W0` frozen, it vmaps K adapters per step through a batched
  matmul, gives each a small pool, and saves each at its lowest loss. Tune K and record it.
- Choose targets that balance OpenMoji against the other vendors, oversample high
  twin-scores, and include the scaffold emoji in other vendors' styles.
- Stay within 10 GPU-hours: about 1,000 adapters, with a floor of 500. If 10,000 updates don't
  fit, cut to about 3,000 with a compressed schedule, and record it.
- Checkpoint in batches to `/workspace` and copy them to the Mac, so a stop loses one batch.

**Acceptance criteria:**

- At least 500 adapters exist, each with a recorded loss, plus a contact sheet of 64.
- `NOTES.md` has coverage by vendor and twin-score, the updates per adapter, and GPU-hours
  (10 or fewer).

## M6: Weight-space analysis

**Tasks:**

- Write `analysis/pca.py`. For every adapter, form `A B` for both kernels, flatten and
  concatenate them (96·96 + 96·32 = 12,288 dimensions), standardize, and run PCA. Briefly
  compare raw-factor and SVD-canonicalized PCA. Correlate PC0 and PC1 with the x and y
  extents, including the cross-correlations.
- Write `analysis/directions.py`: the style and fission directions (pick and record the
  twin-score split threshold), and the PC directions mapped back to raw weight coordinates.
  Compute cosines in full weight space: PC0 and PC1 against the outer products of `Δx` and
  `Δy`, and the style direction against PC0 and PC1.
- Write `analysis/causal.py`. On held-out adapters, add ±β times each direction. For PC0 and
  PC1, measure extents and brightness. For style, measure outline darkness (mean luminance in
  a 2 px ring inside the alpha boundary) and size. For fission, plot twin-score against β
  from 0 to 3, find the threshold, and test whether −β fuses split targets.
- Write `analysis/morphospace.py`: 2D PCA and UMAP coordinates for every shipped adapter.

**Acceptance criteria:**

- `NOTES.md` lists r(PC0, x), r(PC1, y), and the cross-correlations next to the paper's
  −0.71, −0.73, 0.07, and 0.08, and the cosines next to about 0.1, even if yours are weaker.
- The style test darkens outlines on non-OpenMoji adapters, or `NOTES.md` explains why not.
- The fission plot and threshold are in `results/`.

## M7: Export, goldens, and WGSL kernels

Start the step kernel on the Mac during M3, against `nca/ref_numpy.py`. Its acceptance
criteria need the goldens from this milestone.

**Export tasks:**

- Write `export/export_weights.py`: `weights.bin` in fp16 plus `manifest.json` with the names,
  shapes, and offsets of `W0` and its biases, `Δx` and `Δy`, A and B for every shipped adapter,
  and the PC0, PC1, style, and fission directions in raw weight coordinates. Add the 2D map
  coordinates, an emoji thumbnail sprite, and per-emoji attribution metadata.
- Write `export/golden.py`. For 5 targets × 3 knob settings (base, stretched, and fission), it
  dumps the full state after steps 1, 8, 64, and 128, with the hash PRNG mask.

**Kernels:**

- Write the following kernels. Use fp16 storage when the device supports `shader-f16`;
  otherwise use fp32. Accumulate in fp32 in both cases.
  - **Fused NCA step:** periodic Sobel perception, the pre-alive mask (3×3 max-pool of
    alpha > 0.1), 96→96 with ReLU, 96→32, the hash-based update mask (keep probability 0.75),
    the residual add, and the post-alive mask. Use ping-pong storage buffers, and batch up to
    64 creatures, each with its own effective weights.
  - **Weight composition:** `W_eff = W0 + A_k B_k + β_x Δx + β_y Δy + Σ γ_j D_j`, run only when
    a slider changes.
  - **Render pass:** RGBA over a checkerboard, with an optional D'Arcy Thompson grid overlay.

**Acceptance criteria:**

- Reloading `weights.bin` through the manifest reproduces the goldens within fp16 rounding.
- The Playwright tests pass in headless Chromium with WebGPU. At steps 1 and 8, the maximum
  absolute error is 1e-4 or less in fp32 and 1e-2 or less in fp16. At step 128, the final
  RGBA passes MSE and LPIPS thresholds that you record, because float differences compound in
  a dynamical system. The composition kernel matches NumPy within fp16 rounding.
- Performance estimates to verify on an Apple M-series Mac: 64 creatures at 60 steps per
  second or more, and one creature at 240 steps per second or more.
- **Benchmark:** Compare ms per step between these kernels and `onnxruntime-web` with the
  WebGPU execution provider, running the same model. Save the results to
  `results/bench.json`.

## M8: The demo page

**Tasks:**

- **Grow:** An emoji grows from one cell, with a step counter, speed control, and ms per step.
- **Stretch:** Width and height sliders drive `Δx` and `Δy` over s' ∈ [15, 57], with a
  "trained on 🐟 only" badge and the Thompson grid overlay.
- **Zoo:** 16-64 creatures grow side by side, and one knob moves them all.
- **Discovered knobs:** PC0 and PC1 size, OpenMoji style, and fission (β up to about 3), each
  labeled with the paper's number.
- **Morphospace map:** A 2D scatter of the shipped adapters. Click a dot to grow it; drag
  between two dots to blend adapters, labeled as an experiment beyond the paper.
- **Poke:** An eraser, with an honest note that these NCAs weren't trained to regenerate
  (paper, A.5). A regenerative variant is a stretch goal.
- **Inspector:** The cosine similarity between `Δx` and PC0, next to the paper's about 0.1.
- **Fallback and mobile:** Without `navigator.gpu`, show a recorded video and static images.
  On mobile, stack the panels and support touch.
- **Footer:** The paper credit, every emoji set's attribution (with OpenMoji's share-alike
  note), and "not affiliated with the authors."

**Acceptance criteria:**

- A Playwright smoke test loads the page, grows the fish for 128 steps, moves the width
  slider, and runs a zoo of 16.
- The page works in Chrome and Safari on macOS.

## M9: Deploy and write-up

**Tasks:**

- Write `deploy/Dockerfile`: `nginx:alpine` serving `web/dist`, buildable with the legacy
  builder. Serve precompressed `.bin` files with `gzip_static on`, and set
  `application/octet-stream` for `.bin`.
- Write `deploy/compose.yml`. It publishes no ports and joins the external network `edge` with
  the alias `sprout`.
- Ask the user before deploying. The user confirms a domain such as `sprout.ifkash.dev` and
  adds a non-proxied Cloudflare A record.
- After the user approves, copy the build and `deploy/` to `~/docs/sprout`, and run
  `docker compose up -d` there. Then compare the live and host Caddyfiles with
  `docker exec caddy cat /etc/caddy/Caddyfile | diff - ~/docs/caddy/Caddyfile`. If they
  differ, stop and ask the user. Otherwise, back up, append, validate, and reload:

  ```
  cd ~/docs/caddy
  cp Caddyfile "Caddyfile.bak-sprout-$(date +%Y%m%d)"
  printf '\nsprout.ifkash.dev {\n\treverse_proxy sprout:80\n}\n' >> Caddyfile
  docker exec -i caddy caddy validate --config - --adapter caddyfile < Caddyfile
  docker exec -i caddy caddy reload --config - --adapter caddyfile < Caddyfile
  ```

  Omit any `tls` block: for a non-proxied A record, the default ACME HTTP-01 challenge works.
- Write `README.md`. It explains what the project is, how to reproduce each milestone with one
  command per milestone, the results tables, every deviation from the paper, the emoji
  attributions, the paper credit, and the conflict-of-interest note.
- Write `post/draft.md`, a blog post of 1,200-1,800 words. Structure it as follows:
  - The hook: "I found the gene for 'wider' in a neural cellular automaton."
  - D'Arcy Thompson's grids, and what an NCA is, in plain words.
  - LoRA adapters as genes, and the fish knob that stretches a fire extinguisher.
  - The weight-space map, style, fission, and degeneracy: two directions, one job.
  - The WGSL step kernel and the benchmark.
  - Limitations: 2D emoji, not biology; no regeneration; about 1,000 adapters, not 25,000;
    a reimplementation without official code.
- Ask the user before you push anything. After the user approves, push the repo to
  `weights-and-wires/sprout` and upload the checkpoints to the Hugging Face Hub under
  `weights-and-wires/sprout-nca`.
- Terminate every pod, and write the final spend to `NOTES.md`. The owner lists the project on
  projects.dotslasha.me after the deploy; that isn't your job.

**Acceptance criteria:**

- The deployed URL loads over HTTPS, grows an emoji, and stretches it in Chrome.
- Every other site on the VM still responds as it did before the deploy.
- No pods are left running, and the total spend is less than $30.

## Final report to the user

When you finish, report the following:

- The scaffold success rate out of 120, from M3.
- The scale-knob LPIPS in distribution, out of distribution, and on the three transfer
  targets, with the noise floor, from M4.
- The adapter count and the updates per adapter, from M5.
- The PCA correlations next to the paper's, the cosine similarities, and the fission threshold
  you found, from M6.
- The kernel parity numbers and ms per step, with the benchmark, from M7.
- The URL, if the site is deployed.
- The total spend.
- The conflict-of-interest note, and anything that you skipped or that failed.
