sprout
A knob learned on one fish reshapes creatures it has never seen, live in your browser
Part 1: For you
The idea in one line
Tiny neural "cells" grow an emoji from a single seed cell, right in your browser tab. A "wider" knob trained only on a fish also stretches a lizard and a fire extinguisher that it never trained on. You turn these knobs on one creature, or on a zoo of 64 at once.
Why it's exciting
- The hook is strange and true. A September 2026 paper, On Growth and Form, and Function, trains two rank-1 knobs, width and height, on a single blue fish. The same knobs transfer zero-shot to a differently colored fish, a green lizard, and a red fire extinguisher, and keep their internal features intact.
- LoRA adapters act like genes. Every emoji is a small low-rank adapter on one shared update network, the way genes steer one shared developmental program. A rank-1 adapter has only 320 parameters, 2.6% of the base network.
- The knobs fall out of the data. From about 25,000 adapters, plain PCA finds size axes: the first two components correlate with width and height at r = −0.71 and r = −0.73. Simple mean differences give an OpenMoji "style" knob that adds a black outline, and a "fission" knob that splits a creature into symmetric twins.
- Two different directions do the same job. The PCA size axes and the trained fish knobs have a cosine similarity of only about 0.1, yet both stretch creatures ("degeneracy").
- It's a century-old idea, computed. D'Arcy Thompson drew fish on warped grids in 1917. The demo draws his grid over a live, growing creature.
- Nobody has built this demo. The paper has no code link, and GitHub and web searches found only the arXiv pages. It appeared on September 24, 2026, so ship soon.
- It's low-level. You write the growth step in WGSL by hand: one fused kernel runs 64 creatures, each with its own weights, in one dispatch. There's no server and no ML runtime.
What the demo looks like
- A grow panel. Pick an emoji and watch it grow from one cell, with a step counter, a speed control, and milliseconds per step.
- Width and height sliders that drive the fish-trained knobs, with a "trained on 🐟 only" badge and an optional Thompson grid overlay.
- A zoo of 16 to 64 creatures growing side by side. One knob moves them all.
- Discovered knobs: size, OpenMoji style, and fission, each with the paper's number.
- A morphospace map. Click a dot to grow that adapter, or drag between two dots to blend them, which is an experiment beyond the paper.
- A poke tool that cuts the creature, with an honest note that it wasn't trained to heal.
- An inspector that shows the cosine similarity between the fish knob and the PCA size axis, next to the paper's value of about 0.1.
How it works
- Prepare the emoji. On your Mac, shrink openly licensed emoji to 37×37 pixels, pad them to 67×67, and measure each one's size and "twin-score" (how much it looks like two halves).
- Grow one fish. On a GPU pod, reproduce one neural cellular automaton (NCA) that grows the fish, and time it so that the rest of the plan uses real numbers.
- Train the shared scaffold. Train one shared network jointly with a small adapter for each of 120 Noto emoji. The shared network becomes the common "body plan."
- Train the fish knobs. Train rank-1 width and height adapters on stretched copies of the fish only. Then apply them to the lizard, the fire extinguisher, and a second fish.
- Sweep adapters overnight. Train about 1,000 adapters on the frozen scaffold at once.
- Find knobs in weight space. Run PCA, compute the style and fission directions, and push each one to measure what changes.
- Write the kernels. Rewrite the growth step in WGSL and check it step by step against JAX, with the same random numbers on both sides.
- Ship it as a static web page on
vm.ifkash.dev.
Weekend plan
| When | What | Done when |
|---|---|---|
| Saturday morning | Prepare the emoji data on the Mac; start the pod; reproduce one fish NCA and time it | The fish grows; the budget uses measured speed |
| Saturday afternoon | Train the shared scaffold on 120 Noto emoji (runs by itself); start the WGSL step kernel on the Mac against a CPU reference | At least 90% of the 120 emoji grow |
| Saturday evening | Train the fish width and height knobs; test zero-shot transfer; start the overnight adapter sweep | The fish knob stretches the lizard |
| Sunday morning | Weight-space analysis (PCA, style, fission); export weights and goldens; kernel parity | Browser rollouts match JAX |
| Sunday afternoon | Build the demo page, deploy it, record the video, write the post | The link works |
Cost
- GPU: 1× RTX 5090 on RunPod, about $0.69 per hour. The fallback is an RTX PRO 6000, about $1.69 per hour.
- Time on the pod: about 15-20 hours, by estimate. The single-fish run measures the real speed on Saturday morning.
- Total: about $10-15. The plan sets a hard stop at $30.
- Hosting: free. It's a static page on your VM, and it keeps working after the pod is gone.
What you have at the end
- A live link where anyone can grow emoji and stretch a fire extinguisher with a fish knob.
- Zero-shot transfer heatmaps for the lizard, the fire extinguisher, and a second fish, laid out like the paper's Figure 4.
- A weight-space map of about 1,000 adapters, with your PCA correlations next to the paper's.
- A video of the fish knob stretching a whole zoo at once.
- A blog post: "I found the gene for 'wider' in a neural cellular automaton."
What might go wrong
| Problem | What to do |
|---|---|
| Some of the 120 scaffold emoji don't grow | Accept 90% or more. The two fish, the lizard, and the fire extinguisher must be among the ones that grow. |
| The fish knob doesn't transfer | Report it honestly. Check the scaffold quality and the alpha handling first; the in-distribution result is still a demo. |
| About 1,000 adapters give weaker PCA correlations than the paper's | Expected: the paper used about 25,000. Report both, and weight the sweep toward OpenMoji and twin-shaped emoji so that style and fission have enough samples. |
| The sweep runs over its time budget | Cut updates per adapter to about 3,000, and keep at least 500 adapters. |
| Browser and JAX rollouts drift apart | Use the same hash-based random numbers in both. Compare early steps tightly and the final image by MSE and LPIPS, because small float errors grow over 128 steps. |
| Visitors expect a poked creature to heal | It won't: the paper's NCAs weren't trained to regenerate. Say so on the page; a healing variant is a stretch goal. |
| WebGPU isn't available in a visitor's browser | Show a recorded video and static images instead. |
| The paper's conflict-of-interest statement | Astonishing Labs partly funded the research and has certain rights to inventions from it. Keep sprout a personal, non-commercial reimplementation that credits the paper. |
| Emoji licenses | Use only the openly licensed sets, credit each one in the footer, and note OpenMoji's share-alike terms. |
Other ideas the research turned up
- Bolt sprinter. A 136-muscle, 31-DOF skeleton learns to run 100 m in about 10 s with no motion capture (top speed 12.5 ± 0.5 m/s), from a 300K-parameter MLP policy trained with FastTD3. It's the best hook, but porting the Warp simulator to the browser is hard, and the repositories have no license file.
- A GRU that builds a dodecahedron. A recurrent net that tracks a spinning icosahedron puts its 20 coset means almost exactly at the vertices of a regular dodecahedron in its hidden state.
- Groundhog bit-flip. Flipping 3 router bits per expert makes MoE language models never stop; deactivating fewer than 4 experts drives average output inflation to 5,912%. It's unverified at browser scale.
- A 4,000-parameter NCA that grows the cosmic web. Its power-spectrum residual stays under 1% for k ≲ 0.5 h/Mpc, but browser cosmic-web toys already exist.
- Motion-only CAPTCHA. Humans pass 99.6% of the time, and the best GUI agent passes 16.8% (ACM MM 2026).
Reading, if you want it
- On Growth and Form, and Function: the paper this project builds on.
- Growing Neural Cellular Automata (Mordvintsev et al., Distill 2020): the NCA architecture the paper uses.
- CAX and its paper: the JAX cellular automata library the authors used, with a growing-NCA example notebook.
- LoRA: the low-rank adapters that act as genes here.
- On Growth and Form by D'Arcy Thompson, on Project Gutenberg: the grid transformations this demo animates.
- LPIPS: the perceptual metric the paper uses to score stretched creatures.
Part 2: For the coding agent
Mission
Build sprout, a browser demo of neural cellular automata (NCAs) that grow emoji from a
single cell, with low-rank adapters as reusable control knobs, entirely on the client:
- Models: A shared scaffold NCA,
W0, trained jointly with rank-16 adapters on 120 Noto Emoji targets; two rank-1 scale knobs,ΔxandΔy, trained on one fish only; and a sweep of about 1,000 rank-16 adapters on the frozen scaffold. - Analysis: PCA on full weight-space adapters, plus style and fission directions, each with a causal test.
- Kernels: A fused NCA step, a weight composition kernel, and a render pass, in WGSL.
- Demo page: Grow, stretch, zoo, discovered knobs, a morphospace map, poke, and inspector.
The final artifact is a static website with no inference server. The paper released no code,
so you reimplement it from the text. Record every choice that the paper leaves open in
NOTES.md.
Work through milestones M0-M9 in order. Each milestone has acceptance criteria. Don't start a
milestone until the previous one passes, except where a milestone says it runs in parallel.
After each milestone, commit your work and write a short entry in NOTES.md with the results
and numbers.
Hard constraints
- Secrets: Read
RUNPOD_API_KEY,HF_TOKEN, andWANDB_API_KEYfrom environment variables only. Never write a secret into any file, log, commit, or echoed command. Commit a.env.examplethat has placeholder values only, and add.envto.gitignore. - Budget: The hard cap is $30 of RunPod spend. Every pod runs a watchdog that stops the pod
after
MAX_POD_HOURShours. The default is6. - Compute: Run training, the adapter sweep, and LPIPS evaluation only on RunPod. The local Mac runs the emoji data preparation, the NumPy reference step, the web build, and the browser tests.
- Transfer targets:
ΔxandΔytrain on the fish only. No code path in the scale-knob training may read the second fish, the lizard, or the fire extinguisher. Enforce this with an assertion and a test. - Outward actions: Ask the user before you do any of the following: create the GitHub
repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything
publicly. Use the
kashifulhaqueGitHub account (gh auth switch -u kashifulhaque). - Shared VM:
vm.ifkash.devruns other production apps behind one shared Caddy, which owns ports 80 and 443. Its config lives at~/docs/caddy. - The sprout container lives in
~/docs/sproutand must not publish any ports. It joins the external Docker networkedgewith a stable alias, such assprout. - Add the vhost only by appending to
~/docs/caddy/Caddyfilewith>>. Never rewrite, rename, or replace that file. It's a single-file bind mount, and a rewrite orphans the inode, socaddy reloadthen reports "config is unchanged" while serving the old config. - Docker on the VM has no BuildKit for plain
docker build. Don't useCOPY --chmod, Dockerfile heredocs, orRUN --mount. Multi-stage builds andCOPY --fromwork. - Don't stop, restart, reconfigure, or remove any other container, network, or vhost.
- Pod cleanup: When a pod isn't running a job, stop it. At the end of the project, terminate every pod that you created and report the total spend.
- Licenses: Use only the emoji sets in the design decisions, and attribute every set you
use in
README.mdand the page footer. Confirm the CAX license (MIT) during M0, and keep its notice in any file that you copy or adapt. Credit the paper (CC BY 4.0), and state that sprout isn't affiliated with the authors. - Conflict of interest: Per the paper, Astonishing Labs partly funded the research, M.L. is a co-founder, and Astonishing Labs has certain rights to associated inventions. Keep sprout a personal, non-commercial reimplementation, and flag this in the final report.
Science background
These facts come from "On Growth and Form, and Function: Reusable Regulatory Handles Control Phenotypic Variation" by Hartl, Montero, Barylli, Risi, and Levin, on arXiv on September 24, 2026 (cs.NE, CC BY 4.0): https://arxiv.org/abs/2609.29755. Implement the experiments to match them:
- Architecture: The Growing NCA of Mordvintsev et al. (2020): a 67×67 grid with periodic boundaries and C = 32 channels per cell, where the last four are RGBA and alpha also sets viability. Fixed perception is identity plus horizontal and vertical Sobel on the 3×3 Moore neighborhood, which gives 96 features. Two pointwise layers, 96→96 with ReLU, then 96→32, with a zero-initialized output kernel: about 12,416 parameters. Cell-wise dropout of 0.25 runs during training and inference, which acts as a stochastic update mask. Alpha is thresholded at 0.1 over a 3×3 neighborhood before and after each update. The authors used JAX and Flax NNX with CAX v0.2.1.
- Targets: RGBA normalized to [0, 1], resized to 37×37, and padded by 15 px on each side to 67×67. Growth starts from one living central seed cell.
- Per-target recipe: A pool of 1,024 states. Each update samples 8, replaces the highest-loss one with a fresh seed, evolves 128 steps, and takes an RGBA MSE loss at a random step 64 ≤ t < 128. Gradients are clipped to unit norm. AdamW with weight decay 1e-4 and an initial learning rate of 1.1618e-3, decayed linearly to 0.1× over 6,000 updates, then constant for the remaining 4,000 (10,000 in total).
- LoRA: Adapters modulate both pointwise kernels, not the biases or the perception. At rank r an adapter has 320·r parameters: 320 at r = 1 (2.6% of the base net), 5,120 at r = 16.
- Shared scaffold (A.2): The scaffold trains jointly with rank-16 adapters on 120 Noto
Emoji targets: 10 workers × 12 targets. Each outer iteration shuffles and processes 3 groups
of 4; within a group each target's adapter is activated in turn, and the 4 gradients are
averaged before the scaffold and active adapters update. Each target keeps its own pool.
10,000 outer iterations per worker (30,000 four-target updates). Every 100 iterations,
θ_shared ← 0.9 θ_shared + 0.1 θ_worker, the optimizer is reinitialized, and pools are kept.
Afterward the adapters are removed, and
W0stays as a target-independent reference. - Adapter sweep (A.3): With
W0frozen, one rank-16 adapter per emoji-vendor pair trains for 10,000 updates and is saved at its lowest training loss, about 25,000 adapters in total. Emoji identities are the fully-qualified sequences in Unicodeemoji-test.txt. - Scale knobs (3.2-3.3, A.4): The baseline is a blue fish, one of the 120 scaffold targets,
with weights
W0 + ΔW_fishat s0 = 37 px. Two rank-1 adapters,ΔxandΔy, train jointly for 10,000 iterations on 5 × 5 = 25 targets with s_x, s_y ∈ {25, 31, 37, 43, 49} px; nothing else trains. Pool of 8 per target. Learning rate 5e-4, decayed linearly to 10% over 3,000 steps, then constant for 7,000. The knobs generalize to s' ∈ [15, 57] (LPIPS with AlexNet, 16 stochastic rollouts per setting) and transfer zero-shot to a differently colored fish, a green lizard, and a red fire extinguisher. Internal features survive, and anisotropic scaling shears oriented shapes. - Weight-space analysis (3.4, B.2): PCA on the full weight-space
ΔW_φ = A_φ B_φ(layer matrices flattened, concatenated, and standardized per dimension). PCA on raw LoRA factors is uninformative because of the gauge symmetry (A R, R⁻¹ B); SVD canonicalization helps. r(PC0, x) = −0.71 and r(PC1, y) = −0.73; r(PC0, y) = 0.07 and r(PC1, x) = 0.08. Adding ±β ΔW_PC stretches or compresses and also changes brightness. PC directions have a cosine of about 0.1 withΔxandΔy. The style direction, mean(OpenMoji ΔW) − mean(others), adds the black outline and shrinks toward the OpenMoji mean size, with a cosine of about 0.1 with PC0 and PC1. The fission directionΔW_F= mean(split) − mean(non-split), with split set by the twin-score: above a target-specific β ≳ 2 the NCA grows split twins, and the negative direction fuses split targets but loses some detail. High-fidelity reconstruction from PCs needs thousands of PCs. - Twin-score (B.3): Take 8-connected foreground components, and drop fragments under 5% of the foreground area. If there aren't exactly two, T = 0. Otherwise b = min/max area; each component's bounding box is resized to a 32×32 binary mask (nearest); s = max(IoU(C1, C2), IoU(C1, flip_x(C2))); and T = b·s ∈ [0, 1].
- Regeneration (A.5): These NCAs weren't trained to regenerate. A regenerative LoRA loss didn't make the scaffold significantly regenerative; a regeneration-trained adapter had a cosine of about 0 with the normal one, and its difference vector didn't transfer.
The core equations are as follows:
x_{t+1} = alive(x_t + m_t ⊙ f(perceive(x_t))) m_t: cell-wise keep mask, p = 0.75
ΔW = Σ_k β_k A_k B_k on both pointwise kernels
β(s) = β0 · log(s / s0) s0 = 37 px; β0 unstated: pick, record
W = W0 + ΔW_k + Σ_q β(s_q) Δ_q ΔW_k: rank-16 target adapter
Design decisions
- Emoji sets: Use Noto Emoji (images Apache-2.0, fonts OFL-1.1), Twemoji (graphics CC-BY 4.0, code MIT), OpenMoji (CC-BY-SA-4.0; note share-alike for derived images), Fluent UI Emoji (MIT), and Blobmoji (Apache-2.0). GitHub asserts no license for EmojiTwo; include it only if you confirm CC-BY 4.0. Skip JoyPixels, EmojiOne 3, and TossFace, whose licenses are restrictive or unclear.
- Split: Choose the 120 Noto scaffold targets with a fixed seed. The set must include a blue fish as k = 1, a differently colored fish, a lizard, and a fire extinguisher. Pick the exact codepoints and record them. Everything else is sweep-only.
- Alpha: Distill premultiplies RGB by alpha; the paper doesn't say. Check the CAX example, pick one, and record it.
- CAX version: Decide between pinning 0.2.1 (the paper's version) and release 0.4.4 (Python ≥ 3.12), and record the choice and the reason.
- Scaffold on one GPU (estimate): One worker over all 120 targets, with vmapped groups of 4, instead of 10 workers and merging. A 1,024-state fp32 pool is about 590 MB per target, about 70 GB for 120, so shrink the pools (for example, to 128 states) or store them in fp16. Record both deviations. The estimate is 2-4 GPU-hours.
- Sweep scale-down (estimate): vmap many adapters per step, with a shared frozen
W0and per-adapter A and B through a batched matmul. Cap the sweep at 10 GPU-hours. Target about 1,000 adapters, with a floor of 500. If needed, cut updates per adapter, for example to 3,000 with the schedule compressed in proportion. Prioritize OpenMoji against the other vendors, and high twin-score targets, so that style and fission get enough samples. - Speed and cost (estimates): One update is about 8 states × 128 steps × 4,489 cells × about 26k FLOPs forward, or about 1.2e11 FLOPs, and about 3× that with the backward pass: about 3-7 minutes per 10,000-update adapter alone on an RTX 5090. Measure it in M2. The whole project is about 15-20 pod-hours, about $10-15.
- Determinism for parity: Implement the same counter-based hash PRNG, such as PCG, keyed
on (seed, step, cell), for the update mask in JAX and WGSL. Use it for every golden rollout.
Training can use
jax.random.
Tech stack
Pin every version. The stack is as follows:
- Python: Python 3.12, JAX, Flax NNX,
optax, CAX (pinned),numpy,pillow,scikit-learnfor PCA,umap-learn,matplotlib, andwandb, managed withuv. Uselpips(PyTorch) for evaluation only. Start from the CAX examples40_growing_nca.ipynband41_growing_conditional_nca.ipynb. - Browser: TypeScript, bundled with
vite, and raw WebGPU with WGSL. Don't use an ML runtime in the final demo.onnxruntime-webis allowed only as the benchmark baseline. - Browser tests: Playwright with Chromium, with WebGPU enabled.
- Pods: the
runpodPython SDK, withrunpodctlinside pods.
Repository layout
Create the following layout:
sprout/
pyproject.toml .env.example .gitignore README.md NOTES.md
infra/pod.py infra/watchdog.sh infra/bootstrap.sh
data/fetch.py # emoji-test.txt + vendor PNGs -> data/raw/ (Mac)
data/prep.py # resize, pad, bbox extents, twin-scores -> emoji.npz
data/split.py # seeded 120-target scaffold split, forced transfer targets
data/attribution.json # per-emoji vendor, license, source URL
nca/
model.py # perception, update net, alive masks, LoRA hooks (Flax NNX)
hashrng.py # counter-based PCG hash for update masks (mirrors hash.wgsl)
pool.py twin.py # sample pool; twin-score (B.3)
ref_numpy.py # NumPy reference step for the WGSL work
train/ # configs/{fish,scaffold,scale,sweep}.yaml
single.py # M2: one NCA for one target
scaffold.py # M3: shared W0 + rank-16 adapters
scale.py # M4: rank-1 Δx, Δy on the fish only
sweep.py # M5: vmapped rank-16 adapters on frozen W0
lpips_eval.py # LPIPS heatmaps, 16 rollouts per setting
analysis/pca.py analysis/directions.py analysis/causal.py analysis/morphospace.py
export/export_weights.py # -> weights.bin (fp16) + manifest.json + sprite.png
export/golden.py # steps 1, 8, 64, 128 for 5 targets x 3 knob settings
web/
src/gpu/ # WGSL kernels + TS dispatch code
step.wgsl compose.wgsl render.wgsl hash.wgsl step.ts compose.ts render.ts
src/panels/ # grow, stretch, zoo, knobs, map, poke, inspector
src/main.ts index.html
public/ # weights.bin, manifest.json, sprite.png, fallback video
tests/ # Playwright: kernel golden tests, smoke test
bench/ # WGSL vs onnxruntime-web benchmark page
deploy/Dockerfile deploy/compose.yml deploy/nginx.conf
results/ post/draft.md
M0: Infrastructure
Tasks:
- Write
infra/pod.py. It uses therunpodSDK and readsRUNPOD_API_KEYfrom the environment. It supports the following subcommands: -create: Creates a pod. The GPU preference order is RTX 5090, then RTX PRO 6000, then RTX 4090. Resolve the GPU type IDs at run time by querying the available GPU types. Use an official RunPod image with CUDA 12.8 or later. Set a 100 GB volume at/workspaceand expose SSH. -status,stop, andterminate. -ssh-info: Prints the SSH command. -cost: Prints the uptime and spend for every pod whose name has the prefixsprout-. - Write
infra/watchdog.sh. It sleeps forMAX_POD_HOURS, then runsrunpodctl stop pod $RUNPOD_POD_ID. - Write
infra/bootstrap.sh. It clones the repo, runsuv syncwith the CUDA build of JAX, and starts the watchdog.
Acceptance criteria:
- A pod comes up,
bootstrap.shfinishes without errors, andjax.devices()lists the GPU. - A test run with
MAX_POD_HOURS=0.05stops the pod within 5 minutes. NOTES.mdrecords the CAX license check and the CAX version decision.
M1: Emoji data
Do this milestone on the Mac. It needs no pod.
Tasks:
- Write
data/fetch.py. It keeps the fully-qualified sequences from https://unicode.org/Public/emoji/latest/emoji-test.txt and downloads the matching PNGs from each licensed vendor's repository. Confirm the EmojiTwo license before you include it. - Write
data/prep.py. It resizes to 37×37 (record the filter), normalizes RGBA to [0, 1], applies your alpha decision, and pads by 15 px to 67×67. It computes x and y extents from alpha > 0.1 and the B.3 twin-score, and writesemoji.npz. - Write
data/split.pywith the seeded scaffold split and the four forced targets, and commit the list with codepoints. Writedata/attribution.jsonfor every emoji. - Plot a contact sheet of the 120 targets, a twin-score histogram, and counts per vendor.
Acceptance criteria:
NOTES.mddocuments the shapes and dtypes of every array. Every image is 67×67×4 in [0, 1], with no NaNs, and every row has an attribution.- The twin-score has unit tests: one blob gives 0, two identical blobs give 1, a mirrored pair gives 1 through the flip, and unequal areas scale by b.
M2: One fish NCA
Tasks:
- Write
nca/model.py(both alive masks, the 0.25 update mask, LoRA hooks),nca/hashrng.py,nca/pool.py, andtrain/single.pywith the per-target recipe. Train the blue fish for 10,000 updates, and log it to wandb. - Measure the wall-clock time per 10,000 updates, compare it with the 3-7 minute estimate,
and update the M3 and M5 budgets in
NOTES.md. - Write
nca/ref_numpy.py, a NumPy step that matchesmodel.py. M7 uses it during M3.
Acceptance criteria:
- The update network has 12,416 parameters, and the output kernel starts at zero.
- The fish grows recognizably: 16 stochastic rollouts at step 128 in
results/m2_fish.png, with a final-step MSE below a threshold that you record. Record what happens at step 256, but don't require stability past step 128. - The NumPy step matches JAX for one step with the same mask within 1e-5.
M3: Shared scaffold
Tasks:
- Write
train/scaffold.py:W0plus 120 rank-16 adapters, with A small and random and B zero at initialization. Follow A.2 with the single-GPU deviation: shuffle, process groups of 4 with vmap, average the 4 gradients for the scaffold, and update each active adapter with its own gradient. Each target keeps its own reduced pool. - Choose the number of outer iterations from the M2 timing to fit 2-4 GPU-hours, and record it.
- Save
W0and all 120 adapters; M4 needsΔW_fishand the transfer targets' adapters. - Plot
results/m3_contact.png(every target at step 128) and a per-target loss table.
Acceptance criteria:
- At least 90% of the 120 targets grow recognizably, judged from the contact sheet and losses.
- Both fish, the lizard, and the fire extinguisher are among the successes. If one isn't, stop and fix the scaffold before M4.
README.mdrecords the deviations from A.2.
M4: Scale knobs and zero-shot transfer
Tasks:
- Build the 25 scale targets by resizing the fish to s_x, s_y ∈ {25, 31, 37, 43, 49} before
padding. Write
train/scale.py: freezeW0andΔW_fish, and train onlyΔxandΔywith β(s) = β0·log(s/s0) and the A.4 pool and schedule. - Write
train/lpips_eval.py. On a grid over s' ∈ [15, 57] that includes the training sizes, run 16 rollouts per setting, score LPIPS (AlexNet) against the resized target, and plot an s_x × s_y heatmap. - Apply the same
ΔxandΔytoW0 + ΔW_kfor the second fish, the lizard, and the fire extinguisher, and plot the same heatmaps, laid out like the paper's Figure 4. - Compute a noise floor: LPIPS between pairs of rollouts of each unmodified NCA.
- Control: add random rank-1 directions with the norms of
ΔxandΔy, and measure the bounding-box change.
Acceptance criteria:
- In distribution, LPIPS stays close to the noise floor (record the threshold), and
NOTES.mdhas the out-of-distribution numbers over [15, 57]. - On all three transfer targets, the extents follow s' monotonically over the training range.
- The leak test passes, and the random-direction control doesn't stretch comparably.
M5: Adapter sweep
Tasks:
- Write
train/sweep.py. WithW0frozen, it vmaps K adapters per step through a batched matmul, gives each a small pool, and saves each at its lowest loss. Tune K and record it. - Choose targets that balance OpenMoji against the other vendors, oversample high twin-scores, and include the scaffold emoji in other vendors' styles.
- Stay within 10 GPU-hours: about 1,000 adapters, with a floor of 500. If 10,000 updates don't fit, cut to about 3,000 with a compressed schedule, and record it.
- Checkpoint in batches to
/workspaceand copy them to the Mac, so a stop loses one batch.
Acceptance criteria:
- At least 500 adapters exist, each with a recorded loss, plus a contact sheet of 64.
NOTES.mdhas coverage by vendor and twin-score, the updates per adapter, and GPU-hours (10 or fewer).
M6: Weight-space analysis
Tasks:
- Write
analysis/pca.py. For every adapter, formA Bfor both kernels, flatten and concatenate them (96·96 + 96·32 = 12,288 dimensions), standardize, and run PCA. Briefly compare raw-factor and SVD-canonicalized PCA. Correlate PC0 and PC1 with the x and y extents, including the cross-correlations. - Write
analysis/directions.py: the style and fission directions (pick and record the twin-score split threshold), and the PC directions mapped back to raw weight coordinates. Compute cosines in full weight space: PC0 and PC1 against the outer products ofΔxandΔy, and the style direction against PC0 and PC1. - Write
analysis/causal.py. On held-out adapters, add ±β times each direction. For PC0 and PC1, measure extents and brightness. For style, measure outline darkness (mean luminance in a 2 px ring inside the alpha boundary) and size. For fission, plot twin-score against β from 0 to 3, find the threshold, and test whether −β fuses split targets. - Write
analysis/morphospace.py: 2D PCA and UMAP coordinates for every shipped adapter.
Acceptance criteria:
NOTES.mdlists r(PC0, x), r(PC1, y), and the cross-correlations next to the paper's −0.71, −0.73, 0.07, and 0.08, and the cosines next to about 0.1, even if yours are weaker.- The style test darkens outlines on non-OpenMoji adapters, or
NOTES.mdexplains why not. - The fission plot and threshold are in
results/.
M7: Export, goldens, and WGSL kernels
Start the step kernel on the Mac during M3, against nca/ref_numpy.py. Its acceptance
criteria need the goldens from this milestone.
Export tasks:
- Write
export/export_weights.py:weights.binin fp16 plusmanifest.jsonwith the names, shapes, and offsets ofW0and its biases,ΔxandΔy, A and B for every shipped adapter, and the PC0, PC1, style, and fission directions in raw weight coordinates. Add the 2D map coordinates, an emoji thumbnail sprite, and per-emoji attribution metadata. - Write
export/golden.py. For 5 targets × 3 knob settings (base, stretched, and fission), it dumps the full state after steps 1, 8, 64, and 128, with the hash PRNG mask.
Kernels:
- Write the following kernels. Use fp16 storage when the device supports
shader-f16; otherwise use fp32. Accumulate in fp32 in both cases. - Fused NCA step: periodic Sobel perception, the pre-alive mask (3×3 max-pool of alpha > 0.1), 96→96 with ReLU, 96→32, the hash-based update mask (keep probability 0.75), the residual add, and the post-alive mask. Use ping-pong storage buffers, and batch up to 64 creatures, each with its own effective weights.
- Weight composition:
W_eff = W0 + A_k B_k + β_x Δx + β_y Δy + Σ γ_j D_j, run only when a slider changes. - Render pass: RGBA over a checkerboard, with an optional D'Arcy Thompson grid overlay.
Acceptance criteria:
- Reloading
weights.binthrough the manifest reproduces the goldens within fp16 rounding. - The Playwright tests pass in headless Chromium with WebGPU. At steps 1 and 8, the maximum absolute error is 1e-4 or less in fp32 and 1e-2 or less in fp16. At step 128, the final RGBA passes MSE and LPIPS thresholds that you record, because float differences compound in a dynamical system. The composition kernel matches NumPy within fp16 rounding.
- Performance estimates to verify on an Apple M-series Mac: 64 creatures at 60 steps per second or more, and one creature at 240 steps per second or more.
- Benchmark: Compare ms per step between these kernels and
onnxruntime-webwith the WebGPU execution provider, running the same model. Save the results toresults/bench.json.
M8: The demo page
Tasks:
- Grow: An emoji grows from one cell, with a step counter, speed control, and ms per step.
- Stretch: Width and height sliders drive
ΔxandΔyover s' ∈ [15, 57], with a "trained on 🐟 only" badge and the Thompson grid overlay. - Zoo: 16-64 creatures grow side by side, and one knob moves them all.
- Discovered knobs: PC0 and PC1 size, OpenMoji style, and fission (β up to about 3), each labeled with the paper's number.
- Morphospace map: A 2D scatter of the shipped adapters. Click a dot to grow it; drag between two dots to blend adapters, labeled as an experiment beyond the paper.
- Poke: An eraser, with an honest note that these NCAs weren't trained to regenerate (paper, A.5). A regenerative variant is a stretch goal.
- Inspector: The cosine similarity between
Δxand PC0, next to the paper's about 0.1. - Fallback and mobile: Without
navigator.gpu, show a recorded video and static images. On mobile, stack the panels and support touch. - Footer: The paper credit, every emoji set's attribution (with OpenMoji's share-alike note), and "not affiliated with the authors."
Acceptance criteria:
- A Playwright smoke test loads the page, grows the fish for 128 steps, moves the width slider, and runs a zoo of 16.
- The page works in Chrome and Safari on macOS.
M9: Deploy and write-up
Tasks:
- Write
deploy/Dockerfile:nginx:alpineservingweb/dist, buildable with the legacy builder. Serve precompressed.binfiles withgzip_static on, and setapplication/octet-streamfor.bin. - Write
deploy/compose.yml. It publishes no ports and joins the external networkedgewith the aliassprout. - Ask the user before deploying. The user confirms a domain such as
sprout.ifkash.devand adds a non-proxied Cloudflare A record. - After the user approves, copy the build and
deploy/to~/docs/sprout, and rundocker compose up -dthere. Then compare the live and host Caddyfiles withdocker exec caddy cat /etc/caddy/Caddyfile | diff - ~/docs/caddy/Caddyfile. If they differ, stop and ask the user. Otherwise, back up, append, validate, and reload:
cd ~/docs/caddy
cp Caddyfile "Caddyfile.bak-sprout-$(date +%Y%m%d)"
printf '\nsprout.ifkash.dev {\n\treverse_proxy sprout:80\n}\n' >> Caddyfile
docker exec -i caddy caddy validate --config - --adapter caddyfile < Caddyfile
docker exec -i caddy caddy reload --config - --adapter caddyfile < Caddyfile
Omit any tls block: for a non-proxied A record, the default ACME HTTP-01 challenge works.
- Write README.md. It explains what the project is, how to reproduce each milestone with one
command per milestone, the results tables, every deviation from the paper, the emoji
attributions, the paper credit, and the conflict-of-interest note.
- Write post/draft.md, a blog post of 1,200-1,800 words. Structure it as follows:
- The hook: "I found the gene for 'wider' in a neural cellular automaton."
- D'Arcy Thompson's grids, and what an NCA is, in plain words.
- LoRA adapters as genes, and the fish knob that stretches a fire extinguisher.
- The weight-space map, style, fission, and degeneracy: two directions, one job.
- The WGSL step kernel and the benchmark.
- Limitations: 2D emoji, not biology; no regeneration; about 1,000 adapters, not 25,000;
a reimplementation without official code.
- Ask the user before you push anything. After the user approves, push the repo to
weights-and-wires/sprout and upload the checkpoints to the Hugging Face Hub under
weights-and-wires/sprout-nca.
- Terminate every pod, and write the final spend to NOTES.md. The owner lists the project on
projects.dotslasha.me after the deploy; that isn't your job.
Acceptance criteria:
- The deployed URL loads over HTTPS, grows an emoji, and stretches it in Chrome.
- Every other site on the VM still responds as it did before the deploy.
- No pods are left running, and the total spend is less than $30.
Final report to the user
When you finish, report the following:
- The scaffold success rate out of 120, from M3.
- The scale-knob LPIPS in distribution, out of distribution, and on the three transfer targets, with the noise floor, from M4.
- The adapter count and the updates per adapter, from M5.
- The PCA correlations next to the paper's, the cosine similarities, and the fission threshold you found, from M6.
- The kernel parity numbers and ms per step, with the benchmark, from M7.
- The URL, if the site is deployed.
- The total spend.
- The conflict-of-interest note, and anything that you skipped or that failed.