# dreamcart: a video game with no game code, dreamed by a $4 chip

---

# Part 1: For you

## The idea in one line

You train a tiny neural network to *be* a video game, then run it on your RP2040 board. The chip
has no game code. Every frame you see, the network imagines.

## Why it's exciting

- **World models are the hottest topic in AI right now.** Genie 3, Matrix-Game, and SANA-WM are
  all neural networks that imagine an interactive world frame by frame. They all need big GPUs.
- **You go the other way.** You find the smallest world model that can still run a game, and you
  run it on a chip with 264 KB of RAM. I couldn't find anyone who has done this.
- **Only you can do it quickly.** You already got an LLM running on this exact board. This is
  the same skill, pointed at the hottest topic of the year.
- **The demo is easy to share.** A video of a game running on a tiny board, with the caption
  "there's no game on this chip," is very shareable.
- **It also has a real research result.** You get a chart that answers the question: *how many
  parameters does it take to be Breakout?*

## How it works

1. **Write the real game.** It's a tiny Breakout: 32×32 pixels, a paddle, a ball, and bricks.
   You write it in PyTorch so the GPU can run 4,096 games at once. That gives you endless
   training data and nothing to store on disk.
2. **Train the dreamer.** A small network sees the last 2 frames and your button press, and
   draws the next frame.
3. **Teach it not to drift.** Small mistakes pile up over time, and the game "melts." You fix
   this by training the model on its own outputs, a trick from the Self Forcing paper.
4. **Shrink it.** You train 7 sizes, from 8,000 to 1 million parameters, and find the smallest
   one that still plays by the rules.
5. **Put it on the chip.** Convert the model to 8-bit integers, write the math in plain C, and
   flash it to the RP2040. Your terminal becomes the screen, and your keyboard becomes the
   controller.
6. **Put it on the web.** Compile the *same C code* for the browser, so anyone can play it.
   Host it on `vm.ifkash.dev`.

## Weekend plan

| When | What | Done when |
|---|---|---|
| Saturday morning | Write the real game and check it on your Mac | The game runs and looks right |
| Saturday afternoon | Train the dreamer on a GPU | The dreamed game looks like the real one |
| Saturday night | Train on its own outputs, run the size sweep (runs by itself) | You have the size chart |
| Sunday morning | Convert to 8-bit, write the C engine, flash the chip | The game plays on the board |
| Sunday afternoon | Browser version, record the video, write the post | The post is ready |

## What you need

- **Your Shrike-Lite board**, or any RP2040 board, such as a $4 Raspberry Pi Pico.
- **A USB cable.** That's all the hardware you need.
- **Optional:** a small SSD1306 OLED screen (about $3), so the board has its own display.

## Cost

- **GPU:** 1× RTX 5090 on RunPod, about $0.69 per hour. The fallback is an RTX 4090, about
  $0.34 per hour. The models are tiny, so you don't need a big GPU.
- **Time on the GPU:** about 8-12 hours.
- **Total:** about $5-10. The plan sets a hard stop at $25.

## Your VMs

- **`vm.ifkash.dev`:** hosts the browser version. It already runs Dokploy, so the game gets
  deployed as one more app there. It doesn't touch your other apps.
- **`vm.dotslasha.me`:** not needed. Its disk is 93% full, and both VMs are too small (2 CPUs,
  4 GB of RAM, no GPU) for training.

## What you have at the end

- **A board that plays Breakout with no game code on it.** That's the video.
- **A browser version** at something like `dreamcart.ifkash.dev`.
- **The chart:** model size against "does it still follow the game's rules?"
- **A side-by-side mode** that shows the real game and the dreamed game getting the same
  button presses.
- **A blog post:** *"There's no game on this chip."*

## What might go wrong

| Problem | What to do |
|---|---|
| The chip is too slow, under 10 frames per second | Use a smaller model, or run it at 5-8 frames per second. It still makes a great demo. |
| The tiny model breaks the rules, for example the ball goes through bricks | Use the next size up. The chart tells you which size is safe. |
| The game "melts" after a few seconds | That's what step 3 fixes. The plan has numbers to check for it. |
| The 8-bit version plays worse than the full version | Retrain it with 8-bit in mind. The plan covers this. |

## Stretch goals

- **Use the FPGA on your board** to speed up the math. You mapped this out in the vicharak
  project but never started it.
- **Dream a second game,** such as Pong or Snake, with the same code.
- **Fool test:** show friends the real game and the dreamed game and see whether they can tell
  which is which.

## Other ideas I thought about

- **A world model in the browser.** Already done: people have shipped Flappy Bird and
  SuperTuxKart. Putting one on a chip is new.
- **Train an LLM across the internet on spare GPUs.** Already done at scale by Prime Intellect.
- **Byte-level smol-llama with no tokenizer.** Solid, but it doesn't make a good demo.

## Before you start

- **Rotate your RunPod API key.** You pasted it in chat. This file doesn't contain it.
- The earlier `pleasescope.md` is still in this folder. Delete it if you don't want it.

## Reading, if you want it

- [Diffusion models are real-time game engines](https://arxiv.org/abs/2408.14837): GameNGen,
  which ran Doom in a neural network.
- [Self Forcing](https://arxiv.org/abs/2506.08009): training on your own outputs so rollouts
  don't drift.
- [Optimizing a Flappy Bird world model for the browser](https://www.njkumar.com/optimizing-flappy-bird-world-model-to-run-in-a-web-browser/):
  the closest thing to this project, at 5 million parameters in a browser.

---

# Part 2: For the coding agent

## Mission

Build `dreamcart`, an action-conditioned neural world model of a tiny Breakout game. It must run
in real time on an RP2040 microcontroller with no game logic on the device. It must also run in
a browser from the same C inference code, compiled to WebAssembly. The research deliverable is
a curve that shows world-model size against rule fidelity.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a
milestone until the previous one passes. After each milestone, commit your work and write a
short entry in `NOTES.md` with the results and numbers.

## Hard constraints

- **Secrets:** Read `RUNPOD_API_KEY`, `HF_TOKEN`, and `WANDB_API_KEY` from environment
  variables only. Never write a secret into any file, log, commit, notebook output, or command
  line that's echoed. Commit a `.env.example` that has placeholder values only, and add `.env`
  to `.gitignore`.
- **Budget:** The hard cap is $25 of RunPod spend. Every pod runs a watchdog that stops the pod
  after `MAX_POD_HOURS` hours. The default is `6`.
- **Compute:** Run training only on RunPod. The local Mac runs the game prototype, the
  firmware build, the flashing, the WASM build, and the tests.
- **Outward actions:** Ask the user before you do any of the following: create the GitHub
  repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything
  publicly. Use the `kashifulhaque` GitHub account (`gh auth switch -u kashifulhaque`), not
  `kashif-wandai`.
- **Shared VM:** `vm.ifkash.dev` runs other production apps under Dokploy, with Traefik on
  ports 80 and 443. Don't stop, restart, reconfigure, or remove any existing container,
  service, or proxy rule. Add dreamcart only as a new Dokploy application.
- **Pod cleanup:** When a pod isn't running a job, stop it. At the end of the project,
  terminate every pod that you created and report the total spend.

## Target hardware facts

- **Chip:** RP2040, with two Arm Cortex-M0+ cores rated at 133 MHz. It has 264 KB of SRAM, no
  FPU, no SIMD, and a single-cycle 32-bit integer multiplier. The weights live in external QSPI
  flash and are read through the execute-in-place (XIP) cache, which is 16 KB.
- **User's board:** Vicharak Shrike-Lite, an RP2040 plus a Renesas SLG47910 FPGA. Any RP2040
  board, such as a Raspberry Pi Pico, runs the same firmware. Keep the firmware free of
  anything specific to one board, except for an optional pin configuration file.
- **Toolchain:** Pico SDK 2.x, `arm-none-eabi-gcc`, and CMake. Check the SDK documentation for
  the supported maximum system clock. Make the clock a build option that defaults to the SDK
  default, and record the measured frames per second at each clock you test.
- **I/O:** USB CDC serial. The device reads single keypresses: `a` moves left, `d` moves
  right, and `r` resets. Any other key, or no key, means no-op. The device writes ANSI frames.

## Budgets for the device model

These budgets are hard limits for the model that runs on the chip:

| Resource | Limit |
|---|---|
| int8 weights, copied into SRAM at boot | 128 KB or less |
| Peak activation memory | 64 KB or less |
| MACs per frame | 3,000,000 or less |
| Target frame rate | 10 fps or more; 15 fps is the goal |

## Tech stack

Pin every version. The stack is as follows:

- **Python:** Python 3.11, PyTorch 2.x, `numpy`, `wandb`, and `matplotlib`, managed with `uv`.
- **Firmware:** C11 with the Pico SDK. Don't use a machine-learning runtime on the device.
  Write the int8 kernels by hand.
- **Web:** Emscripten for the C-to-WASM build, plain HTML, a canvas, and JavaScript. Don't use
  a frontend framework.
- **Pods:** the `runpod` Python SDK, with `runpodctl` inside pods.

## Repository layout

Create the following layout:

```
dreamcart/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh
  dreamcart/
    game.py          # batched torch Breakout, the ground truth
    render.py        # palette, ANSI and PNG rendering, GIF export
    policy.py        # data-collection policies
    model.py         # world-model architectures, parameterized by size
    train.py         # teacher-forced training + rollout training
    evaluate.py      # frames-to-divergence, rule-validity checker
    quant.py         # QAT, int8 export, bit-exact int8 simulator
    export_c.py      # writes weights.h + golden test vectors
  engine/            # portable C inference engine (host, RP2040, WASM)
    dc_engine.h  dc_engine.c  weights.h  test_golden.c
  firmware/          # Pico SDK project, uses engine/
  web/               # index.html, main.js, build.sh (emcc)
  deploy/            # Dockerfile (static nginx) + Dokploy notes
  results/           # JSON metrics, plots, GIFs (commit these)
  post/draft.md
```

## M0: Infrastructure

**Tasks:**

1. Write `infra/pod.py`. It uses the `runpod` SDK and reads `RUNPOD_API_KEY` from the
   environment. It supports the following subcommands:
   - `create`: Creates a pod. The GPU preference order is RTX 5090, then RTX 4090, then RTX PRO
     4500. Resolve the GPU type IDs at run time by querying the available GPU types. Use an
     official RunPod PyTorch image with CUDA 12.8 or later. Set a 50 GB volume at `/workspace`
     and expose SSH.
   - `status`, `stop`, and `terminate`.
   - `ssh-info`: Prints the SSH command.
   - `cost`: Prints the uptime and spend for every pod whose name has the prefix `dreamcart-`.
2. Write `infra/watchdog.sh`. It sleeps for `MAX_POD_HOURS`, then runs
   `runpodctl stop pod $RUNPOD_POD_ID`.
3. Write `infra/bootstrap.sh`. It clones the repo, installs `uv`, runs `uv sync`, and starts
   the watchdog.

**Acceptance criteria:**

- A pod comes up, and `bootstrap.sh` finishes without errors.
- A test run with `MAX_POD_HOURS=0.05` stops the pod within 5 minutes.

## M1: Ground-truth game

**Game spec:**

The game uses the following rules. Implement them exactly:

- **Frame:** 32×32 pixels. Each pixel has one of 5 classes: `0` background, `1` wall, `2`
  brick, `3` paddle, and `4` ball.
- **Walls:** Row 0, column 0, and column 31 are walls. The bottom edge is open.
- **Bricks:** Rows 3-8 and columns 2-29 form the brick region. The region holds bricks that are
  4 pixels wide and 2 pixels tall on a fixed grid, 7 bricks across and 3 rows. Store the brick
  state as a 21-bit mask.
- **Paddle:** 6 pixels wide on row 29. Its x position is clamped to the area inside the walls.
  The actions are `0` no-op, `1` left, and `2` right, and each moves the paddle 1 pixel.
- **Ball:** 1 pixel. Its velocity is `(dx, dy)`, with each component in {-1, +1}.
  Every frame, the ball moves one pixel on each axis, and the game resolves collisions with
  integer rules:
  - A wall reflects the matching velocity component.
  - A brick is removed, and the ball reflects `dy`.
  - The paddle reflects `dy`. If the ball hits the left third of the paddle, `dx` becomes -1.
    If it hits the right third, `dx` becomes +1. If it hits the middle, `dx` stays the same.
  - Document the order in which collisions resolve, and keep it fixed.
- **Miss:** If the ball passes row 31, the next frame respawns the ball at a fixed position,
  such as `(16, 20)`, with a fixed velocity of `(+1, -1)`.
- **Board clear:** When the brick mask is empty, the next frame refills every brick.
- **No score and no lives.** Every piece of the state must be visible in the frame.

**The Markov property:** The next frame must be a deterministic function of the previous
frame, the current frame, and the action. The velocity comes from the difference between the
two frames. Write a test that runs 10 million random transitions. The test hashes each tuple
`(f_prev, f_cur, a)` and checks that no two identical tuples produce different next frames. If
the test finds a collision, fix the rules. Respawn and refill are the most likely places.

**Implementation:**

- `game.py` holds `N` games as integer tensors on the GPU and steps all of them together.
- Starting states are randomized for coverage. Each game gets a random brick mask with
  Bernoulli probability 0.3-1.0 per brick, a random paddle x, a random ball position in the free
  region, and a random velocity.
- The policies in `policy.py` are as follows:
  - Uniform random actions.
  - Sticky-random actions: the policy repeats its action with probability 0.8.
  - A tracker that follows the ball x position, with noise.
  - A mix of all three, chosen per episode.
- Write `render.py` to show the game in a terminal with ANSI half-block characters, which gives
  32 columns and 16 rows, and to export PNG and GIF files.

**Acceptance criteria:**

- The Markov test passes.
- Stepping 4,096 games on the GPU reaches 1 million or more transitions per second.
- A GIF of 300 frames of tracker play, `results/m1_real.gif`, looks correct.

## M2: Teacher-forced world model

**Architecture:**

- **Input:** A one-hot encoding of `f_prev` and `f_cur`, which gives 10 channels at 32×32.
  The action is injected as a learned bias per channel at each resolution, which is cheap on
  the device.
- **Body:** A small U-Net. Downsample with stride-2 3×3 convolutions from 32 to 16 to 8. Use
  residual 3×3 blocks at 8×8 and 16×16, and upsample with nearest neighbor and a 3×3
  convolution, with skip connections that add rather than concatenate. The output is 5 logits
  per pixel.
- **Activation:** ReLU only, because the device uses integer ReLU. Don't use normalization
  layers at inference. Fold any normalization into the convolutions before export.
- **Size:** A single width multiplier and depth setting control the size. Build 7 presets that
  target about 8K, 16K, 32K, 64K, 128K, 256K, and 1M parameters. Report the parameter count and
  the MACs per frame for each preset.

**Training:**

- Stream the data directly from `game.py`. Don't store any data on disk.
- Use cross-entropy on the next frame. Weight pixels that change between `f_cur` and `f_next`
  by 20×, because static pixels dominate otherwise.
- Use AdamW, cosine decay, and bf16 autocast. Log to wandb.

**Evaluation (`evaluate.py`):**

- **One-step metrics:** Pixel accuracy on changed pixels, and the exact-frame match rate.
- **Frames-to-divergence:** Run 500 episodes of 2,000 frames each. The real game and the model
  receive the same actions, and the model gets only the first 2 real frames. Report the median
  number of frames until the first frame that doesn't match exactly.
- **Rule-validity rate:** Parse each dreamed frame into a state, then check the following:
  - The frame has exactly 1 ball pixel. A miss respawns the ball in the next frame, so no
    valid frame has 0 balls.
  - The paddle is exactly 6 contiguous pixels on row 29.
  - Bricks are whole bricks in their grid slots.
  - The walls are intact.
  - The transition from the previous dreamed frame to this one is exactly what the real rules
    produce from the parsed state.

  The rule-validity rate is the fraction of frames that pass every check. This is the main
  metric, because a dreamed game can diverge from the real game and still be a valid game.

**Acceptance criteria:**

- The 1M preset reaches an exact-frame match rate of 99.9% or more for one step, and a rule
  validity of 99% or more over 2,000-frame rollouts.

## M3: Rollout training

Make rollouts stable, in the spirit of Self Forcing:

- **Context-noise augmentation:** Flip each context pixel to a random class with probability
  `p`. Sample `p` for each sample from the range 0 to 2%.
- **Unrolled training:** Starting from real frames, roll the model forward `K` steps on its own
  predictions. Feed back the argmax with a straight-through estimator, or feed back softmax
  probabilities. Compare the two options and record the winner. Compute the loss at every step
  against the real trajectory under the same actions. Use a curriculum for `K`: 2, then 4, 8,
  and 16. Truncate the gradient to the last 4 steps if memory requires it.

**Acceptance criteria:**

- For every preset from 64K up, median frames-to-divergence and rule validity improve over M2.
  Record the before and after numbers in `NOTES.md`.
- The 1M preset keeps a rule validity of 99.5% or more over 10,000-frame rollouts.

## M4: Size sweep

**Tasks:**

- Train all 7 presets with the full M2 and M3 recipe, with the same token budget per preset.
  Run the presets concurrently on one GPU if they fit.
- Plot `results/m4_scaling.png`. It shows rule validity and median frames-to-divergence
  against the parameter count, on a log x-axis, with a second axis for MACs per frame.
- Mark the device budget of 128 KB of weights and 3M MACs on the plot.
- Select the device model: the smallest preset within budget that reaches a rule validity of
  99% or more over 2,000-frame rollouts. If no preset meets both conditions, choose the preset
  with the best rule validity within budget, and record the gap.

**Acceptance criteria:**

- The plot, `results/m4_sweep.json`, and a side-by-side GIF for each preset exist. Each GIF
  shows the real game and the dreamed game under the same actions.

## M5: int8 quantization

**Tasks:**

- Quantize the weights to int8 with symmetric, per-output-channel scales. Quantize the
  activations to int8 with per-tensor scales calibrated on 10,000 frames. Accumulate in int32,
  and requantize with a fixed-point multiplier and shift, in the style of TFLite. Use int32 for
  biases, and use a uint8 lookup table for each action bias.
- Write `quant.py` with an exact integer simulator in numpy that uses the same arithmetic the C
  code uses. If the int8 model loses more than 0.5 points of rule validity compared with float,
  run quantization-aware training (QAT) with fake-quant, and repeat.
- Write `export_c.py` to produce `engine/weights.h` and golden vectors: 200 input tuples with
  the expected logits argmax and the expected intermediate activation checksums.

**Acceptance criteria:**

- The int8 simulator stays within 0.5 points of float rule validity.
- The golden vectors exist.

## M6: C engine and firmware

**Tasks:**

- Write `engine/`: portable C11 with no dynamic allocation. Use static arena buffers sized from
  the model config. The public API is
  `dc_init()`, `dc_step(const uint8_t *f_prev, const uint8_t *f_cur, uint8_t action, uint8_t *f_next)`.
  The engine runs host-side tests with `make test`.
- The host build must match the golden vectors exactly, bit for bit.
- Optimize for the M0+ in the following ways:
  - Copy the weights from flash to SRAM at boot.
  - Unroll the inner loops by 4, and keep hot loops in RAM (`__not_in_flash_func`).
  - Split the output rows of each layer across both cores, synchronized with the
    `multicore_fifo` or a barrier.
  - Keep the `f_prev` and `f_cur` ring buffer on the device.
- Write `firmware/`:
  - It boots into the dreamed game from a fixed start frame, embedded as a constant.
  - It polls USB CDC for keys without blocking.
  - It sends ANSI output that redraws only the cells that changed. It prints a status line with
    the frame rate and the microseconds per frame, measured with the hardware timer.
  - It caps the frame rate at 15 fps.
  - Pressing `r` resets to the start frame.
- Write a host script, `firmware/play.sh`, that opens the serial port in raw mode, for example
  with `picocom` or a small Python script that uses `pyserial`.
- Print the flash and SRAM usage from the linker map in `NOTES.md`.

**Acceptance criteria:**

- The device output matches the host engine, bit for bit, on a scripted sequence of 1,000
  frames. The device can print a CRC of each frame in a debug mode for this check.
- The measured frame rate is 10 fps or more at the SDK default clock. If it isn't, report the
  frame rate at the highest clock that the SDK supports. Also try the next smaller preset, and
  record the frame rates for both presets in `NOTES.md`.
- Pause for the user: the user needs to flash the board and play. Give the exact build and
  flash commands.

## M7: Browser build and deploy

**Tasks:**

- In `web/build.sh`, compile the same `engine/` sources with `emcc -O3`, and export `dc_step`.
- Write `web/index.html` and `web/main.js`:
  - A canvas that renders at 32×32 and scales up 12× with nearest-neighbor scaling.
  - Arrow keys and A/D keys for control, and touch buttons for phones.
  - A **Side by side** toggle that also runs a JavaScript port of the real game, or the Python
    rules compiled to C and then to WASM. Both games receive the same inputs.
  - A frame-rate counter.
- The WASM output must match the golden vectors. Add a self-test that runs when the page loads,
  in debug mode.
- Write `deploy/Dockerfile`: `nginx:alpine` serving `web/dist`.
- Ask the user before deploying. After the user approves, create a new Dokploy application on
  `vm.ifkash.dev`, with a domain such as `dreamcart.ifkash.dev` that the user confirms. Don't
  touch existing apps. The user might need to add the DNS record.

**Acceptance criteria:**

- The game plays locally in Chrome and Safari at 15 fps or more.
- After deployment, the page loads at the confirmed URL.

## M8: Write-up

**Tasks:**

- Write `README.md`. It explains what the project is, how to reproduce each milestone with one
  command per milestone, the device build and flash steps, the sweep plot, and GIFs.
- Write `post/draft.md`, a blog post of 1,200-1,800 words for ifkash.dev. Structure it as
  follows:
  - The hook: "There's no game on this chip."
  - Why world models are the topic of the year, and why you went small.
  - The Markov-in-two-frames design.
  - Drift, and how rollout training fixed it, with before and after GIFs.
  - The scaling chart: how many parameters it takes to be Breakout.
  - Squeezing the model into 264 KB: the int8 path, dual-core, and measured frame rates.
  - Play it yourself: the link to the browser version.
  - Limitations and next steps: the FPGA MAC unit, other games, and stochastic games.
- Ask the user before you push anything. After the user approves, push the repo to
  `weights-and-wires/dreamcart`, and upload the checkpoints for every preset to the Hugging Face
  Hub under `weights-and-wires/dreamcart`.
- Terminate every pod, and write the final spend to `NOTES.md`.

## Final report to the user

When you finish, report the following:

- The selected device preset, with its parameter count, MACs, and weight size.
- The frame rate on the device and in the browser.
- Rule validity and frames-to-divergence for float and int8.
- The sweep plot.
- The total spend.
- Anything that you skipped or that failed.
