← projects

dreamcart

A video game with no game code, dreamed by a $4 chip

18 min read dreamcart.md

Part 1: For you

The idea in one line

You train a tiny neural network to be a video game, then run it on your RP2040 board. The chip has no game code. Every frame you see, the network imagines.

Why it's exciting

  • World models are the hottest topic in AI right now. Genie 3, Matrix-Game, and SANA-WM are all neural networks that imagine an interactive world frame by frame. They all need big GPUs.
  • You go the other way. You find the smallest world model that can still run a game, and you run it on a chip with 264 KB of RAM. I couldn't find anyone who has done this.
  • Only you can do it quickly. You already got an LLM running on this exact board. This is the same skill, pointed at the hottest topic of the year.
  • The demo is easy to share. A video of a game running on a tiny board, with the caption "there's no game on this chip," is very shareable.
  • It also has a real research result. You get a chart that answers the question: how many parameters does it take to be Breakout?

How it works

  1. Write the real game. It's a tiny Breakout: 32×32 pixels, a paddle, a ball, and bricks. You write it in PyTorch so the GPU can run 4,096 games at once. That gives you endless training data and nothing to store on disk.
  2. Train the dreamer. A small network sees the last 2 frames and your button press, and draws the next frame.
  3. Teach it not to drift. Small mistakes pile up over time, and the game "melts." You fix this by training the model on its own outputs, a trick from the Self Forcing paper.
  4. Shrink it. You train 7 sizes, from 8,000 to 1 million parameters, and find the smallest one that still plays by the rules.
  5. Put it on the chip. Convert the model to 8-bit integers, write the math in plain C, and flash it to the RP2040. Your terminal becomes the screen, and your keyboard becomes the controller.
  6. Put it on the web. Compile the same C code for the browser, so anyone can play it. Host it on vm.ifkash.dev.

Weekend plan

When What Done when
Saturday morning Write the real game and check it on your Mac The game runs and looks right
Saturday afternoon Train the dreamer on a GPU The dreamed game looks like the real one
Saturday night Train on its own outputs, run the size sweep (runs by itself) You have the size chart
Sunday morning Convert to 8-bit, write the C engine, flash the chip The game plays on the board
Sunday afternoon Browser version, record the video, write the post The post is ready

What you need

  • Your Shrike-Lite board, or any RP2040 board, such as a $4 Raspberry Pi Pico.
  • A USB cable. That's all the hardware you need.
  • Optional: a small SSD1306 OLED screen (about $3), so the board has its own display.

Cost

  • GPU: 1× RTX 5090 on RunPod, about $0.69 per hour. The fallback is an RTX 4090, about $0.34 per hour. The models are tiny, so you don't need a big GPU.
  • Time on the GPU: about 8-12 hours.
  • Total: about $5-10. The plan sets a hard stop at $25.

Your VMs

  • vm.ifkash.dev: hosts the browser version. It already runs Dokploy, so the game gets deployed as one more app there. It doesn't touch your other apps.
  • vm.dotslasha.me: not needed. Its disk is 93% full, and both VMs are too small (2 CPUs, 4 GB of RAM, no GPU) for training.

What you have at the end

  • A board that plays Breakout with no game code on it. That's the video.
  • A browser version at something like dreamcart.ifkash.dev.
  • The chart: model size against "does it still follow the game's rules?"
  • A side-by-side mode that shows the real game and the dreamed game getting the same button presses.
  • A blog post: "There's no game on this chip."

What might go wrong

Problem What to do
The chip is too slow, under 10 frames per second Use a smaller model, or run it at 5-8 frames per second. It still makes a great demo.
The tiny model breaks the rules, for example the ball goes through bricks Use the next size up. The chart tells you which size is safe.
The game "melts" after a few seconds That's what step 3 fixes. The plan has numbers to check for it.
The 8-bit version plays worse than the full version Retrain it with 8-bit in mind. The plan covers this.

Stretch goals

  • Use the FPGA on your board to speed up the math. You mapped this out in the vicharak project but never started it.
  • Dream a second game, such as Pong or Snake, with the same code.
  • Fool test: show friends the real game and the dreamed game and see whether they can tell which is which.

Other ideas I thought about

  • A world model in the browser. Already done: people have shipped Flappy Bird and SuperTuxKart. Putting one on a chip is new.
  • Train an LLM across the internet on spare GPUs. Already done at scale by Prime Intellect.
  • Byte-level smol-llama with no tokenizer. Solid, but it doesn't make a good demo.

Before you start

  • Rotate your RunPod API key. You pasted it in chat. This file doesn't contain it.
  • The earlier pleasescope.md is still in this folder. Delete it if you don't want it.

Reading, if you want it


Part 2: For the coding agent

Mission

Build dreamcart, an action-conditioned neural world model of a tiny Breakout game. It must run in real time on an RP2040 microcontroller with no game logic on the device. It must also run in a browser from the same C inference code, compiled to WebAssembly. The research deliverable is a curve that shows world-model size against rule fidelity.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a milestone until the previous one passes. After each milestone, commit your work and write a short entry in NOTES.md with the results and numbers.

Hard constraints

  • Secrets: Read RUNPOD_API_KEY, HF_TOKEN, and WANDB_API_KEY from environment variables only. Never write a secret into any file, log, commit, notebook output, or command line that's echoed. Commit a .env.example that has placeholder values only, and add .env to .gitignore.
  • Budget: The hard cap is $25 of RunPod spend. Every pod runs a watchdog that stops the pod after MAX_POD_HOURS hours. The default is 6.
  • Compute: Run training only on RunPod. The local Mac runs the game prototype, the firmware build, the flashing, the WASM build, and the tests.
  • Outward actions: Ask the user before you do any of the following: create the GitHub repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything publicly. Use the kashifulhaque GitHub account (gh auth switch -u kashifulhaque), not kashif-wandai.
  • Shared VM: vm.ifkash.dev runs other production apps under Dokploy, with Traefik on ports 80 and 443. Don't stop, restart, reconfigure, or remove any existing container, service, or proxy rule. Add dreamcart only as a new Dokploy application.
  • Pod cleanup: When a pod isn't running a job, stop it. At the end of the project, terminate every pod that you created and report the total spend.

Target hardware facts

  • Chip: RP2040, with two Arm Cortex-M0+ cores rated at 133 MHz. It has 264 KB of SRAM, no FPU, no SIMD, and a single-cycle 32-bit integer multiplier. The weights live in external QSPI flash and are read through the execute-in-place (XIP) cache, which is 16 KB.
  • User's board: Vicharak Shrike-Lite, an RP2040 plus a Renesas SLG47910 FPGA. Any RP2040 board, such as a Raspberry Pi Pico, runs the same firmware. Keep the firmware free of anything specific to one board, except for an optional pin configuration file.
  • Toolchain: Pico SDK 2.x, arm-none-eabi-gcc, and CMake. Check the SDK documentation for the supported maximum system clock. Make the clock a build option that defaults to the SDK default, and record the measured frames per second at each clock you test.
  • I/O: USB CDC serial. The device reads single keypresses: a moves left, d moves right, and r resets. Any other key, or no key, means no-op. The device writes ANSI frames.

Budgets for the device model

These budgets are hard limits for the model that runs on the chip:

Resource Limit
int8 weights, copied into SRAM at boot 128 KB or less
Peak activation memory 64 KB or less
MACs per frame 3,000,000 or less
Target frame rate 10 fps or more; 15 fps is the goal

Tech stack

Pin every version. The stack is as follows:

  • Python: Python 3.11, PyTorch 2.x, numpy, wandb, and matplotlib, managed with uv.
  • Firmware: C11 with the Pico SDK. Don't use a machine-learning runtime on the device. Write the int8 kernels by hand.
  • Web: Emscripten for the C-to-WASM build, plain HTML, a canvas, and JavaScript. Don't use a frontend framework.
  • Pods: the runpod Python SDK, with runpodctl inside pods.

Repository layout

Create the following layout:

dreamcart/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh
  dreamcart/
    game.py          # batched torch Breakout, the ground truth
    render.py        # palette, ANSI and PNG rendering, GIF export
    policy.py        # data-collection policies
    model.py         # world-model architectures, parameterized by size
    train.py         # teacher-forced training + rollout training
    evaluate.py      # frames-to-divergence, rule-validity checker
    quant.py         # QAT, int8 export, bit-exact int8 simulator
    export_c.py      # writes weights.h + golden test vectors
  engine/            # portable C inference engine (host, RP2040, WASM)
    dc_engine.h  dc_engine.c  weights.h  test_golden.c
  firmware/          # Pico SDK project, uses engine/
  web/               # index.html, main.js, build.sh (emcc)
  deploy/            # Dockerfile (static nginx) + Dokploy notes
  results/           # JSON metrics, plots, GIFs (commit these)
  post/draft.md

M0: Infrastructure

Tasks:

  1. Write infra/pod.py. It uses the runpod SDK and reads RUNPOD_API_KEY from the environment. It supports the following subcommands: - create: Creates a pod. The GPU preference order is RTX 5090, then RTX 4090, then RTX PRO
    1. Resolve the GPU type IDs at run time by querying the available GPU types. Use an official RunPod PyTorch image with CUDA 12.8 or later. Set a 50 GB volume at /workspace and expose SSH. - status, stop, and terminate. - ssh-info: Prints the SSH command. - cost: Prints the uptime and spend for every pod whose name has the prefix dreamcart-.
  2. Write infra/watchdog.sh. It sleeps for MAX_POD_HOURS, then runs runpodctl stop pod $RUNPOD_POD_ID.
  3. Write infra/bootstrap.sh. It clones the repo, installs uv, runs uv sync, and starts the watchdog.

Acceptance criteria:

  • A pod comes up, and bootstrap.sh finishes without errors.
  • A test run with MAX_POD_HOURS=0.05 stops the pod within 5 minutes.

M1: Ground-truth game

Game spec:

The game uses the following rules. Implement them exactly:

  • Frame: 32×32 pixels. Each pixel has one of 5 classes: 0 background, 1 wall, 2 brick, 3 paddle, and 4 ball.
  • Walls: Row 0, column 0, and column 31 are walls. The bottom edge is open.
  • Bricks: Rows 3-8 and columns 2-29 form the brick region. The region holds bricks that are 4 pixels wide and 2 pixels tall on a fixed grid, 7 bricks across and 3 rows. Store the brick state as a 21-bit mask.
  • Paddle: 6 pixels wide on row 29. Its x position is clamped to the area inside the walls. The actions are 0 no-op, 1 left, and 2 right, and each moves the paddle 1 pixel.
  • Ball: 1 pixel. Its velocity is (dx, dy), with each component in {-1, +1}. Every frame, the ball moves one pixel on each axis, and the game resolves collisions with integer rules:
  • A wall reflects the matching velocity component.
  • A brick is removed, and the ball reflects dy.
  • The paddle reflects dy. If the ball hits the left third of the paddle, dx becomes -1. If it hits the right third, dx becomes +1. If it hits the middle, dx stays the same.
  • Document the order in which collisions resolve, and keep it fixed.
  • Miss: If the ball passes row 31, the next frame respawns the ball at a fixed position, such as (16, 20), with a fixed velocity of (+1, -1).
  • Board clear: When the brick mask is empty, the next frame refills every brick.
  • No score and no lives. Every piece of the state must be visible in the frame.

The Markov property: The next frame must be a deterministic function of the previous frame, the current frame, and the action. The velocity comes from the difference between the two frames. Write a test that runs 10 million random transitions. The test hashes each tuple (f_prev, f_cur, a) and checks that no two identical tuples produce different next frames. If the test finds a collision, fix the rules. Respawn and refill are the most likely places.

Implementation:

  • game.py holds N games as integer tensors on the GPU and steps all of them together.
  • Starting states are randomized for coverage. Each game gets a random brick mask with Bernoulli probability 0.3-1.0 per brick, a random paddle x, a random ball position in the free region, and a random velocity.
  • The policies in policy.py are as follows:
  • Uniform random actions.
  • Sticky-random actions: the policy repeats its action with probability 0.8.
  • A tracker that follows the ball x position, with noise.
  • A mix of all three, chosen per episode.
  • Write render.py to show the game in a terminal with ANSI half-block characters, which gives 32 columns and 16 rows, and to export PNG and GIF files.

Acceptance criteria:

  • The Markov test passes.
  • Stepping 4,096 games on the GPU reaches 1 million or more transitions per second.
  • A GIF of 300 frames of tracker play, results/m1_real.gif, looks correct.

M2: Teacher-forced world model

Architecture:

  • Input: A one-hot encoding of f_prev and f_cur, which gives 10 channels at 32×32. The action is injected as a learned bias per channel at each resolution, which is cheap on the device.
  • Body: A small U-Net. Downsample with stride-2 3×3 convolutions from 32 to 16 to 8. Use residual 3×3 blocks at 8×8 and 16×16, and upsample with nearest neighbor and a 3×3 convolution, with skip connections that add rather than concatenate. The output is 5 logits per pixel.
  • Activation: ReLU only, because the device uses integer ReLU. Don't use normalization layers at inference. Fold any normalization into the convolutions before export.
  • Size: A single width multiplier and depth setting control the size. Build 7 presets that target about 8K, 16K, 32K, 64K, 128K, 256K, and 1M parameters. Report the parameter count and the MACs per frame for each preset.

Training:

  • Stream the data directly from game.py. Don't store any data on disk.
  • Use cross-entropy on the next frame. Weight pixels that change between f_cur and f_next by 20×, because static pixels dominate otherwise.
  • Use AdamW, cosine decay, and bf16 autocast. Log to wandb.

Evaluation (evaluate.py):

  • One-step metrics: Pixel accuracy on changed pixels, and the exact-frame match rate.
  • Frames-to-divergence: Run 500 episodes of 2,000 frames each. The real game and the model receive the same actions, and the model gets only the first 2 real frames. Report the median number of frames until the first frame that doesn't match exactly.
  • Rule-validity rate: Parse each dreamed frame into a state, then check the following:
  • The frame has exactly 1 ball pixel. A miss respawns the ball in the next frame, so no valid frame has 0 balls.
  • The paddle is exactly 6 contiguous pixels on row 29.
  • Bricks are whole bricks in their grid slots.
  • The walls are intact.
  • The transition from the previous dreamed frame to this one is exactly what the real rules produce from the parsed state.

The rule-validity rate is the fraction of frames that pass every check. This is the main metric, because a dreamed game can diverge from the real game and still be a valid game.

Acceptance criteria:

  • The 1M preset reaches an exact-frame match rate of 99.9% or more for one step, and a rule validity of 99% or more over 2,000-frame rollouts.

M3: Rollout training

Make rollouts stable, in the spirit of Self Forcing:

  • Context-noise augmentation: Flip each context pixel to a random class with probability p. Sample p for each sample from the range 0 to 2%.
  • Unrolled training: Starting from real frames, roll the model forward K steps on its own predictions. Feed back the argmax with a straight-through estimator, or feed back softmax probabilities. Compare the two options and record the winner. Compute the loss at every step against the real trajectory under the same actions. Use a curriculum for K: 2, then 4, 8, and 16. Truncate the gradient to the last 4 steps if memory requires it.

Acceptance criteria:

  • For every preset from 64K up, median frames-to-divergence and rule validity improve over M2. Record the before and after numbers in NOTES.md.
  • The 1M preset keeps a rule validity of 99.5% or more over 10,000-frame rollouts.

M4: Size sweep

Tasks:

  • Train all 7 presets with the full M2 and M3 recipe, with the same token budget per preset. Run the presets concurrently on one GPU if they fit.
  • Plot results/m4_scaling.png. It shows rule validity and median frames-to-divergence against the parameter count, on a log x-axis, with a second axis for MACs per frame.
  • Mark the device budget of 128 KB of weights and 3M MACs on the plot.
  • Select the device model: the smallest preset within budget that reaches a rule validity of 99% or more over 2,000-frame rollouts. If no preset meets both conditions, choose the preset with the best rule validity within budget, and record the gap.

Acceptance criteria:

  • The plot, results/m4_sweep.json, and a side-by-side GIF for each preset exist. Each GIF shows the real game and the dreamed game under the same actions.

M5: int8 quantization

Tasks:

  • Quantize the weights to int8 with symmetric, per-output-channel scales. Quantize the activations to int8 with per-tensor scales calibrated on 10,000 frames. Accumulate in int32, and requantize with a fixed-point multiplier and shift, in the style of TFLite. Use int32 for biases, and use a uint8 lookup table for each action bias.
  • Write quant.py with an exact integer simulator in numpy that uses the same arithmetic the C code uses. If the int8 model loses more than 0.5 points of rule validity compared with float, run quantization-aware training (QAT) with fake-quant, and repeat.
  • Write export_c.py to produce engine/weights.h and golden vectors: 200 input tuples with the expected logits argmax and the expected intermediate activation checksums.

Acceptance criteria:

  • The int8 simulator stays within 0.5 points of float rule validity.
  • The golden vectors exist.

M6: C engine and firmware

Tasks:

  • Write engine/: portable C11 with no dynamic allocation. Use static arena buffers sized from the model config. The public API is dc_init(), dc_step(const uint8_t *f_prev, const uint8_t *f_cur, uint8_t action, uint8_t *f_next). The engine runs host-side tests with make test.
  • The host build must match the golden vectors exactly, bit for bit.
  • Optimize for the M0+ in the following ways:
  • Copy the weights from flash to SRAM at boot.
  • Unroll the inner loops by 4, and keep hot loops in RAM (__not_in_flash_func).
  • Split the output rows of each layer across both cores, synchronized with the multicore_fifo or a barrier.
  • Keep the f_prev and f_cur ring buffer on the device.
  • Write firmware/:
  • It boots into the dreamed game from a fixed start frame, embedded as a constant.
  • It polls USB CDC for keys without blocking.
  • It sends ANSI output that redraws only the cells that changed. It prints a status line with the frame rate and the microseconds per frame, measured with the hardware timer.
  • It caps the frame rate at 15 fps.
  • Pressing r resets to the start frame.
  • Write a host script, firmware/play.sh, that opens the serial port in raw mode, for example with picocom or a small Python script that uses pyserial.
  • Print the flash and SRAM usage from the linker map in NOTES.md.

Acceptance criteria:

  • The device output matches the host engine, bit for bit, on a scripted sequence of 1,000 frames. The device can print a CRC of each frame in a debug mode for this check.
  • The measured frame rate is 10 fps or more at the SDK default clock. If it isn't, report the frame rate at the highest clock that the SDK supports. Also try the next smaller preset, and record the frame rates for both presets in NOTES.md.
  • Pause for the user: the user needs to flash the board and play. Give the exact build and flash commands.

M7: Browser build and deploy

Tasks:

  • In web/build.sh, compile the same engine/ sources with emcc -O3, and export dc_step.
  • Write web/index.html and web/main.js:
  • A canvas that renders at 32×32 and scales up 12× with nearest-neighbor scaling.
  • Arrow keys and A/D keys for control, and touch buttons for phones.
  • A Side by side toggle that also runs a JavaScript port of the real game, or the Python rules compiled to C and then to WASM. Both games receive the same inputs.
  • A frame-rate counter.
  • The WASM output must match the golden vectors. Add a self-test that runs when the page loads, in debug mode.
  • Write deploy/Dockerfile: nginx:alpine serving web/dist.
  • Ask the user before deploying. After the user approves, create a new Dokploy application on vm.ifkash.dev, with a domain such as dreamcart.ifkash.dev that the user confirms. Don't touch existing apps. The user might need to add the DNS record.

Acceptance criteria:

  • The game plays locally in Chrome and Safari at 15 fps or more.
  • After deployment, the page loads at the confirmed URL.

M8: Write-up

Tasks:

  • Write README.md. It explains what the project is, how to reproduce each milestone with one command per milestone, the device build and flash steps, the sweep plot, and GIFs.
  • Write post/draft.md, a blog post of 1,200-1,800 words for ifkash.dev. Structure it as follows:
  • The hook: "There's no game on this chip."
  • Why world models are the topic of the year, and why you went small.
  • The Markov-in-two-frames design.
  • Drift, and how rollout training fixed it, with before and after GIFs.
  • The scaling chart: how many parameters it takes to be Breakout.
  • Squeezing the model into 264 KB: the int8 path, dual-core, and measured frame rates.
  • Play it yourself: the link to the browser version.
  • Limitations and next steps: the FPGA MAC unit, other games, and stochastic games.
  • Ask the user before you push anything. After the user approves, push the repo to weights-and-wires/dreamcart, and upload the checkpoints for every preset to the Hugging Face Hub under weights-and-wires/dreamcart.
  • Terminate every pod, and write the final spend to NOTES.md.

Final report to the user

When you finish, report the following:

  • The selected device preset, with its parameter count, MACs, and weight size.
  • The frame rate on the device and in the browser.
  • Rule validity and frames-to-divergence for float and int8.
  • The sweep plot.
  • The total spend.
  • Anything that you skipped or that failed.