# mindseye: a robot that plans in its imagination, live in your browser tab

---

# Part 1: For you

## The idea in one line

A robot pusher learns how the world works just by watching. Then, right in your browser, it
imagines hundreds of possible futures each second and picks the best one. You can mess with it
and watch it adapt. You can also break physics and watch it get *surprised*.

## Why it's exciting

- **It's Yann LeCun's big bet.** He left Meta in late 2025 to build world models that predict
  in an abstract space instead of predicting pixels. The architecture is called *JEPA*. In March
  2026, his group released
  [LeWorldModel](https://arxiv.org/abs/2603.19312), the first JEPA that trains stably from raw
  pixels. It's tiny, about 15M parameters, and trains in hours on one GPU.
- **Nobody has put one in a browser.** The research agents found no browser demo of a JEPA that
  plans. You'd be first.
- **It's physical AI.** It covers a robot, contact physics, and planning. This is the stuff
  robotics labs care about.
- **It's low-level.** You write your own WebGPU kernels so that the whole planning loop runs on
  the viewer's GPU. There's no server and no ONNX runtime, just your shaders.
- **It's brand new for you.** No LLaMA and no FPGA. It uses a new architecture, a new domain, and
  a new runtime.

## What the demo looks like

- **A 3D table with a robot pusher and a T-shaped block.** This is the classic "Push-T"
  robotics task.
- **Drag the ghost target** anywhere. The robot pushes the block there.
- **Grab the block with your mouse and throw it.** The robot replans on the fly.
- **A "what it's imagining" filmstrip** that shows the future the robot is picturing right now.
- **A surprise meter.** Click **Break physics**: teleport the block, turn the floor to ice, or
  let the block pass through a wall. The meter spikes, because the model knows that isn't how
  the world works.
- **A speed counter** that shows how many futures per second your GPU imagines.

## How it works

1. **Build the sim.** A MuJoCo table, pusher, and T-block. The *same* scene runs in Python for
   training and in the browser through the official MuJoCo WebAssembly package.
2. **Collect data.** The pusher moves around randomly and bumps the block. The GPU pod records
   about a million small 64×64 frames.
3. **Train the world model.** Train LeWorldModel on those frames. It learns to predict what
   happens next, but in its own compressed "thought space," not in pixels.
4. **Plan by imagining.** To reach a goal, the robot tries about 300 random action sequences in
   its head, keeps the best ones, and repeats. It then takes the first step of the winner. This
   method is called *CEM*.
5. **Write the kernels.** Rewrite the model and the whole planning loop in WGSL, the WebGPU
   shader language, so that everything stays on the GPU.
6. **Ship it** as a static web page on `vm.ifkash.dev`.

## Weekend plan

| When | What | Done when |
|---|---|---|
| Saturday morning | Run the official LeWorldModel checkpoint, build the MuJoCo scene | Their model plans; your scene runs in Python and in the browser |
| Saturday afternoon | Generate data, start training (runs by itself) | Training is running |
| Saturday evening | Check planning in Python, add the surprise meter | The robot pushes the block to the goal |
| Sunday morning | Write the WebGPU kernels, match them against PyTorch | Browser output matches PyTorch |
| Sunday afternoon | Build the demo page, deploy it, record the video, write the post | The link works |

## Cost

- **GPU:** 1× RTX 5090 on RunPod, about $0.69 per hour. The fallback is an RTX PRO 6000, about
  $1.69 per hour.
- **Time on the GPU:** about 12-18 hours.
- **Total:** about $10-15. The plan sets a hard stop at $40.
- **Hosting:** free. It's a static page on your VM, and it keeps working after the pod is gone.

## What you have at the end

- A live link where anyone can play with a robot that thinks in latent space.
- A speed chart: your WebGPU kernels against ONNX Runtime Web.
- A video of the surprise meter spiking when you break physics.
- A blog post: *"I put LeCun's world model in a browser tab."*

## What might go wrong

| Problem | What to do |
|---|---|
| Your trained model plans badly | Start from the official checkpoint's recipe, and compare with its numbers. |
| The browser can't plan fast enough | Use fewer samples or shorter plans. 5 plans per second still looks smooth. |
| The browser frames look different from the training frames | The plan draws the model's view with a tiny custom renderer that's identical in Python and the browser. |
| WebGPU isn't available in a visitor's browser | Show a recorded video instead. |

## Other ideas the research turned up

- **WarpDrive:** Make an open video world model, Matrix-Game 2.0, run at 60 FPS on one RTX 5090
  with FP4 kernels, then stream it to the browser. This is the most hardcore low-level option,
  but the demo only works while you pay for the GPU.
- **Neural MuJoCo:** Replace a robot's physics engine with a neural network, and switch between
  them live. It's cool, but the port is hard.
- **A Jev-style decision model** playing a browser game. Jev is closed, and open clones such as
  LAYA already exist.

## Reading, if you want it

- [LeWorldModel paper](https://arxiv.org/abs/2603.19312) and
  [code](https://github.com/lucas-maes/le-wm), under the MIT license.
- [Official MuJoCo WebAssembly bindings](https://github.com/google-deepmind/mujoco/tree/main/wasm).

---

# Part 2: For the coding agent

## Mission

Build `mindseye`, a browser demo of an action-conditioned JEPA world model that plans in latent
space and runs entirely on the client:

- **Model:** LeWorldModel (LeWM), trained from pixels on a MuJoCo Push-T task.
- **Planning:** cross-entropy method (CEM) model-predictive control, with the model and the
  full planning loop written as hand-written WGSL compute kernels.
- **Physics:** the official MuJoCo WebAssembly bindings.
- **Surprise meter:** a violation-of-expectation readout based on prediction error.

The final artifact is a static website with no inference server.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a
milestone until the previous one passes. After each milestone, commit your work and write a
short entry in `NOTES.md` with the results and numbers.

## Hard constraints

- **Secrets:** Read `RUNPOD_API_KEY`, `HF_TOKEN`, and `WANDB_API_KEY` from environment
  variables only. Never write a secret into any file, log, commit, or echoed command. Commit a
  `.env.example` that has placeholder values only, and add `.env` to `.gitignore`.
- **Budget:** The hard cap is $40 of RunPod spend. Every pod runs a watchdog that stops the pod
  after `MAX_POD_HOURS` hours. The default is `6`.
- **Compute:** Run training and large-scale data generation only on RunPod. The local Mac runs
  the web build, the browser tests, and the small scripts.
- **Outward actions:** Ask the user before you do any of the following: create the GitHub
  repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything
  publicly. Use the `kashifulhaque` GitHub account (`gh auth switch -u kashifulhaque`).
- **Shared VM:** `vm.ifkash.dev` runs other production apps under Dokploy, with Traefik on
  ports 80 and 443. Don't stop, restart, reconfigure, or remove any existing container,
  service, or proxy rule. Add mindseye only as a new Dokploy application.
- **Pod cleanup:** When a pod isn't running a job, stop it. At the end of the project,
  terminate every pod that you created and report the total spend.
- **Licenses:** LeWM is under the MIT license. Keep its license notice in any file that you copy
  or adapt, and credit the paper in `README.md`.

## Upstream facts

These facts come from the LeWM README and the paper. Confirm the details in the repo's configs
before you rely on them:

- **Repo:** `github.com/lucas-maes/le-wm`, under the MIT license. It depends on the
  `stable-worldmodel` and `stable-pretraining` packages:
  `uv pip install stable-worldmodel[train,env]`.
- **Architecture:** A ViT encoder produces one embedding per frame. An autoregressive
  predictor, `ARPredictor`, is conditioned on an action encoder, and projection MLPs map
  between spaces. The loss is next-embedding prediction plus a Gaussian regularizer on the
  embeddings. The model has about 15M parameters.
- **Commands:** Train with `python train.py data=pusht`. Evaluate planning with
  `python eval.py --config-name=pusht.yaml policy=pusht/lewm`. Checkpoint paths are relative
  to `$STABLEWM_HOME`.
- **Checkpoints on the Hugging Face Hub:** `quentinll/lewm-pusht`, `lewm-cube`,
  `lewm-tworooms`, and `lewm-reacher`.
- **Record the unknowns.** Read the image resolution, patch size, embedding dimension, history
  length, and planner hyperparameters from `config/`, and record them in `NOTES.md` during M1.

## Tech stack

Pin every version. The stack is as follows:

- **Python:** Python 3.11, PyTorch 2.x, `mujoco` (the Python bindings), `h5py`, `numpy`,
  `wandb`, and LeWM's dependencies, managed with `uv`.
- **Browser:**
  - The official MuJoCo WASM package, `@mujoco/mujoco` from
    `google-deepmind/mujoco/wasm`. Use the same MuJoCo version as the Python package, and
    record both versions.
  - `three.js` for the human-facing 3D view.
  - Plain TypeScript, bundled with `vite`, and raw WebGPU with WGSL.
  - Don't use an ML runtime in the final demo. `onnxruntime-web` is allowed only as the
    benchmark baseline.
- **Browser tests:** Playwright with Chromium, with WebGPU enabled.
- **Pods:** the `runpod` Python SDK, with `runpodctl` inside pods.

## Repository layout

Create the following layout:

```
mindseye/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh
  sim/
    pusht.xml            # single source of truth for the scene (Python + browser)
    obs_render.py        # tiny deterministic 64x64 observation rasterizer
    collect.py           # parallel data generation -> HDF5 in LeWM's format
  train/                 # thin wrappers/configs around LeWM
    configs/mj_pusht.yaml
    probe_decoder.py     # latent -> 64x64 frame, for visualization only
    plan_eval.py         # CEM MPC success-rate eval in Python
    export_weights.py    # safetensors -> flat f16/f32 binary + JSON manifest
    golden.py            # golden tensors for kernel tests
  web/
    src/sim.ts           # @mujoco/mujoco wrapper, stepping, perturbations
    src/obs_render.ts    # exact port of obs_render.py
    src/gpu/             # WGSL kernels + TS dispatch code
      matmul.wgsl  layernorm.wgsl  attention.wgsl  gelu.wgsl
      vit_encoder.ts  predictor.ts  cem.ts  surprise.ts
    src/view3d.ts        # three.js scene, ghost goal, imagination filmstrip
    src/main.ts  index.html
    tests/               # Playwright: kernel golden tests, planner smoke test
    bench/               # WGSL vs onnxruntime-web benchmark page
  deploy/Dockerfile      # nginx:alpine serving web/dist
  results/  post/draft.md
```

## M0: Infrastructure

**Tasks:**

1. Write `infra/pod.py`. It uses the `runpod` SDK and reads `RUNPOD_API_KEY` from the
   environment. It supports the following subcommands:
   - `create`: Creates a pod. The GPU preference order is RTX 5090, then RTX PRO 6000, then
     RTX 4090. Resolve the GPU type IDs at run time by querying the available GPU types. Use an
     official RunPod PyTorch image with CUDA 12.8 or later. Set a 100 GB volume at `/workspace`
     and expose SSH. Prefer a pod with 16 or more vCPUs for data generation.
   - `status`, `stop`, and `terminate`.
   - `ssh-info`: Prints the SSH command.
   - `cost`: Prints the uptime and spend for every pod whose name has the prefix `mindseye-`.
2. Write `infra/watchdog.sh`. It sleeps for `MAX_POD_HOURS`, then runs
   `runpodctl stop pod $RUNPOD_POD_ID`.
3. Write `infra/bootstrap.sh`. It clones the repo and LeWM, runs `uv sync`, sets
   `STABLEWM_HOME=/workspace/swm`, sets up MuJoCo offscreen support (EGL) if you need it, and
   starts the watchdog.

**Acceptance criteria:**

- A pod comes up, and `bootstrap.sh` finishes without errors.
- A test run with `MAX_POD_HOURS=0.05` stops the pod within 5 minutes.

## M1: Reproduce upstream

**Tasks:**

- Download `quentinll/lewm-pusht` and the Push-T dataset, and run the upstream planning
  evaluation. Record the success rate and the planning time per step on the pod GPU.
- Record in `NOTES.md` every architecture and planner hyperparameter listed in the upstream
  facts.

**Acceptance criteria:**

- The upstream evaluation runs, and the success rate is within the range the paper reports.
  If it isn't, stop and report the gap before you continue.

## M2: MuJoCo Push-T scene and a renderer that matches everywhere

**The scene (`sim/pusht.xml`):**

- A flat table, with walls that bound a 0.5 m × 0.5 m workspace.
- A T-shaped block built from two box geoms, with a free joint whose motion stays in the plane.
  Restrict it with joint constraints, or keep it flat with friction and a low center of mass.
  Record the choice you make.
- A cylindrical pusher driven by 2D position actuators, or by `mocap` with a weld to a slider.
  The action is `(dx, dy)` per control step, clipped to ±2 cm.
- Pick a control frequency of about 10 Hz and a timestep that divides it evenly.
- Define the goal as a target pose for the block: `(x, y, θ)`.

**Observation renderer:**

The model's input comes from a deterministic top-down orthographic rasterizer, not from the
MuJoCo renderer and not from three.js. This makes the pixels identical in training and in the
browser.

- The rasterizer maps the poses of the pusher and the block to a 64×64 RGB image with flat
  colors and no lighting. It uses point-in-polygon tests for the T-block and a circle test for
  the pusher.
- For anti-aliasing, use 4×4 supersampling with a fixed sample pattern.
- `obs_render.py` and `obs_render.ts` must implement the same arithmetic. Use float32 in numpy
  and `Math.fround` in TypeScript where it matters.
- If M1 shows that LeWM expects a different resolution, match LeWM, and record the change.

**Acceptance criteria:**

- The golden-image test passes: over 1,000 random poses, the Python and TypeScript renderers
  produce the same bytes, or at most 0.1% of pixels differ by 1 LSB or less.
- Physics parity: from the same initial state, run 200 steps with the same actions in Python
  `mujoco` and in the browser. The block pose must stay within 1 mm and 0.5° at every step.
  If it drifts more, first align the MuJoCo versions and the integrator settings. Record the
  result, because it matters less than renderer parity.

## M3: Data generation

**Tasks:**

- Write `sim/collect.py` to generate episodes in parallel on the pod's CPUs. Use one process
  per core.
- Mix the following data sources:
  - 60% random-walk pusher motion with momentum, biased toward the block so that contacts
    happen often.
  - 30% scripted pushes that move toward a random point on the block's edge and push through
    it.
  - 10% pure noise.
- Randomize the starting poses of the block and the pusher.
- Scale: about 20,000 episodes of 100 steps each, which is 2 million frames. Write them to HDF5
  in the exact schema the LeWM dataloader expects. Copy the Push-T dataset's keys, dtypes, and
  shapes from M1.
- Save a held-out set of 500 episodes.

**Acceptance criteria:**

- The upstream LeWM dataloader reads the new dataset without any code changes.
- In at least 50% of steps, the pusher touches the block.
- `results/m3_samples.png` shows a grid of frames.

## M4: Training and probes

**Tasks:**

- Write `train/configs/mj_pusht.yaml`, based on upstream `pusht.yaml` and pointed at the M3
  data.
- Train with the upstream hyperparameters first. Change a hyperparameter only if the training
  collapses or the loss plateaus early, and record every change.
- Write `probe_decoder.py`: a small convolutional decoder that maps a frozen embedding to a
  64×64 frame. Train it with L2 plus a light perceptual or SSIM term. It's for visualization
  only.
- Surprise calibration: compute `s_t = ||pred(z_{<=t}, a_t) - enc(o_{t+1})||²` on held-out
  normal data. Store its mean and standard deviation for z-scoring, and save them to the export
  manifest.

**Acceptance criteria:**

- The training loss curves are in wandb, and there's no representation collapse: the effective
  rank of the embeddings is above 50% of the embedding dimension.
- The probe reconstructions at 1, 5, and 10 steps of imagined rollout are recognizable. Save
  them as `results/m4_probe.png`.
- Surprise separation: on 200 perturbed held-out episodes, the z-scored surprise at the
  perturbation step is above 3 in at least 90% of cases, and above 3 on at most 2% of normal
  steps. The perturbations are teleporting the block, removing friction, and a pass-through.

## M5: Planning in Python

**Tasks:**

- Write `train/plan_eval.py`: CEM MPC in latent space. Start from the planner settings that you
  recorded in M1. If you have to set them yourself, use 300 samples, a horizon of 5,
  3 iterations, 30 elites, and a Gaussian over action sequences. The cost is the distance
  between the predicted final embedding and the goal embedding, where the goal embedding is the
  encoding of the rendered goal frame.
- Run the policy as receding-horizon control: execute the first action, then replan.
- **Evaluation:** Run 200 episodes with random start and goal poses and a limit of 200 steps.
  An episode succeeds when the block is within 2 cm and 15° of the goal.

**Acceptance criteria:**

- The success rate is 60% or more, or within 10 points of the upstream Push-T number from M1,
  whichever is lower.
- `results/m5_planning.json` and 10 GIFs exist.

## M6: WGSL kernels

**Export:**

- Write `export_weights.py` to produce `weights.bin` with fp16 weights, plus `manifest.json`
  with the tensor names, shapes, offsets, the normalization constants, and the surprise
  statistics.
- Write `golden.py` to dump the inputs, the intermediate activations after every block, and the
  outputs for 20 cases, for the encoder, for one predictor step, and for one full CEM iteration
  with fixed noise.

**Kernels:**

- Write the following kernels. Use fp16 storage when the device supports `shader-f16`;
  otherwise use fp32. Accumulate in fp32 in both cases.
  - A tiled matmul that uses workgroup memory.
  - A fused LayerNorm.
  - A fused GELU with bias add.
  - Attention for short sequences: fuse QKᵀ, softmax, and V in one workgroup for each head and
    each sample.
  - The patch embedding.
- Batch the predictor across the CEM samples: the batch is `N_samples`, and the sequence is
  the history plus the horizon.
- Keep the whole CEM loop on the GPU:
  - Sample with a GPU Philox or PCG random number generator.
  - Roll out the predictor, compute the costs, and select the top-k with a bitonic sort or a
    radix select.
  - Update the mean and standard deviation.
  - Repeat for every iteration without a CPU readback.
  - Read back only the final first action, 8 bytes, and the elite trajectory for the
    filmstrip.
- Use one command encoder for each replan.
- Measure the GPU time of each pass with timestamp queries, when the browser supports them.

**Acceptance criteria:**

- The Playwright tests pass in headless Chromium with WebGPU. Against the goldens, every
  intermediate tensor has a maximum relative error of 2e-2 or less in fp16 and 1e-4 or less in
  fp32.
- With fixed noise, the CEM run selects the same first action as PyTorch, within 1e-3.
- **Benchmark:** On the benchmark page, compare replans per second between these kernels and
  `onnxruntime-web` with the WebGPU execution provider, running the same model with the CEM
  loop in JavaScript. Report the numbers on an Apple M-series Mac and, if possible, on one other
  GPU. Target: 10 replans per second or more at the M5 settings on an M-series Mac. Save the
  results to `results/m6_bench.json`.

## M7: The demo page

**Tasks:**

- **Layout:** A three.js 3D view of the table, rendered from MuJoCo state every frame. Next to
  it, show the following:
  - The model's 64×64 input.
  - The "imagination" filmstrip: the probe decoder applied to the elite plan's predicted
    embeddings, for the horizon steps. Run the probe decoder in WGSL too. It's a few
    convolutions or transposed convolutions.
  - The surprise meter as a sparkline and a gauge.
  - A counter for replans per second and futures imagined per second, where futures equals
    samples times iterations times replans.
- **Interactions:**
  - Drag the translucent ghost T-block to set the goal pose. The scroll wheel rotates it.
  - Grab the real block and throw it, using a spring force applied through
    `xfrc_applied`.
  - **Break physics** buttons for teleporting the block, setting the friction to 0, and
    turning off wall collisions. Each one shows the surprise spike.
  - A pause button, and a slider for the number of samples.
- Control loop: run physics at the sim rate, and replan at the rate the GPU sustains. Between
  replans, the pusher executes the rest of the current plan.
- Fallback: if `navigator.gpu` is missing, show a looping MP4 of the demo and a short
  explanation.
- Mobile layout: stack the panels vertically, and support touch drag for the goal.

**Acceptance criteria:**

- A Playwright smoke test loads the page, sets a goal, and checks that the block moves toward
  it within 30 seconds of simulated time.
- The page works in Chrome and Safari on macOS.

## M8: Deploy and write-up

**Tasks:**

- Write `deploy/Dockerfile`: `nginx:alpine` serving `web/dist`, with correct MIME types for
  `.wasm` and `.bin` and the cross-origin isolation headers if the MuJoCo multithreaded build
  needs them.
- Ask the user before deploying. After the user approves, create a new Dokploy application on
  `vm.ifkash.dev`, with a domain such as `mindseye.ifkash.dev` that the user confirms. The user
  might need to add the DNS record.
- Write `README.md`. It explains what the project is, how to reproduce each milestone with one
  command per milestone, the benchmark table, GIFs, and credits and license notes for LeWM and
  MuJoCo.
- Write `post/draft.md`, a blog post of 1,200-1,800 words. Structure it as follows:
  - The hook.
  - What a JEPA is, in plain words.
  - Training on MuJoCo Push-T.
  - Planning by imagining: CEM.
  - Moving the whole planner onto WebGPU, with the kernel diagram and the benchmark.
  - The surprise meter.
  - Limitations: 2D task, a single scene, and a probe decoder that's for display only.
- Ask the user before you push anything. After the user approves, push the repo to
  `weights-and-wires/mindseye` and upload the checkpoint to the Hugging Face Hub under
  `weights-and-wires/mindseye-lewm-mjpusht`.
- Terminate every pod, and write the final spend to `NOTES.md`.

**Acceptance criteria:**

- The deployed URL loads and plans in Chrome.
- No pods are left running, and the total spend is less than $40.

## Final report to the user

When you finish, report the following:

- The upstream reproduction number from M1.
- The success rates for your model, from M5.
- The surprise separation, from M4.
- Replans per second for WGSL against ONNX Runtime Web, from M6.
- The URL, if the site is deployed.
- The total spend.
- Anything that you skipped or that failed.
