← projects

mindseye

A robot that plans in its imagination, live in your browser tab

16 min read mindseye.md

Part 1: For you

The idea in one line

A robot pusher learns how the world works just by watching. Then, right in your browser, it imagines hundreds of possible futures each second and picks the best one. You can mess with it and watch it adapt. You can also break physics and watch it get surprised.

Why it's exciting

  • It's Yann LeCun's big bet. He left Meta in late 2025 to build world models that predict in an abstract space instead of predicting pixels. The architecture is called JEPA. In March 2026, his group released LeWorldModel, the first JEPA that trains stably from raw pixels. It's tiny, about 15M parameters, and trains in hours on one GPU.
  • Nobody has put one in a browser. The research agents found no browser demo of a JEPA that plans. You'd be first.
  • It's physical AI. It covers a robot, contact physics, and planning. This is the stuff robotics labs care about.
  • It's low-level. You write your own WebGPU kernels so that the whole planning loop runs on the viewer's GPU. There's no server and no ONNX runtime, just your shaders.
  • It's brand new for you. No LLaMA and no FPGA. It uses a new architecture, a new domain, and a new runtime.

What the demo looks like

  • A 3D table with a robot pusher and a T-shaped block. This is the classic "Push-T" robotics task.
  • Drag the ghost target anywhere. The robot pushes the block there.
  • Grab the block with your mouse and throw it. The robot replans on the fly.
  • A "what it's imagining" filmstrip that shows the future the robot is picturing right now.
  • A surprise meter. Click Break physics: teleport the block, turn the floor to ice, or let the block pass through a wall. The meter spikes, because the model knows that isn't how the world works.
  • A speed counter that shows how many futures per second your GPU imagines.

How it works

  1. Build the sim. A MuJoCo table, pusher, and T-block. The same scene runs in Python for training and in the browser through the official MuJoCo WebAssembly package.
  2. Collect data. The pusher moves around randomly and bumps the block. The GPU pod records about a million small 64×64 frames.
  3. Train the world model. Train LeWorldModel on those frames. It learns to predict what happens next, but in its own compressed "thought space," not in pixels.
  4. Plan by imagining. To reach a goal, the robot tries about 300 random action sequences in its head, keeps the best ones, and repeats. It then takes the first step of the winner. This method is called CEM.
  5. Write the kernels. Rewrite the model and the whole planning loop in WGSL, the WebGPU shader language, so that everything stays on the GPU.
  6. Ship it as a static web page on vm.ifkash.dev.

Weekend plan

When What Done when
Saturday morning Run the official LeWorldModel checkpoint, build the MuJoCo scene Their model plans; your scene runs in Python and in the browser
Saturday afternoon Generate data, start training (runs by itself) Training is running
Saturday evening Check planning in Python, add the surprise meter The robot pushes the block to the goal
Sunday morning Write the WebGPU kernels, match them against PyTorch Browser output matches PyTorch
Sunday afternoon Build the demo page, deploy it, record the video, write the post The link works

Cost

  • GPU: 1× RTX 5090 on RunPod, about $0.69 per hour. The fallback is an RTX PRO 6000, about $1.69 per hour.
  • Time on the GPU: about 12-18 hours.
  • Total: about $10-15. The plan sets a hard stop at $40.
  • Hosting: free. It's a static page on your VM, and it keeps working after the pod is gone.

What you have at the end

  • A live link where anyone can play with a robot that thinks in latent space.
  • A speed chart: your WebGPU kernels against ONNX Runtime Web.
  • A video of the surprise meter spiking when you break physics.
  • A blog post: "I put LeCun's world model in a browser tab."

What might go wrong

Problem What to do
Your trained model plans badly Start from the official checkpoint's recipe, and compare with its numbers.
The browser can't plan fast enough Use fewer samples or shorter plans. 5 plans per second still looks smooth.
The browser frames look different from the training frames The plan draws the model's view with a tiny custom renderer that's identical in Python and the browser.
WebGPU isn't available in a visitor's browser Show a recorded video instead.

Other ideas the research turned up

  • WarpDrive: Make an open video world model, Matrix-Game 2.0, run at 60 FPS on one RTX 5090 with FP4 kernels, then stream it to the browser. This is the most hardcore low-level option, but the demo only works while you pay for the GPU.
  • Neural MuJoCo: Replace a robot's physics engine with a neural network, and switch between them live. It's cool, but the port is hard.
  • A Jev-style decision model playing a browser game. Jev is closed, and open clones such as LAYA already exist.

Reading, if you want it


Part 2: For the coding agent

Mission

Build mindseye, a browser demo of an action-conditioned JEPA world model that plans in latent space and runs entirely on the client:

  • Model: LeWorldModel (LeWM), trained from pixels on a MuJoCo Push-T task.
  • Planning: cross-entropy method (CEM) model-predictive control, with the model and the full planning loop written as hand-written WGSL compute kernels.
  • Physics: the official MuJoCo WebAssembly bindings.
  • Surprise meter: a violation-of-expectation readout based on prediction error.

The final artifact is a static website with no inference server.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a milestone until the previous one passes. After each milestone, commit your work and write a short entry in NOTES.md with the results and numbers.

Hard constraints

  • Secrets: Read RUNPOD_API_KEY, HF_TOKEN, and WANDB_API_KEY from environment variables only. Never write a secret into any file, log, commit, or echoed command. Commit a .env.example that has placeholder values only, and add .env to .gitignore.
  • Budget: The hard cap is $40 of RunPod spend. Every pod runs a watchdog that stops the pod after MAX_POD_HOURS hours. The default is 6.
  • Compute: Run training and large-scale data generation only on RunPod. The local Mac runs the web build, the browser tests, and the small scripts.
  • Outward actions: Ask the user before you do any of the following: create the GitHub repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything publicly. Use the kashifulhaque GitHub account (gh auth switch -u kashifulhaque).
  • Shared VM: vm.ifkash.dev runs other production apps under Dokploy, with Traefik on ports 80 and 443. Don't stop, restart, reconfigure, or remove any existing container, service, or proxy rule. Add mindseye only as a new Dokploy application.
  • Pod cleanup: When a pod isn't running a job, stop it. At the end of the project, terminate every pod that you created and report the total spend.
  • Licenses: LeWM is under the MIT license. Keep its license notice in any file that you copy or adapt, and credit the paper in README.md.

Upstream facts

These facts come from the LeWM README and the paper. Confirm the details in the repo's configs before you rely on them:

  • Repo: github.com/lucas-maes/le-wm, under the MIT license. It depends on the stable-worldmodel and stable-pretraining packages: uv pip install stable-worldmodel[train,env].
  • Architecture: A ViT encoder produces one embedding per frame. An autoregressive predictor, ARPredictor, is conditioned on an action encoder, and projection MLPs map between spaces. The loss is next-embedding prediction plus a Gaussian regularizer on the embeddings. The model has about 15M parameters.
  • Commands: Train with python train.py data=pusht. Evaluate planning with python eval.py --config-name=pusht.yaml policy=pusht/lewm. Checkpoint paths are relative to $STABLEWM_HOME.
  • Checkpoints on the Hugging Face Hub: quentinll/lewm-pusht, lewm-cube, lewm-tworooms, and lewm-reacher.
  • Record the unknowns. Read the image resolution, patch size, embedding dimension, history length, and planner hyperparameters from config/, and record them in NOTES.md during M1.

Tech stack

Pin every version. The stack is as follows:

  • Python: Python 3.11, PyTorch 2.x, mujoco (the Python bindings), h5py, numpy, wandb, and LeWM's dependencies, managed with uv.
  • Browser:
  • The official MuJoCo WASM package, @mujoco/mujoco from google-deepmind/mujoco/wasm. Use the same MuJoCo version as the Python package, and record both versions.
  • three.js for the human-facing 3D view.
  • Plain TypeScript, bundled with vite, and raw WebGPU with WGSL.
  • Don't use an ML runtime in the final demo. onnxruntime-web is allowed only as the benchmark baseline.
  • Browser tests: Playwright with Chromium, with WebGPU enabled.
  • Pods: the runpod Python SDK, with runpodctl inside pods.

Repository layout

Create the following layout:

mindseye/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh
  sim/
    pusht.xml            # single source of truth for the scene (Python + browser)
    obs_render.py        # tiny deterministic 64x64 observation rasterizer
    collect.py           # parallel data generation -> HDF5 in LeWM's format
  train/                 # thin wrappers/configs around LeWM
    configs/mj_pusht.yaml
    probe_decoder.py     # latent -> 64x64 frame, for visualization only
    plan_eval.py         # CEM MPC success-rate eval in Python
    export_weights.py    # safetensors -> flat f16/f32 binary + JSON manifest
    golden.py            # golden tensors for kernel tests
  web/
    src/sim.ts           # @mujoco/mujoco wrapper, stepping, perturbations
    src/obs_render.ts    # exact port of obs_render.py
    src/gpu/             # WGSL kernels + TS dispatch code
      matmul.wgsl  layernorm.wgsl  attention.wgsl  gelu.wgsl
      vit_encoder.ts  predictor.ts  cem.ts  surprise.ts
    src/view3d.ts        # three.js scene, ghost goal, imagination filmstrip
    src/main.ts  index.html
    tests/               # Playwright: kernel golden tests, planner smoke test
    bench/               # WGSL vs onnxruntime-web benchmark page
  deploy/Dockerfile      # nginx:alpine serving web/dist
  results/  post/draft.md

M0: Infrastructure

Tasks:

  1. Write infra/pod.py. It uses the runpod SDK and reads RUNPOD_API_KEY from the environment. It supports the following subcommands: - create: Creates a pod. The GPU preference order is RTX 5090, then RTX PRO 6000, then RTX 4090. Resolve the GPU type IDs at run time by querying the available GPU types. Use an official RunPod PyTorch image with CUDA 12.8 or later. Set a 100 GB volume at /workspace and expose SSH. Prefer a pod with 16 or more vCPUs for data generation. - status, stop, and terminate. - ssh-info: Prints the SSH command. - cost: Prints the uptime and spend for every pod whose name has the prefix mindseye-.
  2. Write infra/watchdog.sh. It sleeps for MAX_POD_HOURS, then runs runpodctl stop pod $RUNPOD_POD_ID.
  3. Write infra/bootstrap.sh. It clones the repo and LeWM, runs uv sync, sets STABLEWM_HOME=/workspace/swm, sets up MuJoCo offscreen support (EGL) if you need it, and starts the watchdog.

Acceptance criteria:

  • A pod comes up, and bootstrap.sh finishes without errors.
  • A test run with MAX_POD_HOURS=0.05 stops the pod within 5 minutes.

M1: Reproduce upstream

Tasks:

  • Download quentinll/lewm-pusht and the Push-T dataset, and run the upstream planning evaluation. Record the success rate and the planning time per step on the pod GPU.
  • Record in NOTES.md every architecture and planner hyperparameter listed in the upstream facts.

Acceptance criteria:

  • The upstream evaluation runs, and the success rate is within the range the paper reports. If it isn't, stop and report the gap before you continue.

M2: MuJoCo Push-T scene and a renderer that matches everywhere

The scene (sim/pusht.xml):

  • A flat table, with walls that bound a 0.5 m × 0.5 m workspace.
  • A T-shaped block built from two box geoms, with a free joint whose motion stays in the plane. Restrict it with joint constraints, or keep it flat with friction and a low center of mass. Record the choice you make.
  • A cylindrical pusher driven by 2D position actuators, or by mocap with a weld to a slider. The action is (dx, dy) per control step, clipped to ±2 cm.
  • Pick a control frequency of about 10 Hz and a timestep that divides it evenly.
  • Define the goal as a target pose for the block: (x, y, θ).

Observation renderer:

The model's input comes from a deterministic top-down orthographic rasterizer, not from the MuJoCo renderer and not from three.js. This makes the pixels identical in training and in the browser.

  • The rasterizer maps the poses of the pusher and the block to a 64×64 RGB image with flat colors and no lighting. It uses point-in-polygon tests for the T-block and a circle test for the pusher.
  • For anti-aliasing, use 4×4 supersampling with a fixed sample pattern.
  • obs_render.py and obs_render.ts must implement the same arithmetic. Use float32 in numpy and Math.fround in TypeScript where it matters.
  • If M1 shows that LeWM expects a different resolution, match LeWM, and record the change.

Acceptance criteria:

  • The golden-image test passes: over 1,000 random poses, the Python and TypeScript renderers produce the same bytes, or at most 0.1% of pixels differ by 1 LSB or less.
  • Physics parity: from the same initial state, run 200 steps with the same actions in Python mujoco and in the browser. The block pose must stay within 1 mm and 0.5° at every step. If it drifts more, first align the MuJoCo versions and the integrator settings. Record the result, because it matters less than renderer parity.

M3: Data generation

Tasks:

  • Write sim/collect.py to generate episodes in parallel on the pod's CPUs. Use one process per core.
  • Mix the following data sources:
  • 60% random-walk pusher motion with momentum, biased toward the block so that contacts happen often.
  • 30% scripted pushes that move toward a random point on the block's edge and push through it.
  • 10% pure noise.
  • Randomize the starting poses of the block and the pusher.
  • Scale: about 20,000 episodes of 100 steps each, which is 2 million frames. Write them to HDF5 in the exact schema the LeWM dataloader expects. Copy the Push-T dataset's keys, dtypes, and shapes from M1.
  • Save a held-out set of 500 episodes.

Acceptance criteria:

  • The upstream LeWM dataloader reads the new dataset without any code changes.
  • In at least 50% of steps, the pusher touches the block.
  • results/m3_samples.png shows a grid of frames.

M4: Training and probes

Tasks:

  • Write train/configs/mj_pusht.yaml, based on upstream pusht.yaml and pointed at the M3 data.
  • Train with the upstream hyperparameters first. Change a hyperparameter only if the training collapses or the loss plateaus early, and record every change.
  • Write probe_decoder.py: a small convolutional decoder that maps a frozen embedding to a 64×64 frame. Train it with L2 plus a light perceptual or SSIM term. It's for visualization only.
  • Surprise calibration: compute s_t = ||pred(z_{<=t}, a_t) - enc(o_{t+1})||² on held-out normal data. Store its mean and standard deviation for z-scoring, and save them to the export manifest.

Acceptance criteria:

  • The training loss curves are in wandb, and there's no representation collapse: the effective rank of the embeddings is above 50% of the embedding dimension.
  • The probe reconstructions at 1, 5, and 10 steps of imagined rollout are recognizable. Save them as results/m4_probe.png.
  • Surprise separation: on 200 perturbed held-out episodes, the z-scored surprise at the perturbation step is above 3 in at least 90% of cases, and above 3 on at most 2% of normal steps. The perturbations are teleporting the block, removing friction, and a pass-through.

M5: Planning in Python

Tasks:

  • Write train/plan_eval.py: CEM MPC in latent space. Start from the planner settings that you recorded in M1. If you have to set them yourself, use 300 samples, a horizon of 5, 3 iterations, 30 elites, and a Gaussian over action sequences. The cost is the distance between the predicted final embedding and the goal embedding, where the goal embedding is the encoding of the rendered goal frame.
  • Run the policy as receding-horizon control: execute the first action, then replan.
  • Evaluation: Run 200 episodes with random start and goal poses and a limit of 200 steps. An episode succeeds when the block is within 2 cm and 15° of the goal.

Acceptance criteria:

  • The success rate is 60% or more, or within 10 points of the upstream Push-T number from M1, whichever is lower.
  • results/m5_planning.json and 10 GIFs exist.

M6: WGSL kernels

Export:

  • Write export_weights.py to produce weights.bin with fp16 weights, plus manifest.json with the tensor names, shapes, offsets, the normalization constants, and the surprise statistics.
  • Write golden.py to dump the inputs, the intermediate activations after every block, and the outputs for 20 cases, for the encoder, for one predictor step, and for one full CEM iteration with fixed noise.

Kernels:

  • Write the following kernels. Use fp16 storage when the device supports shader-f16; otherwise use fp32. Accumulate in fp32 in both cases.
  • A tiled matmul that uses workgroup memory.
  • A fused LayerNorm.
  • A fused GELU with bias add.
  • Attention for short sequences: fuse QKᵀ, softmax, and V in one workgroup for each head and each sample.
  • The patch embedding.
  • Batch the predictor across the CEM samples: the batch is N_samples, and the sequence is the history plus the horizon.
  • Keep the whole CEM loop on the GPU:
  • Sample with a GPU Philox or PCG random number generator.
  • Roll out the predictor, compute the costs, and select the top-k with a bitonic sort or a radix select.
  • Update the mean and standard deviation.
  • Repeat for every iteration without a CPU readback.
  • Read back only the final first action, 8 bytes, and the elite trajectory for the filmstrip.
  • Use one command encoder for each replan.
  • Measure the GPU time of each pass with timestamp queries, when the browser supports them.

Acceptance criteria:

  • The Playwright tests pass in headless Chromium with WebGPU. Against the goldens, every intermediate tensor has a maximum relative error of 2e-2 or less in fp16 and 1e-4 or less in fp32.
  • With fixed noise, the CEM run selects the same first action as PyTorch, within 1e-3.
  • Benchmark: On the benchmark page, compare replans per second between these kernels and onnxruntime-web with the WebGPU execution provider, running the same model with the CEM loop in JavaScript. Report the numbers on an Apple M-series Mac and, if possible, on one other GPU. Target: 10 replans per second or more at the M5 settings on an M-series Mac. Save the results to results/m6_bench.json.

M7: The demo page

Tasks:

  • Layout: A three.js 3D view of the table, rendered from MuJoCo state every frame. Next to it, show the following:
  • The model's 64×64 input.
  • The "imagination" filmstrip: the probe decoder applied to the elite plan's predicted embeddings, for the horizon steps. Run the probe decoder in WGSL too. It's a few convolutions or transposed convolutions.
  • The surprise meter as a sparkline and a gauge.
  • A counter for replans per second and futures imagined per second, where futures equals samples times iterations times replans.
  • Interactions:
  • Drag the translucent ghost T-block to set the goal pose. The scroll wheel rotates it.
  • Grab the real block and throw it, using a spring force applied through xfrc_applied.
  • Break physics buttons for teleporting the block, setting the friction to 0, and turning off wall collisions. Each one shows the surprise spike.
  • A pause button, and a slider for the number of samples.
  • Control loop: run physics at the sim rate, and replan at the rate the GPU sustains. Between replans, the pusher executes the rest of the current plan.
  • Fallback: if navigator.gpu is missing, show a looping MP4 of the demo and a short explanation.
  • Mobile layout: stack the panels vertically, and support touch drag for the goal.

Acceptance criteria:

  • A Playwright smoke test loads the page, sets a goal, and checks that the block moves toward it within 30 seconds of simulated time.
  • The page works in Chrome and Safari on macOS.

M8: Deploy and write-up

Tasks:

  • Write deploy/Dockerfile: nginx:alpine serving web/dist, with correct MIME types for .wasm and .bin and the cross-origin isolation headers if the MuJoCo multithreaded build needs them.
  • Ask the user before deploying. After the user approves, create a new Dokploy application on vm.ifkash.dev, with a domain such as mindseye.ifkash.dev that the user confirms. The user might need to add the DNS record.
  • Write README.md. It explains what the project is, how to reproduce each milestone with one command per milestone, the benchmark table, GIFs, and credits and license notes for LeWM and MuJoCo.
  • Write post/draft.md, a blog post of 1,200-1,800 words. Structure it as follows:
  • The hook.
  • What a JEPA is, in plain words.
  • Training on MuJoCo Push-T.
  • Planning by imagining: CEM.
  • Moving the whole planner onto WebGPU, with the kernel diagram and the benchmark.
  • The surprise meter.
  • Limitations: 2D task, a single scene, and a probe decoder that's for display only.
  • Ask the user before you push anything. After the user approves, push the repo to weights-and-wires/mindseye and upload the checkpoint to the Hugging Face Hub under weights-and-wires/mindseye-lewm-mjpusht.
  • Terminate every pod, and write the final spend to NOTES.md.

Acceptance criteria:

  • The deployed URL loads and plans in Chrome.
  • No pods are left running, and the total spend is less than $40.

Final report to the user

When you finish, report the following:

  • The upstream reproduction number from M1.
  • The success rates for your model, from M5.
  • The surprise separation, from M4.
  • Replans per second for WGSL against ONNX Runtime Web, from M6.
  • The URL, if the site is deployed.
  • The total spend.
  • Anything that you skipped or that failed.