mindseye
A robot that plans in its imagination, live in your browser tab
Part 1: For you
The idea in one line
A robot pusher learns how the world works just by watching. Then, right in your browser, it imagines hundreds of possible futures each second and picks the best one. You can mess with it and watch it adapt. You can also break physics and watch it get surprised.
Why it's exciting
- It's Yann LeCun's big bet. He left Meta in late 2025 to build world models that predict in an abstract space instead of predicting pixels. The architecture is called JEPA. In March 2026, his group released LeWorldModel, the first JEPA that trains stably from raw pixels. It's tiny, about 15M parameters, and trains in hours on one GPU.
- Nobody has put one in a browser. The research agents found no browser demo of a JEPA that plans. You'd be first.
- It's physical AI. It covers a robot, contact physics, and planning. This is the stuff robotics labs care about.
- It's low-level. You write your own WebGPU kernels so that the whole planning loop runs on the viewer's GPU. There's no server and no ONNX runtime, just your shaders.
- It's brand new for you. No LLaMA and no FPGA. It uses a new architecture, a new domain, and a new runtime.
What the demo looks like
- A 3D table with a robot pusher and a T-shaped block. This is the classic "Push-T" robotics task.
- Drag the ghost target anywhere. The robot pushes the block there.
- Grab the block with your mouse and throw it. The robot replans on the fly.
- A "what it's imagining" filmstrip that shows the future the robot is picturing right now.
- A surprise meter. Click Break physics: teleport the block, turn the floor to ice, or let the block pass through a wall. The meter spikes, because the model knows that isn't how the world works.
- A speed counter that shows how many futures per second your GPU imagines.
How it works
- Build the sim. A MuJoCo table, pusher, and T-block. The same scene runs in Python for training and in the browser through the official MuJoCo WebAssembly package.
- Collect data. The pusher moves around randomly and bumps the block. The GPU pod records about a million small 64×64 frames.
- Train the world model. Train LeWorldModel on those frames. It learns to predict what happens next, but in its own compressed "thought space," not in pixels.
- Plan by imagining. To reach a goal, the robot tries about 300 random action sequences in its head, keeps the best ones, and repeats. It then takes the first step of the winner. This method is called CEM.
- Write the kernels. Rewrite the model and the whole planning loop in WGSL, the WebGPU shader language, so that everything stays on the GPU.
- Ship it as a static web page on
vm.ifkash.dev.
Weekend plan
| When | What | Done when |
|---|---|---|
| Saturday morning | Run the official LeWorldModel checkpoint, build the MuJoCo scene | Their model plans; your scene runs in Python and in the browser |
| Saturday afternoon | Generate data, start training (runs by itself) | Training is running |
| Saturday evening | Check planning in Python, add the surprise meter | The robot pushes the block to the goal |
| Sunday morning | Write the WebGPU kernels, match them against PyTorch | Browser output matches PyTorch |
| Sunday afternoon | Build the demo page, deploy it, record the video, write the post | The link works |
Cost
- GPU: 1× RTX 5090 on RunPod, about $0.69 per hour. The fallback is an RTX PRO 6000, about $1.69 per hour.
- Time on the GPU: about 12-18 hours.
- Total: about $10-15. The plan sets a hard stop at $40.
- Hosting: free. It's a static page on your VM, and it keeps working after the pod is gone.
What you have at the end
- A live link where anyone can play with a robot that thinks in latent space.
- A speed chart: your WebGPU kernels against ONNX Runtime Web.
- A video of the surprise meter spiking when you break physics.
- A blog post: "I put LeCun's world model in a browser tab."
What might go wrong
| Problem | What to do |
|---|---|
| Your trained model plans badly | Start from the official checkpoint's recipe, and compare with its numbers. |
| The browser can't plan fast enough | Use fewer samples or shorter plans. 5 plans per second still looks smooth. |
| The browser frames look different from the training frames | The plan draws the model's view with a tiny custom renderer that's identical in Python and the browser. |
| WebGPU isn't available in a visitor's browser | Show a recorded video instead. |
Other ideas the research turned up
- WarpDrive: Make an open video world model, Matrix-Game 2.0, run at 60 FPS on one RTX 5090 with FP4 kernels, then stream it to the browser. This is the most hardcore low-level option, but the demo only works while you pay for the GPU.
- Neural MuJoCo: Replace a robot's physics engine with a neural network, and switch between them live. It's cool, but the port is hard.
- A Jev-style decision model playing a browser game. Jev is closed, and open clones such as LAYA already exist.
Reading, if you want it
- LeWorldModel paper and code, under the MIT license.
- Official MuJoCo WebAssembly bindings.
Part 2: For the coding agent
Mission
Build mindseye, a browser demo of an action-conditioned JEPA world model that plans in latent
space and runs entirely on the client:
- Model: LeWorldModel (LeWM), trained from pixels on a MuJoCo Push-T task.
- Planning: cross-entropy method (CEM) model-predictive control, with the model and the full planning loop written as hand-written WGSL compute kernels.
- Physics: the official MuJoCo WebAssembly bindings.
- Surprise meter: a violation-of-expectation readout based on prediction error.
The final artifact is a static website with no inference server.
Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a
milestone until the previous one passes. After each milestone, commit your work and write a
short entry in NOTES.md with the results and numbers.
Hard constraints
- Secrets: Read
RUNPOD_API_KEY,HF_TOKEN, andWANDB_API_KEYfrom environment variables only. Never write a secret into any file, log, commit, or echoed command. Commit a.env.examplethat has placeholder values only, and add.envto.gitignore. - Budget: The hard cap is $40 of RunPod spend. Every pod runs a watchdog that stops the pod
after
MAX_POD_HOURShours. The default is6. - Compute: Run training and large-scale data generation only on RunPod. The local Mac runs the web build, the browser tests, and the small scripts.
- Outward actions: Ask the user before you do any of the following: create the GitHub
repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything
publicly. Use the
kashifulhaqueGitHub account (gh auth switch -u kashifulhaque). - Shared VM:
vm.ifkash.devruns other production apps under Dokploy, with Traefik on ports 80 and 443. Don't stop, restart, reconfigure, or remove any existing container, service, or proxy rule. Add mindseye only as a new Dokploy application. - Pod cleanup: When a pod isn't running a job, stop it. At the end of the project, terminate every pod that you created and report the total spend.
- Licenses: LeWM is under the MIT license. Keep its license notice in any file that you copy
or adapt, and credit the paper in
README.md.
Upstream facts
These facts come from the LeWM README and the paper. Confirm the details in the repo's configs before you rely on them:
- Repo:
github.com/lucas-maes/le-wm, under the MIT license. It depends on thestable-worldmodelandstable-pretrainingpackages:uv pip install stable-worldmodel[train,env]. - Architecture: A ViT encoder produces one embedding per frame. An autoregressive
predictor,
ARPredictor, is conditioned on an action encoder, and projection MLPs map between spaces. The loss is next-embedding prediction plus a Gaussian regularizer on the embeddings. The model has about 15M parameters. - Commands: Train with
python train.py data=pusht. Evaluate planning withpython eval.py --config-name=pusht.yaml policy=pusht/lewm. Checkpoint paths are relative to$STABLEWM_HOME. - Checkpoints on the Hugging Face Hub:
quentinll/lewm-pusht,lewm-cube,lewm-tworooms, andlewm-reacher. - Record the unknowns. Read the image resolution, patch size, embedding dimension, history
length, and planner hyperparameters from
config/, and record them inNOTES.mdduring M1.
Tech stack
Pin every version. The stack is as follows:
- Python: Python 3.11, PyTorch 2.x,
mujoco(the Python bindings),h5py,numpy,wandb, and LeWM's dependencies, managed withuv. - Browser:
- The official MuJoCo WASM package,
@mujoco/mujocofromgoogle-deepmind/mujoco/wasm. Use the same MuJoCo version as the Python package, and record both versions. three.jsfor the human-facing 3D view.- Plain TypeScript, bundled with
vite, and raw WebGPU with WGSL. - Don't use an ML runtime in the final demo.
onnxruntime-webis allowed only as the benchmark baseline. - Browser tests: Playwright with Chromium, with WebGPU enabled.
- Pods: the
runpodPython SDK, withrunpodctlinside pods.
Repository layout
Create the following layout:
mindseye/
pyproject.toml .env.example .gitignore README.md NOTES.md
infra/pod.py infra/watchdog.sh infra/bootstrap.sh
sim/
pusht.xml # single source of truth for the scene (Python + browser)
obs_render.py # tiny deterministic 64x64 observation rasterizer
collect.py # parallel data generation -> HDF5 in LeWM's format
train/ # thin wrappers/configs around LeWM
configs/mj_pusht.yaml
probe_decoder.py # latent -> 64x64 frame, for visualization only
plan_eval.py # CEM MPC success-rate eval in Python
export_weights.py # safetensors -> flat f16/f32 binary + JSON manifest
golden.py # golden tensors for kernel tests
web/
src/sim.ts # @mujoco/mujoco wrapper, stepping, perturbations
src/obs_render.ts # exact port of obs_render.py
src/gpu/ # WGSL kernels + TS dispatch code
matmul.wgsl layernorm.wgsl attention.wgsl gelu.wgsl
vit_encoder.ts predictor.ts cem.ts surprise.ts
src/view3d.ts # three.js scene, ghost goal, imagination filmstrip
src/main.ts index.html
tests/ # Playwright: kernel golden tests, planner smoke test
bench/ # WGSL vs onnxruntime-web benchmark page
deploy/Dockerfile # nginx:alpine serving web/dist
results/ post/draft.md
M0: Infrastructure
Tasks:
- Write
infra/pod.py. It uses therunpodSDK and readsRUNPOD_API_KEYfrom the environment. It supports the following subcommands: -create: Creates a pod. The GPU preference order is RTX 5090, then RTX PRO 6000, then RTX 4090. Resolve the GPU type IDs at run time by querying the available GPU types. Use an official RunPod PyTorch image with CUDA 12.8 or later. Set a 100 GB volume at/workspaceand expose SSH. Prefer a pod with 16 or more vCPUs for data generation. -status,stop, andterminate. -ssh-info: Prints the SSH command. -cost: Prints the uptime and spend for every pod whose name has the prefixmindseye-. - Write
infra/watchdog.sh. It sleeps forMAX_POD_HOURS, then runsrunpodctl stop pod $RUNPOD_POD_ID. - Write
infra/bootstrap.sh. It clones the repo and LeWM, runsuv sync, setsSTABLEWM_HOME=/workspace/swm, sets up MuJoCo offscreen support (EGL) if you need it, and starts the watchdog.
Acceptance criteria:
- A pod comes up, and
bootstrap.shfinishes without errors. - A test run with
MAX_POD_HOURS=0.05stops the pod within 5 minutes.
M1: Reproduce upstream
Tasks:
- Download
quentinll/lewm-pushtand the Push-T dataset, and run the upstream planning evaluation. Record the success rate and the planning time per step on the pod GPU. - Record in
NOTES.mdevery architecture and planner hyperparameter listed in the upstream facts.
Acceptance criteria:
- The upstream evaluation runs, and the success rate is within the range the paper reports. If it isn't, stop and report the gap before you continue.
M2: MuJoCo Push-T scene and a renderer that matches everywhere
The scene (sim/pusht.xml):
- A flat table, with walls that bound a 0.5 m × 0.5 m workspace.
- A T-shaped block built from two box geoms, with a free joint whose motion stays in the plane. Restrict it with joint constraints, or keep it flat with friction and a low center of mass. Record the choice you make.
- A cylindrical pusher driven by 2D position actuators, or by
mocapwith a weld to a slider. The action is(dx, dy)per control step, clipped to ±2 cm. - Pick a control frequency of about 10 Hz and a timestep that divides it evenly.
- Define the goal as a target pose for the block:
(x, y, θ).
Observation renderer:
The model's input comes from a deterministic top-down orthographic rasterizer, not from the MuJoCo renderer and not from three.js. This makes the pixels identical in training and in the browser.
- The rasterizer maps the poses of the pusher and the block to a 64×64 RGB image with flat colors and no lighting. It uses point-in-polygon tests for the T-block and a circle test for the pusher.
- For anti-aliasing, use 4×4 supersampling with a fixed sample pattern.
obs_render.pyandobs_render.tsmust implement the same arithmetic. Use float32 in numpy andMath.froundin TypeScript where it matters.- If M1 shows that LeWM expects a different resolution, match LeWM, and record the change.
Acceptance criteria:
- The golden-image test passes: over 1,000 random poses, the Python and TypeScript renderers produce the same bytes, or at most 0.1% of pixels differ by 1 LSB or less.
- Physics parity: from the same initial state, run 200 steps with the same actions in Python
mujocoand in the browser. The block pose must stay within 1 mm and 0.5° at every step. If it drifts more, first align the MuJoCo versions and the integrator settings. Record the result, because it matters less than renderer parity.
M3: Data generation
Tasks:
- Write
sim/collect.pyto generate episodes in parallel on the pod's CPUs. Use one process per core. - Mix the following data sources:
- 60% random-walk pusher motion with momentum, biased toward the block so that contacts happen often.
- 30% scripted pushes that move toward a random point on the block's edge and push through it.
- 10% pure noise.
- Randomize the starting poses of the block and the pusher.
- Scale: about 20,000 episodes of 100 steps each, which is 2 million frames. Write them to HDF5 in the exact schema the LeWM dataloader expects. Copy the Push-T dataset's keys, dtypes, and shapes from M1.
- Save a held-out set of 500 episodes.
Acceptance criteria:
- The upstream LeWM dataloader reads the new dataset without any code changes.
- In at least 50% of steps, the pusher touches the block.
results/m3_samples.pngshows a grid of frames.
M4: Training and probes
Tasks:
- Write
train/configs/mj_pusht.yaml, based on upstreampusht.yamland pointed at the M3 data. - Train with the upstream hyperparameters first. Change a hyperparameter only if the training collapses or the loss plateaus early, and record every change.
- Write
probe_decoder.py: a small convolutional decoder that maps a frozen embedding to a 64×64 frame. Train it with L2 plus a light perceptual or SSIM term. It's for visualization only. - Surprise calibration: compute
s_t = ||pred(z_{<=t}, a_t) - enc(o_{t+1})||²on held-out normal data. Store its mean and standard deviation for z-scoring, and save them to the export manifest.
Acceptance criteria:
- The training loss curves are in wandb, and there's no representation collapse: the effective rank of the embeddings is above 50% of the embedding dimension.
- The probe reconstructions at 1, 5, and 10 steps of imagined rollout are recognizable. Save
them as
results/m4_probe.png. - Surprise separation: on 200 perturbed held-out episodes, the z-scored surprise at the perturbation step is above 3 in at least 90% of cases, and above 3 on at most 2% of normal steps. The perturbations are teleporting the block, removing friction, and a pass-through.
M5: Planning in Python
Tasks:
- Write
train/plan_eval.py: CEM MPC in latent space. Start from the planner settings that you recorded in M1. If you have to set them yourself, use 300 samples, a horizon of 5, 3 iterations, 30 elites, and a Gaussian over action sequences. The cost is the distance between the predicted final embedding and the goal embedding, where the goal embedding is the encoding of the rendered goal frame. - Run the policy as receding-horizon control: execute the first action, then replan.
- Evaluation: Run 200 episodes with random start and goal poses and a limit of 200 steps. An episode succeeds when the block is within 2 cm and 15° of the goal.
Acceptance criteria:
- The success rate is 60% or more, or within 10 points of the upstream Push-T number from M1, whichever is lower.
results/m5_planning.jsonand 10 GIFs exist.
M6: WGSL kernels
Export:
- Write
export_weights.pyto produceweights.binwith fp16 weights, plusmanifest.jsonwith the tensor names, shapes, offsets, the normalization constants, and the surprise statistics. - Write
golden.pyto dump the inputs, the intermediate activations after every block, and the outputs for 20 cases, for the encoder, for one predictor step, and for one full CEM iteration with fixed noise.
Kernels:
- Write the following kernels. Use fp16 storage when the device supports
shader-f16; otherwise use fp32. Accumulate in fp32 in both cases. - A tiled matmul that uses workgroup memory.
- A fused LayerNorm.
- A fused GELU with bias add.
- Attention for short sequences: fuse QKᵀ, softmax, and V in one workgroup for each head and each sample.
- The patch embedding.
- Batch the predictor across the CEM samples: the batch is
N_samples, and the sequence is the history plus the horizon. - Keep the whole CEM loop on the GPU:
- Sample with a GPU Philox or PCG random number generator.
- Roll out the predictor, compute the costs, and select the top-k with a bitonic sort or a radix select.
- Update the mean and standard deviation.
- Repeat for every iteration without a CPU readback.
- Read back only the final first action, 8 bytes, and the elite trajectory for the filmstrip.
- Use one command encoder for each replan.
- Measure the GPU time of each pass with timestamp queries, when the browser supports them.
Acceptance criteria:
- The Playwright tests pass in headless Chromium with WebGPU. Against the goldens, every intermediate tensor has a maximum relative error of 2e-2 or less in fp16 and 1e-4 or less in fp32.
- With fixed noise, the CEM run selects the same first action as PyTorch, within 1e-3.
- Benchmark: On the benchmark page, compare replans per second between these kernels and
onnxruntime-webwith the WebGPU execution provider, running the same model with the CEM loop in JavaScript. Report the numbers on an Apple M-series Mac and, if possible, on one other GPU. Target: 10 replans per second or more at the M5 settings on an M-series Mac. Save the results toresults/m6_bench.json.
M7: The demo page
Tasks:
- Layout: A three.js 3D view of the table, rendered from MuJoCo state every frame. Next to it, show the following:
- The model's 64×64 input.
- The "imagination" filmstrip: the probe decoder applied to the elite plan's predicted embeddings, for the horizon steps. Run the probe decoder in WGSL too. It's a few convolutions or transposed convolutions.
- The surprise meter as a sparkline and a gauge.
- A counter for replans per second and futures imagined per second, where futures equals samples times iterations times replans.
- Interactions:
- Drag the translucent ghost T-block to set the goal pose. The scroll wheel rotates it.
- Grab the real block and throw it, using a spring force applied through
xfrc_applied. - Break physics buttons for teleporting the block, setting the friction to 0, and turning off wall collisions. Each one shows the surprise spike.
- A pause button, and a slider for the number of samples.
- Control loop: run physics at the sim rate, and replan at the rate the GPU sustains. Between replans, the pusher executes the rest of the current plan.
- Fallback: if
navigator.gpuis missing, show a looping MP4 of the demo and a short explanation. - Mobile layout: stack the panels vertically, and support touch drag for the goal.
Acceptance criteria:
- A Playwright smoke test loads the page, sets a goal, and checks that the block moves toward it within 30 seconds of simulated time.
- The page works in Chrome and Safari on macOS.
M8: Deploy and write-up
Tasks:
- Write
deploy/Dockerfile:nginx:alpineservingweb/dist, with correct MIME types for.wasmand.binand the cross-origin isolation headers if the MuJoCo multithreaded build needs them. - Ask the user before deploying. After the user approves, create a new Dokploy application on
vm.ifkash.dev, with a domain such asmindseye.ifkash.devthat the user confirms. The user might need to add the DNS record. - Write
README.md. It explains what the project is, how to reproduce each milestone with one command per milestone, the benchmark table, GIFs, and credits and license notes for LeWM and MuJoCo. - Write
post/draft.md, a blog post of 1,200-1,800 words. Structure it as follows: - The hook.
- What a JEPA is, in plain words.
- Training on MuJoCo Push-T.
- Planning by imagining: CEM.
- Moving the whole planner onto WebGPU, with the kernel diagram and the benchmark.
- The surprise meter.
- Limitations: 2D task, a single scene, and a probe decoder that's for display only.
- Ask the user before you push anything. After the user approves, push the repo to
weights-and-wires/mindseyeand upload the checkpoint to the Hugging Face Hub underweights-and-wires/mindseye-lewm-mjpusht. - Terminate every pod, and write the final spend to
NOTES.md.
Acceptance criteria:
- The deployed URL loads and plans in Chrome.
- No pods are left running, and the total spend is less than $40.
Final report to the user
When you finish, report the following:
- The upstream reproduction number from M1.
- The success rates for your model, from M5.
- The surprise separation, from M4.
- Replans per second for WGSL against ONNX Runtime Web, from M6.
- The URL, if the site is deployed.
- The total spend.
- Anything that you skipped or that failed.