# rewind: an AI weather model that predicts yesterday, live in your browser

---

# Part 1: For you

## The idea in one line

A small AI weather model runs on the whole globe, right in your browser tab. It forecasts
tomorrow, and its twin *backcasts* yesterday: it predicts the past from the present. You scrub
time from 10 days back to 10 days ahead, and you drop a "butterfly" anywhere on Earth to watch
what the AI does with it.

## Why it's exciting

- **The hook is strange and true.** AI can predict yesterday's weather from today, and physics
  says it shouldn't be able to. An August 2026 paper,
  [Missing the Butterfly and Predicting the Past](https://arxiv.org/abs/2608.25835), shows that
  AI weather models backcast with real skill. The authors say this "appears to violate the
  second law of thermodynamics."
- **The butterfly is missing.** In real weather, a tiny nudge grows fast and wrecks the
  forecast within a day or two. That's the butterfly effect. The same paper shows that AI
  weather models don't do this: small and large nudges grow at the same rate. You get to show
  it live.
- **There's one clean explanation.** The training data is *coarse-grained*: it drops the small,
  fast details of the weather. A tiny toy system called *Lorenz 96* shows the whole mechanism,
  and you ship it as a second tab.
- **Nobody has built this demo.** The research found no demo that runs a weather model
  backward or lets you plant a butterfly. Existing sites, such as
  [Google DeepMind Weather Lab](https://blog.google/technology/google-deepmind/weather-lab/),
  [CSU AI Weather](https://aiweather.cira.colostate.edu/), and the
  [Hypatia Zero](https://github.com/hypatia-earth/zero) WebGPU renderer, only show forecasts
  that someone computed earlier.
- **It's low-level.** You write the model's core math in WebGPU shaders by hand. One kernel,
  the *spherical harmonic transform*, powers both the model and the energy charts. There's no
  server and no ML runtime.
- **The clock is ticking.** The authors plan to release their models when the paper is
  published. Ship soon so that your interactive version comes first.

## What the demo looks like

- **A 3D globe** that shows one weather field at a time: the 500 hPa height, temperature, wind
  speed, sea-level pressure, or surface temperature.
- **A timeline from -10 to +10 days.** Drag left to backcast and right to forecast. A wipe
  slider compares the model with what really happened.
- **About a dozen real events** from 2021-2022: a cold wave, a heat dome, hurricanes, and a
  European windstorm.
- **A "Rewind" button.** It starts at a storm's peak and backcasts to its birth.
- **Skill charts** for both directions, with the paper's numbers marked for reference, and a
  speed counter in milliseconds per model step.
- **The butterfly tool.** Click anywhere on the globe to plant a tiny wind bump. The page shows
  how the difference spreads, on the globe and on a chart.
- **A Lorenz 96 lab** in a second tab that explains *why* the butterfly goes missing.

## How it works

1. **Get the data.** Download about 40 years of global weather data (ERA5, from WeatherBench 2)
   onto a GPU pod. It's free to read, and it's about 30 GB at the size you need.
2. **Play with the toy first.** While the data downloads, train four tiny models on Lorenz 96
   on your Mac. This shows the whole effect in minutes.
3. **Train two twins.** Train one model that steps the weather 24 hours forward and an
   identical one that steps it 24 hours backward.
4. **Measure them.** Check how many days each direction stays skillful. Plant butterflies at
   three sizes and measure how fast they grow.
5. **Write the kernels.** Rewrite the model in WGSL, the WebGPU shader language, and check it
   against PyTorch number by number.
6. **Ship it** as a static web page on `vm.ifkash.dev`.

## Weekend plan

| When | What | Done when |
|---|---|---|
| Saturday morning | Start the pod and the data pull; build the Lorenz 96 lab on the Mac while it downloads | The data is on the pod; the toy shows the effect |
| Saturday afternoon | Train the forward and backward models (runs by itself) | Both models finish training |
| Saturday evening | Evaluate: skill curves in both directions, and the butterfly test | Backcasts are skillful, but for fewer days than forecasts |
| Sunday morning | Write the WGSL kernels, match them against PyTorch | Browser output matches PyTorch |
| Sunday afternoon | Build the globe page, deploy it, record the video, write the post | The link works |

## Cost

- **GPU:** 1× RTX 5090 on RunPod, about $0.69 per hour. The fallback is an RTX PRO 6000, about
  $1.69 per hour.
- **Time on the pod:** about 12-16 hours.
- **Total:** about $10-15. The plan sets a hard stop at $30.
- **Hosting:** free. It's a static page on your VM, and it keeps working after the pod is gone.

## What you have at the end

- A live link where anyone can rewind the weather and plant a butterfly.
- Skill curves that show forecasts beating backcasts, the same asymmetry the paper reports.
- A video of the Rewind button taking a hurricane back to its birth.
- A blog post: *"I taught a weather model to predict yesterday."*

## What might go wrong

| Problem | What to do |
|---|---|
| The small model backcasts badly, or its long runs blow up | Use the short multi-step fine-tune in the plan. Accept 4.5 skillful days or more; the asymmetry is the story, not the absolute number. |
| The butterfly result doesn't reproduce at this coarse resolution | Report it honestly; it's still interesting. The Lorenz 96 lab shows the mechanism either way. |
| The page is too heavy to load | Keep the weights in fp16 with 15M parameters or fewer per model, and load each event's data only when someone picks it. |
| WebGPU isn't available in a visitor's browser | Show a recorded video and static charts instead. |
| The authors release their models first | The interactive angle and the butterfly tool are still yours. |

## Other ideas the research turned up

- **Watch a 7M-parameter model think.** Run the
  [Tiny Recursive Model](https://arxiv.org/abs/2510.04871) on Sudoku in WebGPU, with
  [PTRM](https://arxiv.org/abs/2605.19943)'s 64 noisy trajectories picked by the Q head
  (Sudoku-Extreme goes from 87.4% to 98.75%). An int4 switch breaks it: per-tensor 4-bit drops
  Sudoku from 84.1% to 0.0% in [this quantization paper](https://arxiv.org/html/2607.16237).
- **[NIV variable fonts](https://arxiv.org/abs/2606.05261).** Drop in a static `.ttf` and get
  weight, width, and slant axes its designer never drew, with a WASM OpenType writer. Its
  research license might forbid hosted services, so ask the authors.
- **[EGGROLL](https://arxiv.org/abs/2511.16652) in a tab.** Train an int8 language model in the
  browser with evolution strategies and no gradients.
- **[Cool-chic 5.0](https://arxiv.org/abs/2605.02726) in WASM and WGSL.** An image codec that
  claims 11% rate savings over VVC, decoded by a tiny network.
- **[Label-free acoustic keyboard eavesdropping](https://arxiv.org/abs/2607.22094),
  on-device.** It's closer to audio territory.

## Reading, if you want it

- [Missing the Butterfly and Predicting the Past](https://arxiv.org/abs/2608.25835): the paper
  this project builds on.
- [Butterfly Effect and the Kinetic Energy Cascade in Probabilistic Machine Learning Weather Prediction Models](https://arxiv.org/abs/2609.18489):
  a separate September 2026 paper that finds the same missing spread in four ensemble models.
- [Spherical Fourier Neural Operators](https://arxiv.org/abs/2306.03838) (Bonev et al., 2023):
  the model architecture this project uses.
- [torch-harmonics](https://github.com/NVIDIA/torch-harmonics): the SHT and SFNO library.
- [WeatherBench 2 data guide](https://weatherbench2.readthedocs.io/en/latest/data-guide.html):
  where the data comes from.
- [Lorenz 63 forecast and backcast movie](https://doi.org/10.5281/zenodo.22062019): a short
  visual from the paper's authors.

---

# Part 2: For the coding agent

## Mission

Build `rewind`, a browser demo of a global AI weather model that runs forward and backward in
time entirely on the client:

- **Models:** Two Spherical Fourier Neural Operators (SFNOs) trained on ERA5 at 1.5°, with one
  shared config: `fwd` steps 24 hours forward, and `bwd` steps 24 hours backward.
- **Kernels:** The spherical harmonic transform (SHT), the SFNO blocks, and the kinetic energy
  (KE) and difference kinetic energy (DKE) diagnostics, as hand-written WGSL compute kernels.
- **Butterfly tool:** Control and perturbed members at three amplitudes, with DKE against lead.
- **Lorenz 96 lab:** A second tab that reproduces the coarse-graining mechanism on a toy system.

The final artifact is a static website with no inference server. Ship it soon: the paper's
authors plan to release their models on publication.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a
milestone until the previous one passes. After each milestone, commit your work and write a
short entry in `NOTES.md` with the results and numbers.

## Hard constraints

- **Secrets:** Read `RUNPOD_API_KEY`, `HF_TOKEN`, and `WANDB_API_KEY` from environment
  variables only. Never write a secret into any file, log, commit, or echoed command. Commit a
  `.env.example` that has placeholder values only, and add `.env` to `.gitignore`.
- **Budget:** The hard cap is $30 of RunPod spend. Every pod runs a watchdog that stops the pod
  after `MAX_POD_HOURS` hours. The default is `6`.
- **Compute:** Run training and the bulk data download only on RunPod. The local Mac runs the
  web build, the browser tests, and the Lorenz 96 lab.
- **Data splits:** Train on 1979-2018, validate on 2019-2020, and test on 2021-2022. No
  training code path may read a test year. The demo events come from test years only.
- **Outward actions:** Ask the user before you do any of the following: create the GitHub
  repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything
  publicly. Use the `kashifulhaque` GitHub account (`gh auth switch -u kashifulhaque`).
- **Shared VM:** `vm.ifkash.dev` runs other production apps behind one shared Caddy, which owns
  ports 80 and 443. Its config lives at `~/docs/caddy`.
  - The rewind container lives in `~/docs/rewind` and must not publish any ports. It joins the
    external Docker network `edge` with a stable alias, such as `rewind`.
  - Add the vhost only by appending to `~/docs/caddy/Caddyfile` with `>>`. Never rewrite,
    rename, or replace that file. It's a single-file bind mount, and a rewrite orphans the
    inode, so `caddy reload` then reports "config is unchanged" while serving the old config.
  - Docker on the VM has no BuildKit for plain `docker build`. Don't use `COPY --chmod`,
    Dockerfile heredocs, or `RUN --mount`. Multi-stage builds and `COPY --from` work.
  - Don't stop, restart, reconfigure, or remove any other container, network, or vhost.
- **Pod cleanup:** When a pod isn't running a job, stop it. At the end of the project,
  terminate every pod that you created and report the total spend.
- **Licenses:** ERA5 comes under the
  [Copernicus license](https://apps.ecmwf.int/datasets/licences/copernicus/), which requires
  attribution. Put "Contains modified Copernicus Climate Change Service information" in
  `README.md` and in the page footer. Confirm the torch-harmonics license during M0, and keep
  its notice in any file that you copy or adapt. Credit both papers in `README.md`.

## Science background

These facts come from the two papers. Implement the experiments to match them:

- **Paper 1:** "Missing the Butterfly and Predicting the Past: Features or Bugs of Accurate AI
  Weather Models?" by Hassanzadeh, Li, Sun, Wang, Wikner, Finkel, and Weare, on arXiv on
  August 26, 2026: <https://arxiv.org/abs/2608.25835>.
  - AI weather prediction (AIWP) models can backcast skillfully, but backcasts are
    systematically less accurate than forecasts. Backcasting models are architecturally
    identical to forecasting models, with the input and output pairs reversed.
  - **ERA5 experiment:** A deterministic Pangu-Weather-based Transformer at about 100 km, with
    Δt = 24 h and about 40 years of ERA5. With ACC > 0.6 on Z500 as the skill threshold, the
    forecast is skillful to 9.1 days and the backcast to 6.5 days. PlaSim (about 300 km) with a
    Transformer, a GenCast-like diffusion model, and neural operators shows larger asymmetries.
  - **Butterfly metric:** DKE is the area-weighted ensemble variance of the horizontal winds,
    for amplitudes that decrease from a reference ε0. In physics models, the smallest
    amplitudes grow by orders of magnitude within a day or two. AIWP models show
    amplitude-independent exponential DKE growth, both forward and backward.
  - **Single cause:** coarse-graining of the training data, which means filtering small and
    fast scales, dropping variables, or both. With the official Pangu-Weather 24 h, 6 h, 3 h,
    and 1 h models, butterfly-like growth emerges and skill degrades as Δt shrinks.
- **Paper 2:** "Butterfly Effect and the Kinetic Energy Cascade in Probabilistic Machine
  Learning Weather Prediction Models" by Chen, Oskarsson, Driscoll, and Schemm, on arXiv on
  September 16, 2026: <https://arxiv.org/abs/2609.18489>. With KE and DKE spectra,
  NeuralGCM-ENS, FourCastNet 3, AIFS-ENS, and GenCast don't reproduce the small-scale spread
  growth of IFS-ENS.
- **Two-scale Lorenz 96:** The system is defined by the following equations:

  ```
  dX_k/dt     = -X_{k-1} (X_{k-2} - X_{k+1}) - X_k + F - (h c / b) Σ_j Y_{j,k}
  dY_{j,k}/dt = -c b Y_{j+1,k} (Y_{j+2,k} - Y_{j-1,k}) - c Y_{j,k} + (h c / b) X_k
  ```

  The paper compares two training regimes:

  | Regime | Variables | Δt | Paper's finding |
  |---|---|---|---|
  | Real-world (coarse) | X only | 10δt | Better forecast skill; no butterfly-like effect |
  | Perfect-data (full) | X and Y | δt | Skill worse by a factor of about 7; a butterfly-like effect emerges |

## Design decisions

- **Data:** The WeatherBench 2 public GCS bucket, read anonymously:
  `gs://weatherbench2/datasets/era5/1959-2023_01_10-6h-240x121_equiangular_with_poles_conservative.zarr`
  It's 1.5°, a 240×121 equiangular grid with poles, 6-hourly, from 1959 to 2023-01-10, with
  13 pressure levels. For details, see the
  [WeatherBench 2 data guide](https://weatherbench2.readthedocs.io/en/latest/data-guide.html).
- **Architecture:** An SFNO instead of the paper's Pangu-style Transformer. One kernel, the SHT,
  powers both the model and the KE spectrum analysis, and the paper reports the asymmetry in
  neural-operator models too. Record this deviation in `README.md`.
- **State:** 16 prognostic channels on the 121×240 grid: `z`, `t`, `u`, and `v` at 250, 500,
  and 850 hPa, plus `2m_temperature`, `mean_sea_level_pressure`, `10m_u_component_of_wind`,
  and `10m_v_component_of_wind`.
- **Other inputs:** `land_sea_mask`, `geopotential_at_surface`, and sin and cos of latitude.
  Time forcings are sin and cos of the day of year and of the hour of the *target* time, for
  the backward model too.
- **Samples:** 00Z and 12Z states, fp16 on disk, normalized per channel with training-year
  statistics. At 121 × 240 × 16 × 2 bytes, a state is about 0.93 MB, so the dataset is about
  30 GB.
- **Models:** `fwd` maps x_t to x_{t+24h}, and `bwd` maps x_t to x_{t-24h}, with identical
  configs. Both predict the residual. Target 15M parameters or fewer each, which is 30 MB or
  less in fp16, so that the page loads. Stretch: `fwd6h`, a 6 h step rolled out 4 times, to
  reproduce the paper's Pangu Δt trend.

## Tech stack

Pin every version. The stack is as follows:

- **Python:** Python 3.11, PyTorch 2.x, `torch-harmonics`, `xarray`, `zarr`, `gcsfs`, `numpy`,
  `matplotlib`, and `wandb`, managed with `uv`.
- **torch-harmonics:** `th.RealSHT(nlat, nlon, grid="equiangular")`, its inverse and vector
  variants, and the SFNO example models. It also ships a shallow-water example.
- **Browser:** TypeScript, bundled with `vite`, raw WebGPU with WGSL, and `three.js` for the
  globe. Don't use an ML runtime in the final demo. `onnxruntime-web` is allowed only as the
  benchmark baseline.
- **Browser tests:** Playwright with Chromium, with WebGPU enabled.
- **Pods:** the `runpod` Python SDK, with `runpodctl` inside pods.

## Repository layout

Create the following layout:

```
rewind/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh
  data/pull.py           # WB2 zarr (anonymous gcsfs) -> local fp16 memmap/zarr
  data/stats.py          # per-channel normalization stats, climatology
  l96/sim.py             # two-scale L96, RK4 integrator
  l96/train.py           # 4 small nets: {coarse, full} x {fwd, bwd}
  train/
    configs/fwd.yaml  configs/bwd.yaml  configs/fwd6h.yaml
    model.py             # SFNO wrapper around torch-harmonics
    train.py             # one-step + short autoregressive fine-tune
    eval.py              # ACC/RMSE vs lead, both directions
    butterfly.py         # perturbation ensembles, DKE vs lead
    spectra.py           # KE and DKE spectra via vector SHT
    export_weights.py    # -> weights.bin (fp16) + manifest.json + SHT tables
    golden.py            # per-block activations + 10-day rollout
  web/
    src/gpu/             # WGSL kernels + TS dispatch code
      sht.wgsl  sfno.wgsl  mlp.wgsl  spectrum.wgsl  dke.wgsl  colormap.wgsl
      sht.ts  sfno.ts  spectrum.ts  dke.ts
    src/globe.ts         # three.js globe, field textures
    src/timeline.ts      # -10..+10 day scrubber, Rewind button
    src/butterfly.ts     # plant a bump, run members, DKE chart
    src/charts.ts        # ACC chart, KE spectrum, ms/step counter
    src/l96/             # TS integrator + tiny nets for the lab tab
    src/main.ts  index.html
    public/events/       # per-event t0 state + truth frames (precompressed)
    tests/               # Playwright: kernel golden tests, smoke test
    bench/               # WGSL vs onnxruntime-web benchmark page
  deploy/Dockerfile  deploy/compose.yml  deploy/nginx.conf
  results/  post/draft.md
```

## M0: Infrastructure

**Tasks:**

1. Write `infra/pod.py`. It uses the `runpod` SDK and reads `RUNPOD_API_KEY` from the
   environment. It supports the following subcommands:
   - `create`: Creates a pod. The GPU preference order is RTX 5090, then RTX PRO 6000, then
     RTX 4090. Resolve the GPU type IDs at run time by querying the available GPU types. Use an
     official RunPod PyTorch image with CUDA 12.8 or later. Set a 150 GB volume at `/workspace`
     and expose SSH.
   - `status`, `stop`, and `terminate`.
   - `ssh-info`: Prints the SSH command.
   - `cost`: Prints the uptime and spend for every pod whose name has the prefix `rewind-`.
2. Write `infra/watchdog.sh`. It sleeps for `MAX_POD_HOURS`, then runs
   `runpodctl stop pod $RUNPOD_POD_ID`.
3. Write `infra/bootstrap.sh`. It clones the repo, runs `uv sync`, and starts the watchdog.

**Acceptance criteria:**

- A pod comes up, and `bootstrap.sh` finishes without errors.
- A test run with `MAX_POD_HOURS=0.05` stops the pod within 5 minutes.

## M1: Data

**Tasks:**

- Write `data/pull.py`. It opens the WB2 store with `xarray`, `zarr`, and anonymous `gcsfs`,
  selects the 16 prognostic channels at 00Z and 12Z plus the static fields, and writes them to a
  local fp16 memmap or zarr store on the pod volume. Run it on the pod only.
- Write `data/stats.py`. It computes per-channel means and standard deviations from the
  training years only, and a 1990-2019 day-of-year climatology from the same store for ACC
  anomalies. Confirm which climatology WB2 uses, and record it and your choice in `NOTES.md`.
- Make the loader assert that no training or validation code path ever reads a 2021 or 2022
  sample.
- Plot `results/m1_fields.png`, a quick check of one state's main fields.
- Compute persistence and climatology baselines: Z500 and T850 ACC and RMSE against lead from
  1 to 10 days on 2021-2022, so that later numbers have context.

**Acceptance criteria:**

- `NOTES.md` documents the shapes and dtypes of every array, and the dataset has no NaNs.
- The loader assertion exists and has a test, and the baselines are in `NOTES.md`.

## M2: Lorenz 96 lab

Do this milestone on the Mac while the M1 download runs. It's cheap and needs no pod.

**Tasks:**

- Write `l96/sim.py`: the two-scale Lorenz 96 system from the science background, with RK4.
  Start from the standard values K = 36, J = 10, F = 10, h = 1, b = 10, c = 10, and
  δt = 0.005. Check the paper's supplementary settings, use those if it gives them, and record
  your choice in `NOTES.md`.
- Write `l96/train.py`. It trains four small MLPs, or 1D periodic convolutional nets: the
  coarse and full regimes from the science background, each forward and backward.
- Evaluate each net on X skill against lead, and record the skill metric you pick.
- Run the butterfly test on each net: perturb the initial state at three decreasing
  amplitudes, and track the ensemble spread against lead.
- **The "second law" angle:** Show that integrating the true system backward is unstable, yet
  the learned backcaster stays stable. Keep this brief.
- Export the nets' weights to JSON. Write `web/src/l96/` with the true integrator and the tiny
  nets in plain TypeScript. This tab doesn't need WebGPU.

**Acceptance criteria:**

- The coarse forward net beats the full forward net on X skill. Record the ratio; the paper
  reports about 7×.
- Only the full nets show butterfly-like fast initial growth. The coarse nets' growth doesn't
  depend on the amplitude.
- The TypeScript integrator and nets match the Python ones from the same initial state.
  Record the maximum error in `NOTES.md`.

## M3: Train fwd and bwd

**Tasks:**

- Write `train/model.py`, which wraps the torch-harmonics SFNO with the inputs and outputs from
  the design decisions. Confirm the class import path and spectral weight layout.
- Loss: latitude-weighted MSE on the normalized fields. Train one step first, then fine-tune
  autoregressively for 2-4 steps.
- Write `train/configs/fwd.yaml` and `train/configs/bwd.yaml`, identical except for the
  direction. Log both runs to wandb. The estimate is about 4-5 hours per model on an RTX 5090.
- Stretch goal, only if the budget allows: train `fwd6h` and roll it out 4 times per 24 hours.

**Acceptance criteria:**

- Each model has 15M parameters or fewer, and both loss curves are in wandb.
- A 10-day `fwd` rollout from a validation date stays stable, with no blow-up.

## M4: Evaluate, butterfly, and spectra

**Tasks:**

- Write `train/eval.py`: Z500 and T850 ACC and RMSE against lead from 1 to 10 days, in both
  directions, on 2021-2022. Use anomalies against the M1 climatology. Plot
  `results/m4_skill.png`, and add the M1 baselines and the 0.6 line.
- Write `train/butterfly.py`. For `fwd` and `bwd`, at 50 test dates:
  - Perturb `u` and `v` in the initial state with a localized Gaussian bump with a radius of
    about 500 km, at amplitudes ε0, ε0/10, and ε0/100. Pick ε0, for example 1 m/s, and record
    the choice.
  - Run N = 8 members of random bumps per amplitude, and compute DKE against lead per the
    paper's definition.
  - Fit an exponential growth rate for each amplitude over the early leads, and record the
    window you fit.
- Write `train/spectra.py`: KE and DKE spectra with the vector SHT, for ERA5 and the models.
- Record the butterfly result honestly, whatever it is. The paper's expectation is
  amplitude-independent growth.

**Acceptance criteria:**

- `fwd` keeps Z500 ACC above 0.6 for at least 7 days. The paper reached 9.1 days at about
  100 km; this model is coarser and smaller.
- `bwd` stays skillful for at least 4.5 days, and for fewer days than `fwd`. If `bwd` matches
  or beats `fwd`, stop and check for a data leak, such as the time forcings or split overlap.
- The 10-day `fwd` KE spectrum stays within a factor of 3 of ERA5 for l ≤ 60.
- The DKE plots and a growth-rate table exist in `results/`.

## M5: Export and goldens

**Tasks:**

- Write `train/export_weights.py`. It produces `weights.bin` with fp16 weights for each model,
  plus `manifest.json` with the tensor names, shapes, offsets, normalization statistics, and
  the metadata for the SHT tables.
- Export the SHT tables straight from torch-harmonics, so that the browser matches it exactly:
  the longitude DFT twiddles for m ≤ mmax, and the associated Legendre values multiplied by the
  quadrature weights for the equiangular grid.
- Write `train/golden.py`. It dumps the inputs, the activations after every block, and the
  outputs for 10 cases, for both models. It also dumps a 10-day rollout for each model.

**Acceptance criteria:**

- Reloading `weights.bin` through the manifest reproduces the goldens within fp16 rounding, and
  a NumPy SHT built from the exported tables matches torch-harmonics. Record the errors.

## M6: WGSL kernels

**Kernels:**

- Write the following kernels. Use fp16 storage when the device supports `shader-f16`;
  otherwise use fp32. Accumulate in fp32 in both cases.
  - **Forward and inverse real SHT:** first the longitude DFT as a matmul with the precomputed
    twiddles for m ≤ mmax, then the Legendre transform as a batched matmul for each m. The
    grid has 240 longitudes, which isn't a power of two, and a DFT as a matmul is simpler and
    fast enough. Say so in `README.md`.
  - **SFNO spectral channel mixing,** with complex weights per l or per (l, m), matching what
    the pinned torch-harmonics SFNO does. Confirm the layout from the M5 manifest.
  - **The rest of the SFNO:** the pointwise MLP with a fused GELU, the SFNO's normalization
    layer, the encoder and decoder 1×1 convolutions, the residual add, and the denormalization.
  - **Diagnostics:** the KE spectrum through the vector SHT of `(u, v)`, and the DKE reduction
    (the area-weighted ensemble variance).
- Batch the ensemble members in one dispatch.

**Acceptance criteria:**

- The Playwright tests pass in headless Chromium with WebGPU. Against the goldens, every
  intermediate tensor has a maximum relative error of 2e-2 or less in fp16 and 1e-4 or less in
  fp32.
- The day-1 Z500 matches PyTorch within 0.1% RMSE, and the day-10 Z500 ACC is within 0.01 of
  PyTorch's.
- The KE spectrum passes a Parseval check: its sum matches the grid-point area-weighted
  ½|u|² within 1%.
- One 24-hour step takes 250 ms or less on an Apple M-series Mac, so a 10-day rollout takes
  about 3 seconds or less.
- **Benchmark:** Compare ms per step between these kernels and `onnxruntime-web` with the
  WebGPU execution provider, running the same model. Save the results to
  `results/m6_bench.json`.

## M7: The demo page

**Tasks:**

- **Globe:** A three.js globe. A field selector for Z500, T850, wind speed at 850 hPa, MSLP, and
  T2M. A WGSL pass colormaps the selected field to a texture.
- **Events:** About 12 events from the 2021-2022 test period, never from the training years.
  Suggestions: the February 2021 North American cold wave, the June 2021 Pacific Northwest heat
  dome, Hurricane Ida (August 2021), Storm Eunice (February 2022), Hurricane Ian (September
  2022), and the late December 2022 North American winter storm. Verify every date in the data.
  The store ends on 2023-01-10, so check that each +10 day window fits.
- **Event data:** For each event, ship the full 16-channel state at t0 (about 0.93 MB in fp16)
  and daily Z500 and T850 truth for -10 to +10 days (about 2.4 MB). Precompress the files, and
  lazy-load each event.
- **Controls:** A timeline scrubber from -10 to +10 days (left is backcast, right is forecast),
  a truth and model wipe slider, and a **Rewind** button that starts at an event's peak and
  backcasts to its birth.
- **ACC chart:** Live ACC for both directions, with the 0.6 line and the paper's 9.1-day and
  6.5-day marks for reference.
- **Butterfly tool:** Click the globe to plant a bump. The page runs control and perturbed
  members at three amplitudes, shows the difference field on the globe, and plots DKE on a
  log scale.
- **Extras:** A KE spectrum mini-chart, a ms-per-step counter, the M2 lab as a second tab, and
  a footer with the Copernicus attribution and credits for both papers.
- **Fallback and mobile:** If `navigator.gpu` is missing, show a precomputed rollout video and
  static charts. On mobile, stack the panels and support touch to rotate the globe.

**Acceptance criteria:**

- A Playwright smoke test loads the page, runs a 5-day forecast and a 5-day backcast, and
  checks that the ACC chart has 10 points.
- The page works in Chrome and Safari on macOS.

## M8: Deploy and write-up

**Tasks:**

- Write `deploy/Dockerfile`: `nginx:alpine` serving `web/dist`, buildable with the legacy
  builder. Serve the precompressed `.bin` files with `gzip_static on`; serve brotli only if you
  confirm that the image has a brotli module. Set `application/octet-stream` for `.bin` and
  `application/wasm` for `.wasm`.
- Write `deploy/compose.yml`. It publishes no ports and joins the external network `edge` with
  the alias `rewind`.
- Ask the user before deploying. The user confirms a domain such as `rewind.ifkash.dev` and
  adds a non-proxied Cloudflare A record.
- After the user approves, copy the build and `deploy/` to `~/docs/rewind`, and run
  `docker compose up -d` there. Then compare the live and host Caddyfiles with
  `docker exec caddy cat /etc/caddy/Caddyfile | diff - ~/docs/caddy/Caddyfile`. If they
  differ, stop and ask the user. Otherwise, back up, append, validate, and reload:

  ```
  cd ~/docs/caddy
  cp Caddyfile "Caddyfile.bak-rewind-$(date +%Y%m%d)"
  printf '\nrewind.ifkash.dev {\n\treverse_proxy rewind:80\n}\n' >> Caddyfile
  docker exec -i caddy caddy validate --config - --adapter caddyfile < Caddyfile
  docker exec -i caddy caddy reload --config - --adapter caddyfile < Caddyfile
  ```

  Omit any `tls` block: for a non-proxied A record, the default ACME HTTP-01 challenge works.
- Write `README.md`. It explains what the project is, how to reproduce each milestone with one
  command per milestone, the results tables, the SFNO deviation from the paper, and credits and
  license notes for ERA5, torch-harmonics, and both papers.
- Write `post/draft.md`, a blog post of 1,200-1,800 words. Structure it as follows:
  - The hook: "I taught a weather model to predict yesterday."
  - What AI weather models are, in plain words.
  - Training on ERA5.
  - Predicting the past, and why it "shouldn't" work.
  - The butterfly that isn't there.
  - Coarse-graining, explained with Lorenz 96.
  - The SHT in WGSL, with the kernel diagram and the benchmark.
  - Limitations: 1.5° and a 24-hour step, deterministic models only, 16 channels, and not a real
    forecast service.
- Ask the user before you push anything. After the user approves, push the repo to
  `weights-and-wires/rewind` and upload the checkpoints to the Hugging Face Hub under
  `weights-and-wires/rewind-sfno`.
- Terminate every pod, and write the final spend to `NOTES.md`. The owner lists the project on
  projects.dotslasha.me after the deploy; that isn't your job.

**Acceptance criteria:**

- The deployed URL loads over HTTPS and runs a forecast and a backcast in Chrome.
- Every other site on the VM still responds as it did before the deploy.
- No pods are left running, and the total spend is less than $30.

## Final report to the user

When you finish, report the following:

- The persistence and climatology baselines, from M1.
- The skillful days for `fwd` and `bwd` at ACC > 0.6, from M4.
- The butterfly growth-rate table for each amplitude and direction, from M4.
- The Lorenz 96 results: the skill ratio and the butterfly result for each net, from M2.
- The kernel parity numbers and ms per step, with the benchmark, from M6.
- The URL, if the site is deployed.
- The total spend.
- Anything that you skipped or that failed.
