← projects

clatter

Hear what any 3D object sounds like, live in your browser

19 min read clatter.md

Part 1: For you

The idea in one line

Drop any 3D object into a browser scene, hit it, drop it, or throw it, and hear the sound it would really make. The sound comes from a neural network that learned the physics of vibration. Nobody recorded it.

Why it's cool

  • It's physics you can hear. When you hit a mug, it rings at a few exact pitches called modes. The pitches depend on its shape and material. Computing them normally takes a slow physics solve, called FEM, for every object.
  • The neural network skips the solve. You give it a shape, and it predicts the modes in milliseconds. So you can upload any 3D model and hear it instantly.
  • Nobody has shipped this in a browser. Papers such as NeuralSound showed the idea works, but only in videos. The research agents found no live demo where you can drop in your own object.
  • It's physical AI for your ears. It covers simulation, a neural shortcut for a physics solver, and a physics engine that drives the sound.
  • It's low-level. You write a real-time audio engine in C++, compiled to WebAssembly, that rings thousands of modes at once without a single glitch.
  • It's not ASR or TTS, and it doesn't build on any of your past projects.

What the demo looks like

  • A 3D room with a shelf of objects: a mug, a bell, a wine glass, a wrench, and a weird-shaped sculpture.
  • Click anywhere on an object to tap it. Tapping a different spot gives a different sound, just like in real life.
  • Change the material: glass, steel, wood, or ceramic. The sound changes instantly.
  • Make it bigger or smaller. Bigger objects sound lower.
  • Press "Chaos" to drop 30 objects on the floor and hear them all clatter, with the sound coming from the direction of each object.
  • Upload your own 3D model (.glb). In about a second, it has a voice.
  • An A/B toggle: hear the neural version against the slow "true" physics version, to show how close it is.

How it works

  1. Make the answers. On a GPU pod, run the slow physics solve on about 15,000 random 3D shapes. This gives you the true modes of each shape.
  2. Train the shortcut. A small 3D network looks at a shape, as 32×32×32 blocks, and predicts its modes: the pitches, and how loud each pitch is at each spot on the surface.
  3. Material is free. Pitch scales with a simple formula for stiffness and density. So you train on one material and get every material for free.
  4. Build the sound engine. Each mode is a tiny ringing filter. The engine runs thousands of them in C++ compiled to WebAssembly, inside the browser's audio thread.
  5. Connect the physics. A physics engine (Rapier) detects every hit: where it happened and how hard. That hit rings the right modes.
  6. Ship it as a static web page on vm.ifkash.dev.

Weekend plan

When What Done when
Saturday morning Write the physics solver, check it on a few shapes A metal bar rings at the pitch the textbook predicts
Saturday afternoon Generate 15,000 shapes and their modes on the pod (runs by itself) The dataset is done
Saturday night Train the network (runs by itself) Predicted pitches match the true ones
Sunday morning Build the WebAssembly sound engine and the physics scene Tapping an object makes a sound
Sunday afternoon Upload support, Chaos mode, deploy, record the video, write the post The link works

Cost

  • GPU: 1× RTX 5090 on RunPod, about $0.69 per hour. Pick a pod with lots of CPU cores, because the physics solve runs on the CPU.
  • Time on the pod: about 10-15 hours.
  • Total: about $8-15. The plan sets a hard stop at $30.
  • Hosting: free. It's a static page, and it keeps working after the pod is gone.

What you have at the end

  • A live link where anyone can hear their own 3D models.
  • A chart: neural pitches against true pitches, and how fast each one is.
  • A video of 30 objects clattering, with no recorded sound anywhere.
  • A blog post: "No sound files: every sound here is predicted from physics."

What might go wrong

Problem What to do
The physics solver takes all Saturday Use NeuralSound's open-source code to make the data.
Symmetric objects confuse the network, because some modes come in pairs Train on "how the sound spectrum looks at each spot" instead of on single modes. The plan explains how.
The audio glitches when too many objects ring Turn off quiet modes early, and cap how many modes ring at once.
Uploaded models are broken, for example with holes or no inside Fill them into solid blocks before the network sees them.

Other ideas the research turned up

  • "Mm-hm." A browser tab that listens like a person does: it says "mm-hm" at the right moments and knows when you've finished talking. It's small, runs in the browser, and costs about $10. It's the best pick if you want speech specifically.
  • A playable piano world model. Play keys in the browser, and a model predicts the piano sound. Record a phrase, and the model works out which keys made it. It's based on a July 2026 paper, Music-JEPA.
  • Walk into a photo. Take a photo of a room, then move around inside it and hear your voice echo the way it would there.

Reading, if you want it


Part 2: For the coding agent

Mission

Build clatter, a browser app that synthesizes physically based impact sounds for arbitrary 3D objects in real time:

  • Surrogate model: A neural network, trained on FEM modal analysis ground truth, maps a voxelized shape to its modal parameters. It runs in the browser.
  • Material: Material and scale are applied analytically.
  • Audio engine: A C++ to WASM SIMD modal resonator bank in an AudioWorklet renders audio, excited by contact events from a Rapier physics simulation, with per-object spatial panning.

The final artifact is a static website with no server.

Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a milestone until the previous one passes. After each milestone, commit your work and write a short entry in NOTES.md with the results and numbers.

Hard constraints

  • Secrets: Read RUNPOD_API_KEY, HF_TOKEN, and WANDB_API_KEY from environment variables only. Never write a secret into any file, log, commit, or echoed command. Commit a .env.example that has placeholder values only, and add .env to .gitignore.
  • Budget: The hard cap is $30 of RunPod spend. Every pod runs a watchdog that stops the pod after MAX_POD_HOURS hours. The default is 6.
  • Compute: Run large-scale data generation and training only on RunPod. The local Mac runs the web build, the browser tests, the audio engine tests, and small solver checks.
  • Outward actions: Ask the user before you do any of the following: create the GitHub repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything publicly. Use the kashifulhaque GitHub account (gh auth switch -u kashifulhaque).
  • Shared VM: vm.ifkash.dev runs other production apps under Dokploy, with Traefik on ports 80 and 443. Don't stop, restart, reconfigure, or remove any existing container, service, or proxy rule. Add clatter only as a new Dokploy application.
  • Pod cleanup: When a pod isn't running a job, stop it. At the end of the project, terminate every pod that you created and report the total spend.
  • Licenses: Check the license of every mesh dataset and of NeuralSound's code before you use them. Record the licenses in README.md. Bundle only meshes whose license allows redistribution in the demo.

Physics background

Implement the physics exactly as described here:

  • Linear modal analysis. Discretize the solid with finite elements. Solve the generalized eigenproblem K u = λ M u, where K is the stiffness matrix and M is the mass matrix. The free-free boundary gives the object no fixed points. Discard the 6 rigid-body modes, where λ ≈ 0, and keep the next K_modes = 64 modes.
  • Frequencies. The mode frequency is f_i = sqrt(λ_i) / (2π).
  • Material scaling. For a linear isotropic material with Poisson ratio ν, the modes depend on Young's modulus E, density ρ, and length scale s only through a scalar: f_i(E, ρ, s) = f_i^unit * sqrt(E/ρ) / s. The mode shapes don't change. Train at unit E/ρ and unit scale, and fix ν at a value such as 0.3. Treat ν per material as an approximation, and record this in the limitations.
  • Damping. Use Rayleigh damping, C = α M + β K. This gives a per-mode decay rate of d_i = (α + β ω_i²) / 2, where ω_i = 2π f_i. Each material preset sets α and β.
  • Excitation. An impulse J at a surface point p with normal n excites mode i with amplitude a_i = (u_i(p) · n) * J / (ω_i), up to a global constant. Use the mass-normalized mode shapes u_i.
  • Rendering. Each mode is a damped sinusoid: y(t) = Σ a_i e^{-d_i t} sin(ω_i t). Implement each mode as a two-pole resonator: y[n] = 2 r cos(θ) y[n-1] - r² y[n-2] + g x[n], where r = e^{-d/fs} and θ = ω/fs.
  • Hearing range. Drop modes above 20 kHz or below 20 Hz after material scaling.
  • Acoustic radiation. Don't model radiation in the core scope. Treat it as a stretch goal (FFAT maps, as in NeuralSound).

Material presets

Start from published values, and tune α and β by ear in M5. The table lists the starting values:

Material E (Pa) ρ (kg/m³) ν α β
Steel 2.0e11 7850 0.29 5 3e-8
Glass 6.2e10 2600 0.20 1 1e-8
Ceramic 7.2e10 2700 0.19 6 1e-7
Wood 1.1e10 750 0.25 60 2e-6
Plastic 1.4e9 1070 0.35 30 1e-6

Tech stack

Pin every version. The stack is as follows:

  • Python: Python 3.11, PyTorch 2.x, numpy, scipy, trimesh, and wandb, managed with uv. For voxelization, use trimesh or a custom solid voxelizer. Use cupy only if GPU LOBPCG is needed.
  • Browser:
  • TypeScript and vite, with three.js for rendering.
  • @dimforge/rapier3d-compat for physics.
  • onnxruntime-web with the WebGPU execution provider, falling back to WASM, for the surrogate.
  • Audio engine: C++17 compiled with Emscripten to WASM with SIMD (-msimd128), run inside an AudioWorklet. Share memory through a SharedArrayBuffer ring buffer if cross-origin isolation is available; otherwise use MessagePort.
  • Tests: Playwright with Chromium, and native C++ unit tests for the engine, using the same source code.
  • Pods: the runpod Python SDK, with runpodctl inside pods.

Repository layout

Create the following layout:

clatter/
  pyproject.toml  .env.example  .gitignore  README.md  NOTES.md
  infra/pod.py  infra/watchdog.sh  infra/bootstrap.sh
  fem/
    voxelize.py        # mesh -> solid 32^3 occupancy (+ scale normalization)
    hexfem.py          # voxel hex8 linear-elastic K, M assembly (sparse)
    modal.py           # eigensolve, rigid-mode removal, mass-normalization
    validate.py        # analytic checks (bar, plate, cube)
    make_dataset.py    # parallel over meshes -> shards (npz/zarr)
  train/
    model.py           # 3D CNN / UNet surrogate
    losses.py          # frequency loss + permutation-invariant spectral loss
    train.py  eval.py
    export_onnx.py     # + golden I/O for browser tests
  engine/              # C++ modal resonator bank (native + WASM)
    modal_bank.h  modal_bank.cpp  test_modal_bank.cpp  build_wasm.sh
  web/
    src/voxelize.ts    # exact port of voxelize.py
    src/surrogate.ts   # ORT-web wrapper, feature -> modes
    src/worklet.ts     # AudioWorklet processor hosting the WASM bank
    src/physics.ts     # Rapier world, contact events -> excitations
    src/scene.ts  src/ui.ts  src/main.ts  index.html
    public/objects/    # bundled demo meshes + precomputed FEM "truth" modes
    tests/
  deploy/Dockerfile    # nginx:alpine, COOP/COEP headers for SharedArrayBuffer
  results/  post/draft.md

M0: Infrastructure

Tasks:

  1. Write infra/pod.py. It uses the runpod SDK and reads RUNPOD_API_KEY from the environment. It supports the following subcommands: - create: Creates a pod. The GPU preference order is RTX 5090, then RTX 4090, then RTX PRO 4500. Resolve the GPU type IDs at run time by querying the available GPU types. Prefer a pod with the most vCPUs available, 16 or more, because the eigensolves run on the CPU. Use an official RunPod PyTorch image with CUDA 12.8 or later. Set a 100 GB volume at /workspace and expose SSH. - status, stop, and terminate. - ssh-info: Prints the SSH command. - cost: Prints the uptime and spend for every pod whose name has the prefix clatter-.
  2. Write infra/watchdog.sh. It sleeps for MAX_POD_HOURS, then runs runpodctl stop pod $RUNPOD_POD_ID.
  3. Write infra/bootstrap.sh. It clones the repo, runs uv sync, and starts the watchdog.

Acceptance criteria:

  • A pod comes up, and bootstrap.sh finishes without errors.
  • A test run with MAX_POD_HOURS=0.05 stops the pod within 5 minutes.

M1: FEM solver and validation

Tasks:

  • In voxelize.py, turn a mesh into a solid voxel grid:
  • Normalize the mesh so that its bounding-box diagonal equals 1.
  • Center it in a 32³ grid.
  • Fill the inside with parity ray casting, or with a winding-number test for meshes that aren't watertight.
  • Remove floating islands, and keep the largest 6-connected component.
  • Return the occupancy and the physical voxel size h.
  • In hexfem.py, assemble the stiffness matrix K and the consistent mass matrix M as sparse CSR matrices:
  • Use trilinear hexahedral (hex8) elements, one per occupied voxel.
  • Share nodes between neighboring voxels.
  • Use 2×2×2 Gauss quadrature for the element matrices.
  • Precompute a single reference element stiffness matrix for unit h, and scale it by h, because every voxel is the same cube.
  • Use unit E and ρ, with ν = 0.3.
  • In modal.py, find the smallest 70 eigenpairs with scipy.sparse.linalg.eigsh(K, k=70, M=M, sigma=-small, which='LM') in shift-invert mode.
  • Drop the 6 modes where λ < 1e-6 * λ_7.
  • Keep 64 modes, and mass-normalize them so that uᵀ M u = 1.
  • Record the number of nodes and the solve time.
  • If NeuralSound's repo, hellojxt/NeuralSound, provides a working voxel FEM and dataset pipeline, you may use it instead. Record the choice, and still run the validation.
  • In validate.py, check the solver against analytic results:
  • The first bending frequencies of a free-free slender bar, at 32×2×2 voxels, compared with Euler-Bernoulli beam theory.
  • The ratio of the first two frequencies of a thin square plate.
  • Frequencies scaling as 1/s when the voxel size changes.

Acceptance criteria:

  • For the bar, the first frequency is within 5% of Euler-Bernoulli theory, or within 10% for the voxel approximation if the error is explained in NOTES.md.
  • The scaling test is exact to 1e-6.
  • A single 32³ solve finishes in 10 seconds or less on one CPU core, for a typical object with about 10,000 occupied voxels.

M2: Dataset

Tasks:

  • Meshes: Pick a source with a license that allows use, such as Objaverse objects under CC-BY, Thingi10K, or the ABC dataset, which NeuralSound uses. Mix in procedural primitives so that thin shells and bars are well represented: cylinders, hollow cups, bowls, plates, bars, bells made from a lathe profile, and tori.
  • Aim for about 15,000 shapes, with 60% scanned or CAD shapes and 40% procedural shapes. Filter out shapes with fewer than 500 or more than 25,000 occupied voxels.
  • Parallel runs: Write make_dataset.py to run the solves with one process per core. Write shards that contain the following for each shape:
  • The occupancy grid, packed as bits.
  • The 64 unit-material eigenvalues.
  • For each surface voxel, the normal, which comes from the occupancy gradient, and the 64-vector of u_i(p) · n, evaluated at the voxel center by averaging the nodes.
  • Split: Use 90% for training, 5% for validation, and 5% for testing, split by source mesh.

Acceptance criteria:

  • At least 12,000 valid shapes.
  • A histogram of the first-mode frequency, results/m2_f1_hist.png.
  • The dataset generation stays within the budget.

M3: The surrogate model

Architecture:

  • Input: The 32³ occupancy grid. Add the normalized coordinate channels, which gives 4 channels in total.
  • Body: A 3D U-Net with widths 32, 64, 128, and 256, and GroupNorm with SiLU. Keep the total size to 10M parameters or less, so that the model runs well in ONNX Runtime Web.
  • Frequency head: Global pooling of the bottleneck, followed by an MLP that outputs 64 values of log λ. Enforce the ascending order by predicting log λ_1 plus softplus increments.
  • Mode-gain head: A per-voxel output of 64 values, used only at surface voxels. This predicts |u_i(p) · n|.

Losses:

  • Frequency loss: L1 on log f, sorted.
  • Sign, and ambiguity from mode pairs: the mode shapes are defined only up to sign, and nearly degenerate modes are defined only up to rotation within their group. So supervise the gains as follows:
  • Predict magnitudes only.
  • Add a spectral loss that is invariant to how modes are permuted within groups. For a random sample of surface points, render the log-magnitude spectrum of the impulse response, Σ_i a_i e^{-d_i t} sin(ω_i t), with a fixed reference damping. Render it analytically in the frequency domain as a sum of Lorentzians on 256 log-spaced bins from 20 Hz to 20 kHz, at a reference material. Take the L1 distance to the same rendering of the ground truth.
  • Use (total) = L_freq + λ_s L_spec + λ_g L_gain. Here L_gain is an L1 loss on the per-mode magnitudes, computed only for modes whose neighbors are more than 2% apart in frequency.

Evaluation (eval.py):

  • The median relative error of f_1 through f_10.
  • The spectral loss on the test set.
  • Inference time.
  • Listening pairs: 20 test shapes, each tapped at 3 points, rendered as WAV files for both the ground truth and the prediction, at the steel and glass presets. Save them to results/m3_ab/.

Acceptance criteria:

  • The median relative frequency error is 5% or less for modes 1-10, and 10% or less for modes 11-64.
  • The spectral loss on the test set is within 1.5 times the training value.
  • A single inference on the pod GPU takes 20 ms or less.

M4: Export

Tasks:

  • Export the model to ONNX, with opset 17 or later and a static shape of (1, 4, 32, 32, 32).
  • Check the model in onnxruntime on the CPU against PyTorch. The maximum relative error on the outputs must be 1e-3 or less.
  • Save 10 golden input and output pairs for the browser tests.
  • Save a quantized or fp16 variant only if it keeps the M3 metrics within 1 point. Otherwise, ship fp32.
  • For the 8-12 bundled demo objects, precompute the true FEM modes too. The web app uses them for the A/B toggle.

Acceptance criteria:

  • The ONNX outputs match PyTorch, and the golden files exist.

M5: The WASM modal engine

engine/modal_bank.cpp:

  • Use a structure-of-arrays layout for up to 4,096 active modes. Store r·2cos θ, -r², and two state values per mode.
  • Process 128-sample blocks with WASM SIMD, 4 modes per vector.
  • Handle excitation events with the fields {object_id, sample_offset, amplitude[64]}. The contact pulse is a raised-cosine force pulse whose width, from 0.1 to 2 ms, comes from the material hardness. This follows a Hertz contact model: harder materials give shorter pulses and more high-frequency content.
  • Cull modes whose envelope falls below -80 dBFS, and reuse their slots.
  • Output stereo. Each object has a gain and a pan, derived from the listener-relative azimuth and distance, which the main thread updates each frame. Use equal-power panning, and attenuate with distance as 1/max(d, d0).
  • Apply a final soft limiter.
  • Allocate no memory inside process().

The worklet:

  • Load the WASM module inside the AudioWorkletProcessor.
  • Pass events through a lock-free single-producer single-consumer ring buffer in a SharedArrayBuffer when the page is cross-origin isolated. Otherwise, fall back to port.postMessage.

Acceptance criteria:

  • The native unit tests show that a single mode's output matches the analytic damped sinusoid within -60 dB error.
  • Load test in Chrome on an Apple M-series Mac: 2,000 active modes run without underruns for 60 seconds. Log the process() time per block, and keep the 99th percentile at 50% of the block deadline or less.

M6: Browser voxelizer and surrogate

Tasks:

  • Write voxelize.ts as an exact port of voxelize.py. For 50 meshes, the occupancy must match Python's exactly. Allow up to 0.5% of voxels to differ, and record the causes, such as floating-point parity edge cases.
  • Write surrogate.ts. It loads the ONNX model with the WebGPU execution provider and falls back to WASM. It runs one inference per unique shape and caches the result by mesh hash.
  • Map a hit point in the object's local frame to the nearest surface voxel. Interpolate the gains from its 2×2×2 neighborhood, to avoid clicks between voxels.
  • Apply the material and scale analytically, then send the modes to the worklet.

Acceptance criteria:

  • The browser outputs match the golden files within 1e-3.
  • From selecting a .glb file to hearing its first tap takes 1.5 seconds or less on an M-series Mac, for a mesh with 100,000 triangles or fewer.

M7: The demo app

Tasks:

  • Build a three.js scene: a room with a shelf and 8-12 bundled objects. Include a mug, a wine glass, a bell, a wrench or bar, a bowl, a plate, a vase, and one odd sculpture.
  • Physics: build Rapier rigid bodies from convex decompositions, or from convex hulls for speed. Enable contact-force events. When a force crosses a threshold, send an excitation at the contact point, with an impulse equal to the force times dt. Rate-limit excitations to one per object pair per 15 ms, to avoid buzzing from resting contacts.
  • Interactions:
  • Click to tap. The hammer impulse depends on how long the click is held.
  • Drag to throw.
  • A material dropdown, and a scale slider from 0.2× to 5×.
  • Chaos: Spawn 30 random objects from the shelf above the floor.
  • Upload .glb: Voxelize the model, run the surrogate, add it to the scene, and let the user tap it.
  • A/B: For bundled objects, switch between the neural modes and the true FEM modes.
  • A live spectrum analyzer that uses an AnalyserNode.
  • Fallback: if the browser lacks WebGPU, use the WASM execution provider. If it lacks AudioWorklet or WASM SIMD, show a message and a video.
  • Mobile layout: taps work with touch, and the page handles the requirement to resume the AudioContext after a user gesture.

Acceptance criteria:

  • A Playwright smoke test loads the page, taps an object, and checks that the worklet receives an excitation and produces output above -60 dBFS.
  • Chaos mode with 30 objects runs at 50 fps or more with no audio underruns, on an M-series Mac.
  • The page works in Chrome and Safari on macOS.

M8: Deploy and write-up

Tasks:

  • Write deploy/Dockerfile: nginx:alpine serving web/dist, with the correct wasm MIME type and these headers:
  • Cross-Origin-Opener-Policy: same-origin
  • Cross-Origin-Embedder-Policy: require-corp
  • Ask the user before deploying. After the user approves, create a new Dokploy application on vm.ifkash.dev, with a domain such as clatter.ifkash.dev that the user confirms. The user might need to add the DNS record.
  • Write README.md. It explains what the project is, how to reproduce each milestone with one command per milestone, the accuracy and speed tables, links to the A/B audio, and credits and licenses for NeuralSound, the mesh sources, and Rapier.
  • Write post/draft.md, a blog post of 1,200-1,800 words. Structure it as follows:
  • The hook: "No sound files."
  • Why objects ring: modes, in plain words.
  • FEM on voxels.
  • The surrogate, and the mode-pair problem with the spectral-loss fix.
  • Why material is free.
  • The WASM audio engine and its real-time budget.
  • Physics-driven clatter.
  • Limitations: no radiation model, linear elasticity only, 32³ resolution, and no rolling or sliding sounds.
  • Stretch goals, only if time remains:
  • RealImpact comparison: compare predicted and recorded frequencies for its objects, with a material fit.
  • FFAT radiation maps.
  • Rolling and sliding sounds.
  • Custom WGSL 3D-convolution kernels to replace ONNX Runtime Web.
  • Ask the user before you push anything. After the user approves, push the repo to weights-and-wires/clatter and upload the model to the Hugging Face Hub under weights-and-wires/clatter-modal-surrogate.
  • Terminate every pod, and write the final spend to NOTES.md.

Acceptance criteria:

  • The deployed URL loads, and taps make sound in Chrome.
  • No pods are left running, and the total spend is less than $30.

Final report to the user

When you finish, report the following:

  • The FEM validation errors, from M1.
  • The dataset size, from M2.
  • The surrogate's frequency errors and inference time, from M3.
  • The engine's capacity in modes, and its 99th-percentile block time, from M5.
  • The time from upload to first tap, from M6.
  • The URL, if the site is deployed.
  • The total spend.
  • Anything that you skipped or that failed.