clatter
Hear what any 3D object sounds like, live in your browser
Part 1: For you
The idea in one line
Drop any 3D object into a browser scene, hit it, drop it, or throw it, and hear the sound it would really make. The sound comes from a neural network that learned the physics of vibration. Nobody recorded it.
Why it's cool
- It's physics you can hear. When you hit a mug, it rings at a few exact pitches called modes. The pitches depend on its shape and material. Computing them normally takes a slow physics solve, called FEM, for every object.
- The neural network skips the solve. You give it a shape, and it predicts the modes in milliseconds. So you can upload any 3D model and hear it instantly.
- Nobody has shipped this in a browser. Papers such as NeuralSound showed the idea works, but only in videos. The research agents found no live demo where you can drop in your own object.
- It's physical AI for your ears. It covers simulation, a neural shortcut for a physics solver, and a physics engine that drives the sound.
- It's low-level. You write a real-time audio engine in C++, compiled to WebAssembly, that rings thousands of modes at once without a single glitch.
- It's not ASR or TTS, and it doesn't build on any of your past projects.
What the demo looks like
- A 3D room with a shelf of objects: a mug, a bell, a wine glass, a wrench, and a weird-shaped sculpture.
- Click anywhere on an object to tap it. Tapping a different spot gives a different sound, just like in real life.
- Change the material: glass, steel, wood, or ceramic. The sound changes instantly.
- Make it bigger or smaller. Bigger objects sound lower.
- Press "Chaos" to drop 30 objects on the floor and hear them all clatter, with the sound coming from the direction of each object.
- Upload your own 3D model (
.glb). In about a second, it has a voice. - An A/B toggle: hear the neural version against the slow "true" physics version, to show how close it is.
How it works
- Make the answers. On a GPU pod, run the slow physics solve on about 15,000 random 3D shapes. This gives you the true modes of each shape.
- Train the shortcut. A small 3D network looks at a shape, as 32×32×32 blocks, and predicts its modes: the pitches, and how loud each pitch is at each spot on the surface.
- Material is free. Pitch scales with a simple formula for stiffness and density. So you train on one material and get every material for free.
- Build the sound engine. Each mode is a tiny ringing filter. The engine runs thousands of them in C++ compiled to WebAssembly, inside the browser's audio thread.
- Connect the physics. A physics engine (Rapier) detects every hit: where it happened and how hard. That hit rings the right modes.
- Ship it as a static web page on
vm.ifkash.dev.
Weekend plan
| When | What | Done when |
|---|---|---|
| Saturday morning | Write the physics solver, check it on a few shapes | A metal bar rings at the pitch the textbook predicts |
| Saturday afternoon | Generate 15,000 shapes and their modes on the pod (runs by itself) | The dataset is done |
| Saturday night | Train the network (runs by itself) | Predicted pitches match the true ones |
| Sunday morning | Build the WebAssembly sound engine and the physics scene | Tapping an object makes a sound |
| Sunday afternoon | Upload support, Chaos mode, deploy, record the video, write the post | The link works |
Cost
- GPU: 1× RTX 5090 on RunPod, about $0.69 per hour. Pick a pod with lots of CPU cores, because the physics solve runs on the CPU.
- Time on the pod: about 10-15 hours.
- Total: about $8-15. The plan sets a hard stop at $30.
- Hosting: free. It's a static page, and it keeps working after the pod is gone.
What you have at the end
- A live link where anyone can hear their own 3D models.
- A chart: neural pitches against true pitches, and how fast each one is.
- A video of 30 objects clattering, with no recorded sound anywhere.
- A blog post: "No sound files: every sound here is predicted from physics."
What might go wrong
| Problem | What to do |
|---|---|
| The physics solver takes all Saturday | Use NeuralSound's open-source code to make the data. |
| Symmetric objects confuse the network, because some modes come in pairs | Train on "how the sound spectrum looks at each spot" instead of on single modes. The plan explains how. |
| The audio glitches when too many objects ring | Turn off quiet modes early, and cap how many modes ring at once. |
| Uploaded models are broken, for example with holes or no inside | Fill them into solid blocks before the network sees them. |
Other ideas the research turned up
- "Mm-hm." A browser tab that listens like a person does: it says "mm-hm" at the right moments and knows when you've finished talking. It's small, runs in the browser, and costs about $10. It's the best pick if you want speech specifically.
- A playable piano world model. Play keys in the browser, and a model predicts the piano sound. Record a phrase, and the model works out which keys made it. It's based on a July 2026 paper, Music-JEPA.
- Walk into a photo. Take a photo of a room, then move around inside it and hear your voice echo the way it would there.
Reading, if you want it
- NeuralSound and its code: the neural modal solver this project builds on.
- Differentiable modal resonators: the ringing-filter idea.
- RealImpact: real recordings of objects being hit, which you can use as a stretch-goal reality check.
Part 2: For the coding agent
Mission
Build clatter, a browser app that synthesizes physically based impact sounds for arbitrary 3D
objects in real time:
- Surrogate model: A neural network, trained on FEM modal analysis ground truth, maps a voxelized shape to its modal parameters. It runs in the browser.
- Material: Material and scale are applied analytically.
- Audio engine: A C++ to WASM SIMD modal resonator bank in an
AudioWorkletrenders audio, excited by contact events from a Rapier physics simulation, with per-object spatial panning.
The final artifact is a static website with no server.
Work through milestones M0-M8 in order. Each milestone has acceptance criteria. Don't start a
milestone until the previous one passes. After each milestone, commit your work and write a
short entry in NOTES.md with the results and numbers.
Hard constraints
- Secrets: Read
RUNPOD_API_KEY,HF_TOKEN, andWANDB_API_KEYfrom environment variables only. Never write a secret into any file, log, commit, or echoed command. Commit a.env.examplethat has placeholder values only, and add.envto.gitignore. - Budget: The hard cap is $30 of RunPod spend. Every pod runs a watchdog that stops the pod
after
MAX_POD_HOURShours. The default is6. - Compute: Run large-scale data generation and training only on RunPod. The local Mac runs the web build, the browser tests, the audio engine tests, and small solver checks.
- Outward actions: Ask the user before you do any of the following: create the GitHub
repository, push to the Hugging Face Hub, deploy to a VM, change DNS, or post anything
publicly. Use the
kashifulhaqueGitHub account (gh auth switch -u kashifulhaque). - Shared VM:
vm.ifkash.devruns other production apps under Dokploy, with Traefik on ports 80 and 443. Don't stop, restart, reconfigure, or remove any existing container, service, or proxy rule. Add clatter only as a new Dokploy application. - Pod cleanup: When a pod isn't running a job, stop it. At the end of the project, terminate every pod that you created and report the total spend.
- Licenses: Check the license of every mesh dataset and of NeuralSound's code before you
use them. Record the licenses in
README.md. Bundle only meshes whose license allows redistribution in the demo.
Physics background
Implement the physics exactly as described here:
- Linear modal analysis. Discretize the solid with finite elements. Solve the generalized
eigenproblem
K u = λ M u, whereKis the stiffness matrix andMis the mass matrix. The free-free boundary gives the object no fixed points. Discard the 6 rigid-body modes, whereλ ≈ 0, and keep the nextK_modes = 64modes. - Frequencies. The mode frequency is
f_i = sqrt(λ_i) / (2π). - Material scaling. For a linear isotropic material with Poisson ratio
ν, the modes depend on Young's modulusE, densityρ, and length scalesonly through a scalar:f_i(E, ρ, s) = f_i^unit * sqrt(E/ρ) / s. The mode shapes don't change. Train at unitE/ρand unit scale, and fixνat a value such as 0.3. Treatνper material as an approximation, and record this in the limitations. - Damping. Use Rayleigh damping,
C = α M + β K. This gives a per-mode decay rate ofd_i = (α + β ω_i²) / 2, whereω_i = 2π f_i. Each material preset setsαandβ. - Excitation. An impulse
Jat a surface pointpwith normalnexcites modeiwith amplitudea_i = (u_i(p) · n) * J / (ω_i), up to a global constant. Use the mass-normalized mode shapesu_i. - Rendering. Each mode is a damped sinusoid:
y(t) = Σ a_i e^{-d_i t} sin(ω_i t). Implement each mode as a two-pole resonator:y[n] = 2 r cos(θ) y[n-1] - r² y[n-2] + g x[n], wherer = e^{-d/fs}andθ = ω/fs. - Hearing range. Drop modes above 20 kHz or below 20 Hz after material scaling.
- Acoustic radiation. Don't model radiation in the core scope. Treat it as a stretch goal (FFAT maps, as in NeuralSound).
Material presets
Start from published values, and tune α and β by ear in M5. The table lists the starting
values:
| Material | E (Pa) | ρ (kg/m³) | ν | α | β |
|---|---|---|---|---|---|
| Steel | 2.0e11 | 7850 | 0.29 | 5 | 3e-8 |
| Glass | 6.2e10 | 2600 | 0.20 | 1 | 1e-8 |
| Ceramic | 7.2e10 | 2700 | 0.19 | 6 | 1e-7 |
| Wood | 1.1e10 | 750 | 0.25 | 60 | 2e-6 |
| Plastic | 1.4e9 | 1070 | 0.35 | 30 | 1e-6 |
Tech stack
Pin every version. The stack is as follows:
- Python: Python 3.11, PyTorch 2.x,
numpy,scipy,trimesh, andwandb, managed withuv. For voxelization, usetrimeshor a custom solid voxelizer. Usecupyonly if GPU LOBPCG is needed. - Browser:
- TypeScript and
vite, withthree.jsfor rendering. @dimforge/rapier3d-compatfor physics.onnxruntime-webwith the WebGPU execution provider, falling back to WASM, for the surrogate.- Audio engine: C++17 compiled with Emscripten to WASM with SIMD (
-msimd128), run inside anAudioWorklet. Share memory through aSharedArrayBufferring buffer if cross-origin isolation is available; otherwise useMessagePort. - Tests: Playwright with Chromium, and native C++ unit tests for the engine, using the same source code.
- Pods: the
runpodPython SDK, withrunpodctlinside pods.
Repository layout
Create the following layout:
clatter/
pyproject.toml .env.example .gitignore README.md NOTES.md
infra/pod.py infra/watchdog.sh infra/bootstrap.sh
fem/
voxelize.py # mesh -> solid 32^3 occupancy (+ scale normalization)
hexfem.py # voxel hex8 linear-elastic K, M assembly (sparse)
modal.py # eigensolve, rigid-mode removal, mass-normalization
validate.py # analytic checks (bar, plate, cube)
make_dataset.py # parallel over meshes -> shards (npz/zarr)
train/
model.py # 3D CNN / UNet surrogate
losses.py # frequency loss + permutation-invariant spectral loss
train.py eval.py
export_onnx.py # + golden I/O for browser tests
engine/ # C++ modal resonator bank (native + WASM)
modal_bank.h modal_bank.cpp test_modal_bank.cpp build_wasm.sh
web/
src/voxelize.ts # exact port of voxelize.py
src/surrogate.ts # ORT-web wrapper, feature -> modes
src/worklet.ts # AudioWorklet processor hosting the WASM bank
src/physics.ts # Rapier world, contact events -> excitations
src/scene.ts src/ui.ts src/main.ts index.html
public/objects/ # bundled demo meshes + precomputed FEM "truth" modes
tests/
deploy/Dockerfile # nginx:alpine, COOP/COEP headers for SharedArrayBuffer
results/ post/draft.md
M0: Infrastructure
Tasks:
- Write
infra/pod.py. It uses therunpodSDK and readsRUNPOD_API_KEYfrom the environment. It supports the following subcommands: -create: Creates a pod. The GPU preference order is RTX 5090, then RTX 4090, then RTX PRO 4500. Resolve the GPU type IDs at run time by querying the available GPU types. Prefer a pod with the most vCPUs available, 16 or more, because the eigensolves run on the CPU. Use an official RunPod PyTorch image with CUDA 12.8 or later. Set a 100 GB volume at/workspaceand expose SSH. -status,stop, andterminate. -ssh-info: Prints the SSH command. -cost: Prints the uptime and spend for every pod whose name has the prefixclatter-. - Write
infra/watchdog.sh. It sleeps forMAX_POD_HOURS, then runsrunpodctl stop pod $RUNPOD_POD_ID. - Write
infra/bootstrap.sh. It clones the repo, runsuv sync, and starts the watchdog.
Acceptance criteria:
- A pod comes up, and
bootstrap.shfinishes without errors. - A test run with
MAX_POD_HOURS=0.05stops the pod within 5 minutes.
M1: FEM solver and validation
Tasks:
- In
voxelize.py, turn a mesh into a solid voxel grid: - Normalize the mesh so that its bounding-box diagonal equals 1.
- Center it in a 32³ grid.
- Fill the inside with parity ray casting, or with a winding-number test for meshes that aren't watertight.
- Remove floating islands, and keep the largest 6-connected component.
- Return the occupancy and the physical voxel size
h. - In
hexfem.py, assemble the stiffness matrixKand the consistent mass matrixMas sparse CSR matrices: - Use trilinear hexahedral (hex8) elements, one per occupied voxel.
- Share nodes between neighboring voxels.
- Use 2×2×2 Gauss quadrature for the element matrices.
- Precompute a single reference element stiffness matrix for unit
h, and scale it byh, because every voxel is the same cube. - Use unit
Eandρ, withν = 0.3. - In
modal.py, find the smallest 70 eigenpairs withscipy.sparse.linalg.eigsh(K, k=70, M=M, sigma=-small, which='LM')in shift-invert mode. - Drop the 6 modes where
λ < 1e-6 * λ_7. - Keep 64 modes, and mass-normalize them so that
uᵀ M u = 1. - Record the number of nodes and the solve time.
- If NeuralSound's repo,
hellojxt/NeuralSound, provides a working voxel FEM and dataset pipeline, you may use it instead. Record the choice, and still run the validation. - In
validate.py, check the solver against analytic results: - The first bending frequencies of a free-free slender bar, at 32×2×2 voxels, compared with Euler-Bernoulli beam theory.
- The ratio of the first two frequencies of a thin square plate.
- Frequencies scaling as
1/swhen the voxel size changes.
Acceptance criteria:
- For the bar, the first frequency is within 5% of Euler-Bernoulli theory, or within 10% for
the voxel approximation if the error is explained in
NOTES.md. - The scaling test is exact to 1e-6.
- A single 32³ solve finishes in 10 seconds or less on one CPU core, for a typical object with about 10,000 occupied voxels.
M2: Dataset
Tasks:
- Meshes: Pick a source with a license that allows use, such as Objaverse objects under CC-BY, Thingi10K, or the ABC dataset, which NeuralSound uses. Mix in procedural primitives so that thin shells and bars are well represented: cylinders, hollow cups, bowls, plates, bars, bells made from a lathe profile, and tori.
- Aim for about 15,000 shapes, with 60% scanned or CAD shapes and 40% procedural shapes. Filter out shapes with fewer than 500 or more than 25,000 occupied voxels.
- Parallel runs: Write
make_dataset.pyto run the solves with one process per core. Write shards that contain the following for each shape: - The occupancy grid, packed as bits.
- The 64 unit-material eigenvalues.
- For each surface voxel, the normal, which comes from the occupancy gradient, and the
64-vector of
u_i(p) · n, evaluated at the voxel center by averaging the nodes. - Split: Use 90% for training, 5% for validation, and 5% for testing, split by source mesh.
Acceptance criteria:
- At least 12,000 valid shapes.
- A histogram of the first-mode frequency,
results/m2_f1_hist.png. - The dataset generation stays within the budget.
M3: The surrogate model
Architecture:
- Input: The 32³ occupancy grid. Add the normalized coordinate channels, which gives 4 channels in total.
- Body: A 3D U-Net with widths 32, 64, 128, and 256, and GroupNorm with SiLU. Keep the total size to 10M parameters or less, so that the model runs well in ONNX Runtime Web.
- Frequency head: Global pooling of the bottleneck, followed by an MLP that outputs 64
values of
log λ. Enforce the ascending order by predictinglog λ_1plus softplus increments. - Mode-gain head: A per-voxel output of 64 values, used only at surface voxels. This
predicts
|u_i(p) · n|.
Losses:
- Frequency loss: L1 on
log f, sorted. - Sign, and ambiguity from mode pairs: the mode shapes are defined only up to sign, and nearly degenerate modes are defined only up to rotation within their group. So supervise the gains as follows:
- Predict magnitudes only.
- Add a spectral loss that is invariant to how modes are permuted within groups. For a
random sample of surface points, render the log-magnitude spectrum of the impulse response,
Σ_i a_i e^{-d_i t} sin(ω_i t), with a fixed reference damping. Render it analytically in the frequency domain as a sum of Lorentzians on 256 log-spaced bins from 20 Hz to 20 kHz, at a reference material. Take the L1 distance to the same rendering of the ground truth. - Use
(total) = L_freq + λ_s L_spec + λ_g L_gain. HereL_gainis an L1 loss on the per-mode magnitudes, computed only for modes whose neighbors are more than 2% apart in frequency.
Evaluation (eval.py):
- The median relative error of
f_1throughf_10. - The spectral loss on the test set.
- Inference time.
- Listening pairs: 20 test shapes, each tapped at 3 points, rendered as WAV files for both the
ground truth and the prediction, at the steel and glass presets. Save them to
results/m3_ab/.
Acceptance criteria:
- The median relative frequency error is 5% or less for modes 1-10, and 10% or less for modes 11-64.
- The spectral loss on the test set is within 1.5 times the training value.
- A single inference on the pod GPU takes 20 ms or less.
M4: Export
Tasks:
- Export the model to ONNX, with opset 17 or later and a static shape of
(1, 4, 32, 32, 32). - Check the model in
onnxruntimeon the CPU against PyTorch. The maximum relative error on the outputs must be 1e-3 or less. - Save 10 golden input and output pairs for the browser tests.
- Save a quantized or fp16 variant only if it keeps the M3 metrics within 1 point. Otherwise, ship fp32.
- For the 8-12 bundled demo objects, precompute the true FEM modes too. The web app uses them for the A/B toggle.
Acceptance criteria:
- The ONNX outputs match PyTorch, and the golden files exist.
M5: The WASM modal engine
engine/modal_bank.cpp:
- Use a structure-of-arrays layout for up to 4,096 active modes. Store
r·2cos θ,-r², and two state values per mode. - Process 128-sample blocks with WASM SIMD, 4 modes per vector.
- Handle excitation events with the fields
{object_id, sample_offset, amplitude[64]}. The contact pulse is a raised-cosine force pulse whose width, from 0.1 to 2 ms, comes from the material hardness. This follows a Hertz contact model: harder materials give shorter pulses and more high-frequency content. - Cull modes whose envelope falls below -80 dBFS, and reuse their slots.
- Output stereo. Each object has a gain and a pan, derived from the listener-relative azimuth
and distance, which the main thread updates each frame. Use equal-power panning, and
attenuate with distance as
1/max(d, d0). - Apply a final soft limiter.
- Allocate no memory inside
process().
The worklet:
- Load the WASM module inside the
AudioWorkletProcessor. - Pass events through a lock-free single-producer single-consumer ring buffer in a
SharedArrayBufferwhen the page is cross-origin isolated. Otherwise, fall back toport.postMessage.
Acceptance criteria:
- The native unit tests show that a single mode's output matches the analytic damped sinusoid within -60 dB error.
- Load test in Chrome on an Apple M-series Mac: 2,000 active modes run without underruns for
60 seconds. Log the
process()time per block, and keep the 99th percentile at 50% of the block deadline or less.
M6: Browser voxelizer and surrogate
Tasks:
- Write
voxelize.tsas an exact port ofvoxelize.py. For 50 meshes, the occupancy must match Python's exactly. Allow up to 0.5% of voxels to differ, and record the causes, such as floating-point parity edge cases. - Write
surrogate.ts. It loads the ONNX model with the WebGPU execution provider and falls back to WASM. It runs one inference per unique shape and caches the result by mesh hash. - Map a hit point in the object's local frame to the nearest surface voxel. Interpolate the gains from its 2×2×2 neighborhood, to avoid clicks between voxels.
- Apply the material and scale analytically, then send the modes to the worklet.
Acceptance criteria:
- The browser outputs match the golden files within 1e-3.
- From selecting a
.glbfile to hearing its first tap takes 1.5 seconds or less on an M-series Mac, for a mesh with 100,000 triangles or fewer.
M7: The demo app
Tasks:
- Build a three.js scene: a room with a shelf and 8-12 bundled objects. Include a mug, a wine glass, a bell, a wrench or bar, a bowl, a plate, a vase, and one odd sculpture.
- Physics: build Rapier rigid bodies from convex decompositions, or from convex hulls for
speed. Enable contact-force events. When a force crosses a threshold, send an excitation at
the contact point, with an impulse equal to the force times
dt. Rate-limit excitations to one per object pair per 15 ms, to avoid buzzing from resting contacts. - Interactions:
- Click to tap. The hammer impulse depends on how long the click is held.
- Drag to throw.
- A material dropdown, and a scale slider from 0.2× to 5×.
- Chaos: Spawn 30 random objects from the shelf above the floor.
- Upload .glb: Voxelize the model, run the surrogate, add it to the scene, and let the user tap it.
- A/B: For bundled objects, switch between the neural modes and the true FEM modes.
- A live spectrum analyzer that uses an
AnalyserNode. - Fallback: if the browser lacks WebGPU, use the WASM execution provider. If it lacks
AudioWorkletor WASM SIMD, show a message and a video. - Mobile layout: taps work with touch, and the page handles the requirement to resume the
AudioContextafter a user gesture.
Acceptance criteria:
- A Playwright smoke test loads the page, taps an object, and checks that the worklet receives an excitation and produces output above -60 dBFS.
- Chaos mode with 30 objects runs at 50 fps or more with no audio underruns, on an M-series Mac.
- The page works in Chrome and Safari on macOS.
M8: Deploy and write-up
Tasks:
- Write
deploy/Dockerfile:nginx:alpineservingweb/dist, with the correctwasmMIME type and these headers: Cross-Origin-Opener-Policy: same-originCross-Origin-Embedder-Policy: require-corp- Ask the user before deploying. After the user approves, create a new Dokploy application on
vm.ifkash.dev, with a domain such asclatter.ifkash.devthat the user confirms. The user might need to add the DNS record. - Write
README.md. It explains what the project is, how to reproduce each milestone with one command per milestone, the accuracy and speed tables, links to the A/B audio, and credits and licenses for NeuralSound, the mesh sources, and Rapier. - Write
post/draft.md, a blog post of 1,200-1,800 words. Structure it as follows: - The hook: "No sound files."
- Why objects ring: modes, in plain words.
- FEM on voxels.
- The surrogate, and the mode-pair problem with the spectral-loss fix.
- Why material is free.
- The WASM audio engine and its real-time budget.
- Physics-driven clatter.
- Limitations: no radiation model, linear elasticity only, 32³ resolution, and no rolling or sliding sounds.
- Stretch goals, only if time remains:
- RealImpact comparison: compare predicted and recorded frequencies for its objects, with a material fit.
- FFAT radiation maps.
- Rolling and sliding sounds.
- Custom WGSL 3D-convolution kernels to replace ONNX Runtime Web.
- Ask the user before you push anything. After the user approves, push the repo to
weights-and-wires/clatterand upload the model to the Hugging Face Hub underweights-and-wires/clatter-modal-surrogate. - Terminate every pod, and write the final spend to
NOTES.md.
Acceptance criteria:
- The deployed URL loads, and taps make sound in Chrome.
- No pods are left running, and the total spend is less than $30.
Final report to the user
When you finish, report the following:
- The FEM validation errors, from M1.
- The dataset size, from M2.
- The surrogate's frequency errors and inference time, from M3.
- The engine's capacity in modes, and its 99th-percentile block time, from M5.
- The time from upload to first tap, from M6.
- The URL, if the site is deployed.
- The total spend.
- Anything that you skipped or that failed.