lfm2-350m-webgpu · benchmark
Runs the standard performance suite on your GPU and produces a shareable .log file — load time, prefill/decode throughput, per-kernel GPU time, batched throughput. Weights: 269 MB int4 (downloaded once, cached by the browser). Needs WebGPU with shader-f16 (subgroups optional — iOS Safari runs a fallback path). · chat · 64-chat wall · community results
Reference — Mac mini M4, 10-core GPU, Chrome
| decode | prefill (256) | batch B=64 | weights | |
|---|---|---|---|---|
| q4-GPTQ | ~390 tok/s | ~2330 tok/s | ~1640 tok/s | 269 MB |
Run
idle — press Run bench