lfm2-350m-webgpu · benchmarks
Community results of the standard suite (LFM2.5-350M, GPTQ-int4 269 MB, hand-written WGSL kernels). Click a row for its per-kernel GPU profile. · chat · 64-chat wall · run it on your GPU
| device | decode tok/s | best | prefill tok/s | batch B=64 | GPU ms/tok | date |
|---|
decode = single-stream greedy, median of 5 (chunk 4) · prefill = 256 tokens, median of 3 · batch = 64 parallel chats, total throughput · GPU ms/tok = timestamp-query sum per decode step · «fallback» = the device has no WGSL subgroups (e.g. iOS Safari) and runs the shared-memory reduction path.
▶ run the bench on your GPU