lfm2-350m-webgpu · benchmarks

Community results of the standard suite (LFM2.5-350M, GPTQ-int4 269 MB, hand-written WGSL kernels). Click a row for its per-kernel GPU profile.  ·  chat  ·  64-chat wall  ·  run it on your GPU

devicedecode tok/sbestprefill tok/s batch B=64GPU ms/tokdate

decode = single-stream greedy, median of 5 (chunk 4) · prefill = 256 tokens, median of 3 · batch = 64 parallel chats, total throughput · GPU ms/tok = timestamp-query sum per decode step · «fallback» = the device has no WGSL subgroups (e.g. iOS Safari) and runs the shared-memory reduction path.

▶ run the bench on your GPU