lfm2-350m-webgpu · benchmark

Runs the standard performance suite on your GPU and produces a shareable .log file — load time, prefill/decode throughput, per-kernel GPU time, batched throughput. Weights: 269 MB int4 (downloaded once, cached by the browser). Needs WebGPU with shader-f16 (subgroups optional — iOS Safari runs a fallback path).  ·  chat  ·  64-chat wall  ·  community results

Reference — Mac mini M4, 10-core GPU, Chrome

decodeprefill (256)batch B=64weights
q4-GPTQ~390 tok/s~2330 tok/s~1640 tok/s269 MB

Run

idle — press Run bench

Results