MicroLLM Lab: Run 7 Tiny LLMs in Your Browser

Original: MicroLLM Lab – Try 7 tiny LLM's in the browser

Why This Matters

On-device SLM inference with zero server cost signals a real shift in how lightweight AI tasks can be deployed at the edge.

MicroLLM Lab is a browser-based tool that lets users load, chat with, and benchmark up to 7 small language models (25M–360M parameters) using WebGPU and Q4 quantization — entirely on-device, with no server, no account, and no data leaving the user's machine.

MicroLLM Lab, published by stateofutopia.com, is an experimental browser environment for running Small Language Models (SLMs) directly on consumer hardware via the WebGPU API. The tool supports seven models ranging from 25M to 360M parameters, each quantized to 4-bit (Q4), which shrinks memory footprint by roughly 75% — allowing 100M+ parameter models to fit within 50–84 MB of browser memory. Models cache to the browser's IndexedDB rather than writing to disk.

The interface has three main sections: a Chat panel for live inference, a Benchmarks tab for running objective speed and accuracy tests (regex/exact token matching, not writing quality), and a Compare & Certificate view that generates a shareable performance card. The lab claims sub-10ms time-to-first-token is achievable, with sustained decode speeds recorded per device.

The stated use cases focus on edge computing: classifying queries, filtering spam, extracting intent, and routing decisions to heavier cloud models only when necessary — eliminating API costs for lightweight tasks. WebGPU leverages Apple Silicon Metal, DirectX 12, or Vulkan depending on the platform, keeping all compute local. A custom JavaScript eval panel also allows users to write their own benchmark checks against model output.

Source

stateofutopia.com — Read original →