AI NewsModels & agentsAnnouncement

MicroLLM Lab runs seven tiny language models in the browser on WebGPU, and reached 225 Hacker News points on the day it launched

MicroLLM Lab is a browser page that loads and runs seven small language models between 25 million and 360 million parameters entirely on the local GPU through WebGPU, with a benchmark suite that scores speed and accuracy per model.

AI News

Editorial2 min read

LinkedInX
MicroLLM Lab, a browser playground for small language models over WebGPU

Why it mattersA working team can now test whether a small on-device model is enough for a routing, tagging or filtering step, before paying for a cloud call, without setting up a local runtime or a server.

Any developer can open one browser tab and try seven small language models running entirely on their own machine, with no install and no server. That is what MicroLLM Lab, posted to Hacker News on 28 September and up to 225 points on the day it launched, offers. The models load into the browser's IndexedDB, run on WebGPU, and stay on the device.

The lab describes itself as a WebGPU playground for small language models. Its page names three parts that make it work: models compressed to four bits per weight, so a hundred-million-parameter model fits in about fifty to eighty-four megabytes of browser memory; the WebGPU standard, which runs compute shaders on the machine's own GPU; and a browser cache that keeps models inside the site's private storage so nothing writes to the download folder.

Which models are included

The page names PetitGPT and SmolLM2 among the seven models it hosts, and the source repository, at robss2020/microllm-lab, describes the lab as "inspired by yangqi0/petitgpt". Parameter counts run from 25 million to 360 million, which is a size small enough to fit in a browser tab and large enough to produce fluent output. The whole model archive is downloadable as a 589 MB zip if the reader wants to run the lab locally, and the source is on GitHub with no licence declared yet.

The benchmark is the point

The page's own copy is clear on what the lab measures: objective checks such as regex matches and exact tokens, on the machines that ran them. It does not score writing quality. The suite runs on the active model or across every loaded model, prints a live tokens-per-second speed, and produces a shareable certificate with the reader's own hardware fingerprint, peak and sustained throughput. The output is one number that can be repeated on another machine.

The lab makes two claims worth reporting as its own claims, without independent verification. It says compact models between 25 and 360 million parameters can serve edge tasks with "sub-10ms time-to-first-token" on the client GPU, and that this lets a team classify queries or filter spam before a cloud call is spent. Both figures depend on the reader's hardware. The lab's benchmark is where a team would check them on their own machine.

For a team already sending every user prompt to a cloud model, the useful question is where a small on-device model would do the job well enough on its own: a routing decision, an intent classifier or a spam filter, with no per-call cost, no server latency, and no user data leaving the machine. MicroLLM Lab makes that question answerable this week with real measurements, in a browser tab a product manager can open alongside a working developer.

Source

Primary source: MicroLLM Lab, stateofutopia.com/experiments/microllmlab.

Discussion: Show HN: MicroLLM Lab, Hacker News.

Source repository: robss2020/microllm-lab on GitHub.

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX