MicroLLM Lab is a browser-based workbench for running quantized small language models — think 135M parameters, Q4 precision — directly in your browser tab. No server, no API key, no data leaving your machine. The tool measures two things: sustained decode speed in tokens per second, and pass rates on objective tests you define or select from built-in suites. Every number stays local. The interface is organized around a tight loop: pick a model, run a benchmark suite, read the results. Benchmark suites are written in JavaScript and eval()'d in the page's origin, with each check running against the model's raw decoded text. This means you can write regex-based or exact-token checks — objective grading, not vibes. The tool is explicitly comfortable with failure: a 135M model is allowed to fail, and measuring that failure is the point. Speed measurement captures tokens per second during sustained decode, plus wall-clock time for the full suite. Accuracy is reported as a pass rate on the objective tests. Charts render per-model using the latest suite run, so you can compare across models on the same hardware without trusting anyone else's numbers. The emphasis on local-only measurement is deliberate — this is a tool for people who want to know what actually runs on their specific device. The custom benchmark editor is the most interesting surface. You write JavaScript that defines checks, the tool evals it, and each check executes against the model output. This is closer to a unit-test harness for language models than a traditional leaderboard. It rewards users who can think precisely about what 'correct' means for a given task at this scale. The tool identifies your device and hardware automatically and logs it alongside results, giving you a hardware-tagged performance profile. This is useful for anyone trying to answer the practical question: can this model run usefully on this laptop, this phone, this tablet? The answer is often no — but knowing that precisely, with numbers, is more useful than guessing. MicroLLM Lab occupies a specific niche: it's for tinkerers, educators, and edge-deployment engineers who want ground-truth measurements of what tiny quantized models can actually do on consumer hardware. It does not pretend these models are good. It gives you the tools to find out exactly how they fail, and how fast they fail, on your machine.