Strata is an open-source inference engine that runs Qwen3.8-Flash-Next — a 125-billion-parameter mixture-of-experts model — on consumer hardware. An RTX 5070 with 64 GB of system RAM generates 53-94 tokens per second depending on quantization level, fast enough that the bottleneck is reading speed, not compute. The project ships as a one-click installer for Windows and Linux, exposes an OpenAI/Anthropic-compatible API on localhost, and requires zero cloud connectivity. Nothing leaves the user's PC. The engineering trick is a heterogeneous offloading scheme the project describes with a kitchen analogy: the model's 24,576 mixture-of-experts specialists live in system RAM, the most-used few thousand stay resident on the GPU, and the CPU handles overflow in parallel. A speculative decoding layer — a small helper model guesses the next few tokens, the big model checks them in batch — delivers a 1.6-1.8x throughput multiplier for free. Long context ingestion runs at 1,000-2,650 tokens per second, meaning a 32K-token document is processed in 12-30 seconds. Hardware requirements are genuine consumer-grade: 12 GB VRAM minimum (RTX 20-series or AMD RX 6800 and up), 32 GB system RAM for the coding-specific variant, 64 GB for the full model at recommended quality. An RTX 3090 with 24 GB VRAM is projected at 100-140 tokens per second. AMD support spans RX 6000, 7000, and 9000 series on both Windows and Linux, with image input currently Linux-only on AMD. Multi-GPU configurations are supported. The model menu covers several quantization tiers — Q20 (fastest, most compressed), IQ2XS, IQ3XXS, IQ3S (best quality), plus a coding-specific half-expert variant that fits 32 GB RAM and scores 91% of the full model on SWE-bench Verified. Unsloth contributes a ~4-bit UD-IQ4XS and an experimental UD-Q4KXL that approaches full-model quality but drops to 7-8.5 tokens/second on 64 GB systems due to SSD spillover. An uncensored IQ3XXS from OrcaRouter exists but requires manual setup. The competitive context matters. Running a 125B MoE model locally at usable speeds was functionally impossible on consumer hardware 18 months ago. The combination of aggressive quantization from ISTA-DASLab and Unsloth, speculative decoding, and intelligent GPU/CPU/RAM tiering collapses the cost of inference from cloud API pricing to electricity costs. For coding agents, local deployment eliminates per-token charges, latency to remote servers, and data exfiltration risk simultaneously. The project is MIT-licensed, accepts donations, and integrates with the major AI coding assistants — Claude Code, Cursor, Codex, GitHub Copilot — via its OpenAI-compatible endpoint or MCP server. Setup can be delegated to the coding assistant itself using a setup prompt. The installer handles driver detection, model selection, and download resumption automatically. The structural significance is straightforward: every local inference run is a cloud API call that doesn't happen. At scale, tools like Strata shift the inference cost curve from per-token rent to one-time hardware investment, which is why this category of project — not this specific project, but the pattern it represents — is the most direct threat to the cloud inference business model.