Salvatore Sanfilippo — antirez, the creator of Redis — has released DwarfStar 4 (ds4), a purpose-built C inference engine that runs frontier open-weight models locally on high-memory Macs, CUDA, and ROCm machines. The supported model families are DeepSeek V4 and V4.1 Flash, GLM 5.x, and Qwen3.8 Flash Next, covering both text and vision. The engine is MIT-licensed and ships with a CLI, an OpenAI/Anthropic-compatible API server, and a native coding agent — all sharing model state and KV cache. The core technical move is asymmetric 2-bit quantization with importance matrices. Rather than uniformly crushing the entire model, ds4 compresses the routed MoE experts aggressively while preserving critical shared pathways at higher precision. This is what makes a 284-billion-parameter model fit on a 128 GB machine at all. The project ships its own validated GGUF files rather than accepting arbitrary ones — a deliberate narrowing that trades generality for end-to-end correctness. Performance benchmarks on an M5 Max with 128 GB show 790 tokens/second prefill and 39.4 tokens/second generation at 2K context, dropping to 398 t/s prefill and 27.6 t/s generation at 65K context. On the DGX Spark (also 128 GB), prefill holds at 823-825 t/s across context lengths but generation falls to 13.8-18.1 t/s. These numbers are usable for real work — not parity with cloud H100 clusters, but sufficient for coding agents and interactive sessions. The KV cache design treats SSD as a first-class citizen. Long prefixes can be saved to disk and resumed by prompt hash, meaning restarts don't require full re-prefill. For agent workloads that revisit large codebases repeatedly, this is a meaningful architectural choice — it turns the cache into persistent state rather than volatile memory. Ds4 is explicitly not a general-purpose GGUF runner. It follows a "small, opportunistic set of model families" and validates each layout end to end. This is the opposite of llama.cpp's strategy of broad model support. The tradeoff is clear: you get fewer models but higher confidence that the ones you get actually work correctly against official outputs. The strategic signal here is about control and resilience. Every API call to a cloud LLM is a dependency on someone else's pricing, availability, and policy decisions. Local inference at frontier quality — even at reduced precision — eliminates that dependency for developers, researchers, and organizations that need predictable, private, zero-marginal-cost inference. The 128 GB hardware requirement is steep but falling. Antirez's track record (Redis reached ubiquity as infrastructure) lends credibility to the engineering claims, though the project is early. The real question is whether the narrow-model strategy scales: as new frontier models ship quarterly, can a small team keep validated builds current? The answer will determine whether ds4 becomes infrastructure or remains an impressive one-off.