You know how a competent new hire can look incompetent if you hand them a terrible onboarding process — confusing documentation, broken dev environments, meetings that eat their entire day? Then you replace them, and the next hire fails the same way, and eventually someone asks: is the problem the people or the process? That's the core mechanism of this paper. Small open-weight language models (2-9B parameters) running on ordinary laptops consistently fail at real agent tasks under cloud-scale harnesses. The authors argue — and provide controlled evidence — that the harness is the bottleneck, not the model. The committed claim: a purpose-built local-first agent harness with ten compensatory mechanisms can close most of the task-completion gap between small open models and frontier models. Mingbird, targeting Windows + Ollama, introduces three representative mechanisms: a byte-level net-zero prefill budget that prevents tool descriptions from overflowing the model's context window, a finish gate that re-reads the original task before accepting completion (preventing premature declaration of victory), and signature-level loop detection that catches the model repeating the same failed action pattern. On their own LRAB benchmark — 4 harnesses × 4 models (2B-35B) × 18 real tasks, all 288 cells published — Mingbird scores 0.886 overall versus 0.631 (goose), 0.479 (opencode), and 0.405 (agent-mini). On the third-party τ²-bench (278 tasks), it totals 0.856 against 0.791 and 0.737. The architectural family here is agent scaffolding engineering, not model architecture research. Mingbird doesn't change the model — it changes everything around the model. The ten mechanisms are essentially guardrails that prevent the specific failure modes small models exhibit: context overflow from tool prefill, divergent self-correction loops, silent task abandonment, and looping tool demonstrations. The compute property being exploited is that these models already have sufficient reasoning capability for many tasks; the overhead of local inference is acceptable if the harness doesn't waste context tokens. Integrity is the paper's most interesting tension. The authors are unusually transparent about limits: LRAB is self-built, tested on a single machine, with single-trial scoring. They report that same-night replications of identical arms shift mean scores by up to 0.069 — which is the same magnitude as every single-mechanism ablation delta. This means the ablation results are directional at best, not statistically separable from noise. The one batch-matched comparison (full stack vs. text re-read alone) shows a paired +0.10 across three replications, which is the strongest internal signal. The τ²-bench result on 278 tasks provides independent validation, but margins are tighter. All 288 per-cell results and scoring code are published — a meaningful transparency gesture. The frontier-model probe is revealing: running GPT-4-class models through the same four harnesses produces scores spanning 0.997 to 0.478, with well-formed scaffolds clustering within 0.072 of each other. This is the paper's strongest structural argument — even frontier models suffer massive performance drops under bad scaffolding, which supports the thesis that harness quality dominates model quality for task completion. The milestone question is whether this approach scales beyond 18 tasks and a single machine. The immediate next number is whether Mingbird's mechanisms hold across 100+ diverse real-world tasks with varied tool ecosystems on heterogeneous consumer hardware. The gap between LRAB's controlled 18-task setup and production-grade local agent deployment is substantial. The successor experiment the authors didn't run — a multi-machine, multi-OS replication with independent scorers — is the obvious next step. The honest read: they built this for Windows + Ollama on one machine, and broadening the test matrix is a resource constraint, not a hidden failure. This paper matters most as a diagnostic argument: stop blaming the model and start auditing the scaffolding. For practitioners running local agents, the specific mechanisms (prefill budgets, finish gates, loop detection) are immediately actionable design patterns regardless of whether Mingbird itself becomes the standard harness.