Imagine you're training a pilot exclusively in a classroom with paper diagrams of cockpits, then on graduation day you hand them a real aircraft with live instruments, radio chatter, and weather. They know the theory but have never touched the controls. That's roughly the state of text-to-SQL model training: models learn to map questions to static SQL strings, but in production they operate inside a multi-turn execution harness — inspecting schemas, running probe queries, reading error messages, and revising. HarnessSQL's core insight is embarrassingly simple: stop training outside the cockpit. The committed claim is that training database agents directly within their deployment harness — not just evaluating them there — is essential for complex, long-horizon SQL workflows. This isn't a new architecture or a new model family; it's a training-regime paper. The authors take existing compact models (Qwen3-8B and Qwen3-14B) and show that the right post-training pipeline can close the gap between small open models and much larger proprietary systems on realistic benchmarks. The pipeline has three stages. First, they build isolated, executable database sandboxes paired with hidden execution oracles — ground-truth query results that never leak into the training signal directly. Second, they roll out teacher trajectories inside the actual SQL harness, filtering to keep only verified trajectories (ones whose final answer matches the oracle). These become supervised fine-tuning data that preserves the full multi-turn interaction structure: schema inspection, probe queries, error recovery, hypothesis revision. Third, they run reinforcement learning with execution-reward signals — the model gets credit for producing correct results, not for matching a reference SQL string. The ladder numbers are striking. On Spider 2.0-SQLite, Qwen3-8B jumps from 15.5% to 45.2% execution accuracy — nearly tripling. Qwen3-14B goes from 22.2% to 54.8%, a 2.5× improvement. These are compact models; the fact that a 14B parameter model can reach 54.8% on a benchmark designed to be hard for frontier models is the headline. The paper also shows transfer to out-of-distribution benchmarks (BIRD-Interact and LiveSQLBench), which is the real test of whether the method teaches generalizable interactive skills or just overfits to Spider's distribution. The architecture family is straightforward: decoder-only transformer LLMs (Qwen3 series) post-trained with a two-phase regime — full-sequence SFT on filtered expert trajectories followed by execution-reward RL. The key structural choice is retaining the entire interaction trace as the training sequence, rather than collapsing it to a single (question, SQL) pair. This means the model learns to use the harness — schema lookups, intermediate queries, error messages — as part of its reasoning chain, not as an afterthought bolted on at inference. Integrity is reasonable but not airtight. Spider 2.0-SQLite is a community benchmark, which is good. The transfer experiments on BIRD-Interact and LiveSQLBench provide out-of-distribution validation. However, the authors are grading their own homework — no independent replication exists, and there's no pre-registration. The teacher trajectories are generated by the authors' own pipeline, and the oracle filtering introduces a selection effect: we only see what worked. The RL reward is execution match, which is a strong signal but can reward correct-by-accident queries. The obvious next experiment is scaling beyond 14B parameters — does the harness-native training regime help or become redundant for 70B+ models that might already have enough capacity to figure out the harness at inference time? The authors almost certainly have access to larger Qwen3 variants. My honest read: they either ran it and the gains were smaller (making the "compact model" framing the better story), or they're saving the scaling curve for the next paper. Either way, this is the experiment that would tell you whether harness-native training is a permanent feature of the pipeline or a crutch for smaller models.