Imagine you run a restaurant kitchen where the grill cook, the sauté cook, and the prep cook all use the same set of knives, cutting boards, and burners — just with different ingredients loaded onto the same stations. That is what Whistle does with neural network blocks. Instead of training a separate speech model from scratch, it literally shares the same Simple Attention blocks and Monarch Hadamard MLP code as its sibling language model Needle. The speech-specific addition is exactly one gated cross-attention per decoder layer. Everything else — the self-attention, the convolutions, the engram lookups — is Needle's code running Whistle's weights. The committed claim: a 16.9 MB speech recognition model, running on CPU with zero dependencies, that beats Whisper base (145.3 MB) on LibriSpeech test-clean, test-other, SPGISpeech, Earnings-22, and FLEURS, while decoding at 1,319 tokens per second versus Whisper's 266. Seven languages. Word timestamps. Speech embeddings. All from one .cact file that loads into the same binary as the text model. The architecture is a classic encoder-decoder with specific choices that explain the compression. The encoder runs eight non-causal Simple Attention blocks with four multi-head-channel residual lanes and Monarch Hadamard MLPs — the Monarch structure gives you dense-equivalent mixing at sub-dense parameter cost. The decoder runs eight Laddered Simple Attention blocks with grouped-query attention (8q:2kv), 3-tap causal convolutions, and engram lookups at layers 3 and 7. The ladder training means every depth from 2 layers up is a valid model, selectable at load time via --audio-depth. The encoder always runs all eight. Cross-attention K and V are projected once per clip and shared across all five beams, so beam search costs five transcript caches, not five encoder passes. The ladder against baselines is honest and specific. Whistle leads on LibriSpeech clean and other, SPGISpeech, Earnings-22, and FLEURS average. Whisper base leads on TED-LIUM, AMI, and MLS average. The authors flag that Moonshine is English-only and that Whisper's AMI figure is AMI-IHM (a different subset), rather than hiding the comparison gaps. On speed, the gap is stark: 11.1 ms to first token versus 73.2 ms (Whisper) and 22.8 ms (Moonshine tiny v2). Decode throughput is 1,319/s versus 266/s and 262/s respectively. All measured on Apple M4 Pro CPU with each model's official runtime at defaults. Integrity is above average for an industry release. The 86,174-utterance evaluation uses Whisper's own normalizers. The authors verified no test audio appears in training or validation data by comparing audio checksums and speaker IDs — a step many papers skip. Code ships on GitHub, weights on Hugging Face, and the engine compiles for 17 targets. The weak spot is the absence of pre-registration and independent replication: these are the authors' own numbers on their own runtime. The practical unlock is not the WER numbers — it is the architecture sharing. One C binary, one quantization format, one container loads both a language model and a speech model. The third command-line example shows the real product: audio in, tool calls out, single JSON response. For edge deployment on wearables, robots, automotive, and microcontrollers, collapsing two model runtimes into one eliminates an entire integration surface. The 16.9 MB footprint fits in the L3 cache of most modern SoCs. The obvious next experiment is scaling: more languages, longer audio (the 30-second cap is a real constraint for meetings and podcasts), and noisy/far-field conditions that stress real-world deployment. The authors almost certainly have multilingual expansion in progress — the seven-language-token vocabulary design is clearly built to grow. Whether the Monarch Hadamard MLP and engram architecture hold up at 50+ languages and 5-minute clips is the open question that will determine whether this is a clever small model or the seed of a new ASR stack.