Imagine you're in a dark room with a row of candles stretching away from you. The nearest candle is brightest; each one further away is slightly dimmer. You don't need to count candles or read labels to know which is closest — the brightness gradient tells you. That's the core mechanism this paper identifies in hybrid transformer architectures: local mixing layers act like those candles, creating a decaying signal that global attention reads as implicit positional information. The committed claim: sliding window attention (SWA) and gated linear attention, when interleaved with global NoPE (No Position Encoding) layers, induce a recency bias in the residual stream that functions as an implicit position encoding — and this bias persists across long sequences, unlike the position signal that emerges from the causal mask alone in pure NoPE models. This is a mechanistic explanation for something practitioners already knew worked empirically but couldn't explain. The architecture story matters here. Modern LLMs like Gemma and Jamba have quietly moved to hybrid designs that alternate local mixing layers (SWA, gated linear attention) with global attention layers that carry no explicit position encoding like RoPE. These hybrids match or beat pure-RoPE models at scale, but the field has been flying blind on why. This paper places the mechanism squarely in the residual stream: local layers apply position-dependent transformations that leave a decaying fingerprint. Global attention's dot-product logits then select on this fingerprint. The result is a smooth recency bias — tokens know roughly how far apart they are without anyone explicitly telling them. The theoretical contribution involves showing that SWA applies a banded, position-dependent linear transformation to the residual stream. Because SWA's window is finite, tokens inside the window get transformed differently depending on relative distance, and this asymmetry accumulates across layers. Gated linear attention achieves something similar through its recurrent structure — an exponential decay that weights recent tokens more heavily. Both mechanisms create what the authors call a "recency bias" that global attention logits can exploit. On the empirical side, the authors train small-scale hybrid models and measure attention logit distributions, confirming that global NoPE layers in hybrid architectures exhibit a smooth, monotonically decaying recency bias. They contrast this with pure NoPE models (global attention only, no local mixing), where positional information comes solely from the causal mask — a much weaker signal that degrades as sequence length grows. The hybrid recency bias, by contrast, is maintained across long sequences because each local mixing layer refreshes it. The length extrapolation angle is the paper's most practically significant thread. RoPE and other explicit position encodings famously struggle when inference sequences exceed training length. The authors argue that the recency bias mechanism in hybrid models doesn't have this limitation — the local layers refresh the positional signal at every application, so it doesn't decay or distort at lengths never seen during training. This is an insight, not a proof of indefinite extrapolation, but it points toward a design principle: position encoding through local mixing may be inherently more length-robust than explicit encodings. What this paper doesn't do is validate the theory at frontier scale. The experiments are on small models designed to isolate the mechanism. The authors don't train a 7B+ hybrid and show that their explanation predicts the attention patterns observed there. They also don't directly benchmark length extrapolation against RoPE at scale. This is a mechanistic explanation paper, not an engineering paper — and that's fine, but it means the practical implications remain to be confirmed by the groups building production hybrids.