Imagine you're assembling IKEA furniture, but every time you finish a step, someone wipes your short-term memory and hands you only the last screw you placed. That's how autoregressive Transformers generate text: the only information flowing from one token to the next is the decoded token itself. Every intermediate representation — every half-formed thought about alternative continuations — gets discarded. LIFT breaks this constraint by giving the model a side channel: a latent state vector that carries deep-layer information forward across generation steps. The committed claim: Transformers can learn to exploit deep-to-shallow feedback during pretraining, via a scalable teacher-supervision scheme, and this produces consistent gains on language modeling and reasoning tasks across model sizes from 135M to 1B parameters. This is not the first paper to propose recurrent state for Transformers — MEGALODON, Mamba, and various state-space models have explored this territory — but it is a genuinely novel training method. The trick is converting recurrent-state learning into a prediction problem: pair each input token with a state derived from the next-token distribution of a pretrained teacher LM, then train the student to predict both the next token and the next state. Because teacher states are precomputed, training stays fully parallel. At inference, the model feeds back its own predicted states, with overhead that shrinks as model size grows. The ladder here is honest but modest. LIFT consistently beats standard Transformers and baselines under token-matched budgets — same data, same training tokens, LIFT wins. Under compute-matched conditions (accounting for the overhead of the extra state prediction head), LIFT is on par with or slightly ahead. The standout result is a controlled state-tracking experiment where a tiny LIFT model trained on X data outperforms a same-size standard Transformer trained on 8× more data. That's the kind of sample-efficiency result that makes architecture papers interesting. But no comparison to Mamba or other SSMs is reported, which is a notable gap. Architecturally, LIFT lives in the attention-based Transformer family but grafts on a recurrent state pathway — making it a hybrid. The teacher is any off-the-shelf pretrained LM; the student adds a small prediction head for state vectors. The key structural choice is that the state is a compressed representation of the teacher's next-token distribution, not raw hidden states. This is clever because distributions are more information-dense and more transferable than arbitrary internal representations. The compute overhead is a few extra parameters and one additional loss term during training, plus one state-prediction pass per step at inference. Integrity is mixed. The authors run experiments across multiple model sizes (135M, 350M, 1B), which is good practice for an architecture paper. They report both token-matched and compute-matched comparisons, acknowledging that LIFT has overhead. The state-tracking controlled study is well-designed — including the remarkable detail that LIFT learns state-tracking even when its teacher Transformer fails the task, suggesting the student extracts structural signal the teacher didn't explicitly have. However, benchmarks are not pre-registered, no code release is mentioned, and the baselines don't include the most aggressive recent competitors in the recurrent-state space (Mamba-2, RWKV-6, Griffin). The milestone question: at 1B parameters, LIFT shows clear gains. The natural next number is 7B — the scale where most open-source practitioners operate and where the overhead fraction becomes negligible. If LIFT's gains hold or widen at 7B, this becomes a practical training recipe, not just a research contribution. If gains flatten, this stays an interesting architectural insight. The gap is probably 6-12 months and primarily a compute budget question. The obvious experiment not run: scaling to 7B+ and comparing head-to-head against Mamba-2 or Griffin on the same benchmarks. The honest read is (a) compute budget — training a 7B model is expensive, and this is a three-author academic paper, not a lab with 10,000 GPUs. The absence of SSM comparisons is harder to excuse; that's probably (c) — strategic scoping to keep the narrative clean and the paper focused on the teacher-supervision mechanism rather than getting dragged into the attention-vs-recurrence horse race.