Imagine you're packing for a trip with two suitcases: one is deep (you can stack items vertically — layers of shirts, pants, jackets) and one is wide (you can spread items across a huge flat surface — one layer but enormous). Packing theory has always modeled depth and width separately. This paper writes the first packing formula that tells you how deep AND wide to go simultaneously, given a fixed total luggage allowance. The committed claim: this is the first scaling law that jointly models recurrence (looping) and sparsity (MoE routing) alongside model size and data. The core mathematical object is a bounded, sparsity-conditional recurrence mapping — essentially a function that tells you how many effective parameters you get from a given number of loops, and how that gain changes as you add more experts. Prior scaling laws from Chinchilla (Hoffmann et al., 2022) and MoE-specific work (Clark et al., 2022) treated these axes in isolation. This paper nests both as special cases: set loops to 1 and you recover standard MoE scaling; set experts to 1 and you recover dense looped scaling. That's a genuinely clean theoretical contribution. The numbers tell a clear story. Sparsity delivers roughly 3× active-parameter efficiency — meaning a sparse model performs like a dense model three times its active parameter count. Recurrence delivers roughly 2× total-parameter efficiency specifically on reasoning benchmarks. Crucially, these gains compose: a looped MoE with law-derived recurrence depth matches a non-looped MoE roughly 2× its size on reasoning tasks, at matched training compute, validated at trillion-token scale. The laws also predict held-out loss more accurately than prior scaling law alternatives, which is the real test — if your law can't predict loss on unseen configurations, it's curve-fitting, not science. Architecturally, this sits at the intersection of two well-established families. Looped transformers (weight-sharing across layers) trace back to Universal Transformers and ALBERT, trading parameter count for depth via iterative refinement — think of it as unrolling an ODE solver. MoE routing (Switch, GShard, Mixtral lineage) trades dense compute for sparse capacity by activating only a subset of experts per token. The paper's key structural insight is that sparsity raises the ceiling on recurrence gains — more experts make each additional loop more valuable, likely because diverse expert representations give the recurrence richer material to refine. On integrity, this is a Meta AI paper evaluated on held-out perplexity and downstream reasoning benchmarks. The scaling law is fitted to training runs across a grid of model sizes, loop counts, and expert counts, then tested on configurations outside the training grid — the right validation approach. Downstream evaluations include reasoning benchmarks, which is where the 2× efficiency claims land. No pre-registration, and the benchmark suite appears chosen post-hoc, but the perplexity predictions are the load-bearing validation and those are harder to cherry-pick. No independent replication yet. The practical implication is architectural design guidance. If you're building large language models under a fixed compute and memory budget, this law gives you a principled way to allocate between more experts, more loops, or more base parameters — rather than grid-searching blindly. The test-time scaling angle is particularly interesting: recurrence enables spending more compute at inference by running more loops, which is the same lever that chain-of-thought and iterative refinement exploit but at the architecture level. The obvious next experiment they didn't run: scaling this to frontier-class models (hundreds of billions of active parameters) and testing whether the law's predictions hold or break down. At the demonstrated scale, the law works — but scaling laws famously develop phase transitions and break points at larger scales. My read is (a) compute budget: training enough frontier-scale configurations to validate the law would cost tens of millions of dollars. They're likely saving this for a follow-up with more resources, or waiting for the next hardware generation to make the experiments tractable.