Imagine you're juggling. Two balls? Easy — your hands alternate in a simple nested pattern. Three balls in a cascade? Still manageable, because each ball follows the one before it in a predictable nested stack. Now try juggling where ball 1 must be caught after ball 3, and ball 2 after ball 4 — a cross-serial pattern where the order of catches doesn't mirror the order of throws. Your hands can't track it without conscious effort. This paper asks whether neural language models have the same limitation, and whether that limitation explains something deep about human language. The committed claim: stack-based language models (SLMs) — neural architectures augmented with an explicit push-pop stack — fail to generalize on cross-serial dependencies (the formal upper bound of attested syntactic complexity), but SLMs with deliberately limited working memory perform better on these constructions. This suggests that constrained memory could be a source of the inductive biases that explain why certain word orders (like SOV) dominate the world's languages while others don't. The lineage here runs through formal language theory and computational typology. Cross-serial dependencies sit at the boundary of mildly context-sensitive languages — the formal class believed to capture the outer edge of natural language syntax (think Swiss-German verb constructions). Prior work by Deletang et al. (2023) and others used standard LMs (LSTMs, Transformers) to probe learning biases on artificial languages. This paper extends that program in two directions: using a harder formal class (cross-serial rather than just nested dependencies) and using a more linguistically motivated architecture (stack-augmented models rather than vanilla sequence models). The architecture family is stack-augmented RNNs — specifically, models in the line of Joulin & Mikolov (2015) and Suzgun et al. (2019) — where a differentiable stack provides explicit hierarchical memory on top of a recurrent backbone. The key structural choice is varying the stack depth (working memory capacity). The paper tests these against standard LSTMs and Transformers on artificial languages encoding different word-order typologies with cross-serial dependency patterns. The finding that limited-depth stacks outperform unlimited ones on generalization is the paper's load-bearing result. Integrity is mixed. The evaluation is entirely on artificial languages — synthetic grammars designed to isolate specific formal properties. This is standard methodology for this subfield, but it means the validation is self-contained: the authors designed the test, trained on it, and evaluated on it. There's no independent benchmark, no pre-registration, and no natural language validation. The gap between artificial grammar generalization and actual typological explanation remains a significant inferential leap. The paper is honest about this, but the reader should hold the conclusions loosely. The milestone question is where computational typology meets cognitive science: can these inductive bias findings predict actual typological distributions — not just that SOV is common, but the specific frequency ranking of attested word orders? The field needs to move from 'model X struggles with pattern Y' to 'model X's error distribution correlates with cross-linguistic frequency at r > 0.8 on a held-out language sample.' That's probably 3-5 years out and requires collaboration between computational linguists and typologists with access to large-scale crosslinguistic databases like WALS. The obvious experiment not run: testing these stack-based models on actual natural language data with cross-serial constructions (Swiss-German, Dutch verb clusters). My honest read is (a) — the artificial language setup is cleaner and publishable, and natural language evaluation would require significant additional annotation effort and introduce confounds that would muddy the formal story. It's methodologically defensible to defer, but it's also the experiment that would turn this from a formal result into an empirical one.