Imagine you're a restaurant manager with two kinds of staff: a veteran line cook who can plate a burger in 90 seconds without thinking, and a pastry chef who needs 40 minutes of careful measurement for a soufflé. Your job isn't just having both — it's knowing which one to call for each order. That's the core mechanism this paper proposes for AI: a metacognitive controller that decides whether a problem gets the fast, reflexive treatment or the slow, deliberate one. The committed claim: AI systems need a dual-process architecture — fast reactive agents (System 1) and slow deliberative agents (System 2) — mediated by metacognition, a self-model that monitors confidence and competence to decide which system to invoke. The authors argue this is necessary to move beyond narrow AI toward systems with broader, more human-like intelligence. The architecture draws directly from Daniel Kahneman's "Thinking, Fast and Slow" cognitive framework. System 1 agents exploit cached experience — pattern matching, learned heuristics, quick classification. System 2 agents engage search, planning, and explicit reasoning when System 1 signals low confidence. Both are grounded by two models: a world model (domain knowledge about the environment) and a self-model (records of past performance and solver capabilities). The self-model is the genuinely interesting piece — it's what distinguishes this from a simple ensemble or fallback chain. It enables the system to reason about its own competence boundaries. The ladder problem is stark: this is a position paper. There are no benchmarks, no experimental results, no comparison to any baseline. The authors cite related work — metacognitive architectures like MISM, CARINA, and various cognitive science foundations — but provide no empirical evidence that their proposed multi-agent architecture outperforms anything. In 2021 terms, the relevant baselines would be ensemble methods, confidence-calibrated neural networks, or existing cognitive architectures like SOAR and ACT-R. None are compared quantitatively. The integrity picture reflects the paper's genre. This is explicitly a framework proposal, not an experimental contribution. There are no datasets, no metrics, no ablations. The validation consists entirely of the theoretical coherence of the architecture and its mapping to cognitive science literature. That's legitimate for what it is — but it means the paper asks you to accept the architecture on conceptual merit alone. The self-model concept is compelling but unfalsifiable as presented. The milestone question reveals the paper's real limitation. What concrete metric would tell us this architecture works? The authors don't specify one. A reasonable next milestone would be: demonstrate that a metacognitive controller reduces total compute by X% on a mixed-difficulty benchmark while maintaining accuracy within Y% of an always-deliberate baseline. Without such a number, we cannot track progress. The field has since moved toward this territory with LLM routing, mixture-of-experts, and adaptive compute (e.g., early exit transformers), but this 2021 paper predates those developments. The obvious experiment not run: implement the architecture on any concrete domain and measure whether the metacognitive routing actually improves the cost-accuracy tradeoff. The honest read is (a) — this was a conceptual contribution, likely from a team oriented toward AI theory and cognitive science rather than large-scale systems engineering. The ideas have since been partially validated by the LLM era's interest in adaptive compute and chain-of-thought prompting, but the paper itself remains a blueprint without a building.