Imagine you're solving a jigsaw puzzle, but instead of starting from the picture on the box, you've been working with the individual cardboard fibers. Every piece is defined by its edges (links), and you've been trying to learn the correlations between fibers that span the entire puzzle. What if you could work with the assembled patches (plaquettes) instead — the actual picture fragments that snap together locally? The catch: not every combination of patches is a legal puzzle. There are geometric constraints (Bianchi identities) that say "these four patches around a cube must be consistent." This paper's trick is a multilevel construction that builds the puzzle coarse-to-fine, so at each step you only need to check consistency among at most four new patches at a time, regardless of how big the puzzle gets. The committed claim: a normalizing-flow sampler that operates directly in plaquette space, satisfies Bianchi constraints exactly (not approximately), and whose constraint-solving cost per refinement step is independent of lattice volume. This is a genuine architectural novelty. Prior generative-model approaches to lattice gauge theory have worked in link space, where the action involves products around plaquettes and correlations become increasingly nonlocal as you approach the continuum limit (weak coupling). Working in plaquette space is the obvious move — the action is local there — but the exact constraints have blocked direct generative modeling until now. The construction factorizes the joint distribution over plaquettes via a coarse-to-fine hierarchy. At the coarsest level, a small set of plaquettes is sampled by a normalizing flow. At each refinement, new plaquettes are generated conditioned on the coarser level, and the Bianchi constraints are solved locally — each determined plaquette depends on at most four newly generated variables. The key engineering insight is that the constraint problem does not grow with lattice size; it stays O(1) per local solve. This is the multilevel Plaquette-Space Sampler (PSS). They validate on U(1) in 2D and 4D, and SU(2) in 2D. The headline result: PSS substantially outperforms link-space normalizing-flow baselines, and the advantage grows toward weak coupling — precisely the regime where link-space methods struggle most and where continuum physics lives for asymptotically free theories like QCD. In 2D U(1), they show effective sample sizes (ESS) near 100% where link-space flows collapse. In 4D U(1), the advantage persists on lattices up to the sizes tested, with PSS maintaining high ESS at couplings where the link-space baseline is essentially useless. The integrity picture is solid for a methods paper. They compare head-to-head against link-space normalizing flows using the same flow architecture and training protocol — the only variable is the representation space. This is the right comparison. They report ESS, acceptance rates, and KL divergence across multiple coupling values. The 4D U(1) and 2D SU(2) results extend beyond toy demonstrations, though they acknowledge the big gap: SU(3) in 4D, which is the actual target for QCD. No code release is mentioned. The milestone that matters is SU(3) in 4D at physically relevant lattice sizes — say 16⁴ or larger with β values near the continuum scaling window. The current paper demonstrates the architecture on simpler gauge groups and lower dimensions. The gap between SU(2) in 2D and SU(3) in 4D is not just quantitative; SU(3) constraint solving is algebraically harder, and 4D lattices explode the number of plaquettes. A realistic estimate for closing this gap is 2-4 years, assuming the constraint-solving machinery generalizes and GPU memory/compute keeps pace. The obvious experiment they didn't run is SU(3) in 4D. The honest read: this is almost certainly (a) — the constraint solver for SU(3) requires nontrivial group-theoretic machinery and the 4D volume scaling demands serious compute. They likely wanted to establish the conceptual framework and demonstrate it on tractable groups first. This is a reasonable sequencing decision, not a red flag. The second missing experiment is a direct comparison against state-of-the-art HMC with gradient flow or multi-level HMC methods, rather than only against link-space normalizing flows.