Imagine you're an architect, but every wall, beam, and window must come from a fixed catalog of standardized pieces, each with exact connection points that physically snap together or don't. You can't fudge a dimension or invent a custom part. Your design has to look right AND actually stand up. That's the core constraint BrickBench imposes on AI agents: design LEGO sets from text descriptions using real parts from a discrete library, where every connection must be physically valid. The committed claim is that no existing benchmark tests this particular cocktail of skills — semantic understanding of what to build, combinatorial part selection from a finite library, and physical validity checking — in a single agentic loop. BrickBench fills that gap with three evaluation settings that vary in scale and part availability, scored across validity, alignment, and design quality. The authors also release BrickAgent, an environment where coding agents can construct, inspect, and iteratively validate their brick assemblies. Where does this sit against prior work? There's a lineage here from procedural LEGO generation (Luo et al., 2015; Kim et al., 2020) and text-to-3D benchmarks, but those either ignore physical validity or don't constrain to discrete part libraries. The closest conceptual relatives are assembly planning benchmarks in robotics and constraint-satisfaction problems in combinatorial optimization. BrickBench's contribution is stitching these concerns together under a single evaluation protocol. The paper reports that leading agents — presumably frontier LLM-based coding agents — "largely satisfy verifiable physical and semantic requirements, but fall short of human designs," which is an honest framing: the agents pass the physics checks but produce designs that are less creative and less coherent than what humans build. The architecture is a coding-agent loop: the agent writes code to place parts, queries the BrickAgent environment for validation feedback (collision detection, connection checks), and iterates. This is closer to tool-use agent evaluation than to end-to-end neural generation. The method leans on the LLM's code generation ability and the environment's physics engine for grounding. It's a benchmark, not a new model — which means the contribution lives or dies on whether the evaluation protocol is well-designed and adopted. Integrity is mixed. The benchmark itself is released publicly with code and a project page, which is strong. But the paper is evaluating agents built by others on tasks the authors designed — there's an inherent circularity risk in benchmark design where task difficulty can be tuned to produce the desired narrative. The comparison to human designs provides a useful ceiling, but we don't know the details of how human baselines were collected or how many humans participated. The three evaluation settings (varying scale and part availability) add robustness, but without independent adoption by other labs, the benchmark's signal quality is unverified. The milestone question is clear: current agents satisfy physical constraints but fall short on design quality. The next meaningful threshold is an agent that produces designs indistinguishable from human LEGO builders in blind evaluation — not just valid, but aesthetically and structurally competitive. That's probably 2-4 years out given the trajectory of agentic reasoning capabilities, and it would signal real progress in combinatorial creative reasoning. The obvious experiment not run: testing agents on reconstruction of known official LEGO sets from instructions or images, which would provide a ground-truth comparison without subjective design scoring. My read is (c) — they're saving it. A reconstruction benchmark is a natural follow-up paper, and the environment they've built supports it. The current paper establishes the open-ended design task first, which is the harder and more publishable framing.