Imagine you're building a house out of Lego instead of clay. Clay gives you any shape, but once it's fired you can't change the internal structure — you'd have to start over. Lego constrains you to discrete bricks, but every brick stays individually removable and replaceable. MatLoom makes the same trade for material generation: instead of outputting a dense pixel grid like diffusion models, it outputs a short program — a stack of named, alpha-masked layers with explicit PBR channels — that an interpreter renders into the final material maps. The program IS the asset. The committed claim: a compact domain-specific language (MatLoom DSL) plus parser-guided repair and preview-based critique can produce text-to-material outputs that beat diffusion baselines on prompt alignment, while retaining the full construction logic as editable source code. No fine-tuning of the underlying LLM is required. This is not the first procedural material system, but it's a genuine first in framing LLM-driven material generation as code synthesis in a purpose-built layer language with PBR semantics baked in. On the ladder: the paper benchmarks against three diffusion baselines across 141 prompts using six LLM backbones. The best-performing MatLoom configuration beats all three baselines on all four flat-layout prompt-alignment metrics (mean scores). Initial programs — before any critique or seed search — already exceed all baselines on mean BLIPScore. In a blind user study with 30 participants and 20 prompts, MatLoom renders captured 59.2% of preferences versus 19.3% for the strongest baseline. Those are substantial margins, though the evaluation is entirely prompt-alignment and perceptual preference, not ground-truth PBR accuracy against scanned materials. Architecturally, this sits in the program-synthesis-via-LLM family — closer to AlphaCode or Voyager than to Stable Diffusion. The DSL constrains the output space to composable layers with shared spatial expressions, so the LLM's job is to fill a structured template rather than hallucinate arbitrary code. Parser-guided repair catches syntax errors; preview-based critique uses a vision model to compare rendered output against the prompt and suggest revisions. Noise seed search then optimizes stochastic parameters while keeping the program structure fixed. The key compute property: because programs are short (median 21 lines), the LLM generation cost is trivial compared to diffusion inference, and the interpreter is deterministic. Integrity is mixed. The curated 141-prompt benchmark and the user study are both author-constructed. Six LLM backbones provide breadth, and the user study is blind and four-way, which is good practice. But there's no community benchmark for text-to-material (the field is young), no pre-registration, and no code release is mentioned in the abstract. The baselines are three diffusion methods — reasonable choices for the current field, but the paper doesn't name them in the abstract, making independent triage harder. The 59.2% preference number is strong but comes from 30 participants × 20 prompts, a modest sample. The milestone question is interesting. Right now, the median program is 21 lines and handles flat-layout materials. The next concrete threshold is whether this approach can handle multi-scale, non-flat materials — think weathered brick with mortar variation, or fabric with weave-level detail — where the layer abstraction might break down. If MatLoom-style programs can reach ~100 lines handling 3-4 scales of detail while remaining human-editable, that would make it a credible replacement for Substance Designer graphs in production pipelines. The obvious experiment not run: integration into an actual game or film art pipeline with artist feedback on editability. The paper claims retained programs are editable, but no artist study validates this. My read: this is being saved for a follow-up or industry collaboration paper. The infrastructure for an artist study is different from the ML evaluation pipeline, and it's a natural second publication.