Imagine you're a building inspector checking blueprints for hidden rooms. You scan every wall and floor plan for suspicious voids. Now imagine the builder hides the contraband room not in the structure but in the furniture placement notes — the legal, free-text part of the plan you don't scan with your ruler. That's Pretext. It moves the malicious payload from parseable code into the natural-language instructions that the skill legitimately needs, so the static checker sees nothing and the LLM judge reads it as the skill's stated purpose. The committed claim: current LLM-agent skill-scanning frameworks — static analysis plus an LLM semantic judge, exemplified by NVIDIA's SkillSpector — can be systematically defeated by an attacker with white-box access to the detector. The attack, called Pretext, iteratively rewrites malicious skills until they pass both layers while still delivering the payload and performing the benign task. Up to 97% evasion against a frozen detector, 77% against a co-adaptive one that tries to catch up. The mechanism is elegant and uncomfortable. Static analyzers look for code-level signatures — suspicious API calls, eval patterns, data exfiltration functions. Pretext lifts the malicious intent out of code and embeds it in the natural-language skill description, where it reads as the skill's legitimate operating instructions. The LLM judge, trained to flag semantically suspicious skills, gets fooled because the payload is framed as the skill's purpose and split across multiple files, keeping each individual chunk below the blocking threshold. It's a classic adversarial-examples playbook applied to the meta-layer: you don't attack the agent, you attack the safety system guarding the agent's marketplace. The ladder matters here. The baseline is NVIDIA's SkillSpector, which is essentially the only published defense in this space. There isn't a deep bench of competing frameworks. Pretext doesn't beat a crowded field of strong baselines — it beats the only real entrant. That's informative but also means the comparison is narrow. The 97% frozen-detector number is white-box, meaning the attacker has full access to the detection model's weights and architecture. The 77% co-adaptive number — where the detector is allowed to retrain against detected evasions — is the more meaningful figure, and it's still devastating. Architecturally, Pretext belongs to the iterative adversarial-generation family: generate candidate, test against target model, refine. The LLM attacker is the generator; the detector is the oracle. The key structural choice is the modality shift — moving payloads from code (where static analysis is strong) to natural language (where detection is fundamentally harder because intent is ambiguous). This exploits a basic asymmetry: code can be formally analyzed, but natural language descriptions of what a skill should do are evaluated by a model that must balance helpfulness with suspicion. Three open-source models were tested as attacker backbones, suggesting the technique is model-agnostic. The integrity picture is mixed. The evaluation runs across three open-source attacker models, which is good for generalizability. But validation is self-contained: the authors built the attack, ran it against a detector they reimplemented, and reported the results. No independent red-team validated the claims. The co-adaptive experiment is a nice addition that raises the bar beyond the easy frozen-detector win, but it's still the same team on both sides. Accepted at AIWild@NeurIPS 2026, which is a workshop track — peer-reviewed but not main-conference rigor. The milestone question is where this gets real. Right now, agent skill marketplaces (OpenClaw, Claude Code's tool ecosystem) are early. The question is whether a robust scanning defense can be built before these marketplaces scale to millions of skills. If the answer is no — if the static+LLM-judge architecture is fundamentally broken — the field needs a different paradigm entirely: runtime sandboxing, capability-based permissions, or formal verification of skill intent. Pretext doesn't answer what works; it demonstrates what doesn't. The obvious next experiment not run: testing against closed-source, proprietary detectors with unknown architectures (black-box attack), and testing against runtime monitoring rather than pre-installation scanning. The honest read is (a) — black-box attacks are harder to demonstrate convincingly without cooperation from the defender, and runtime monitoring is a different paper entirely. The authors scoped tightly and delivered on the scoped claim.