Cloudflare has entered the decision-model market with Clef and Clef-flash, two models that produce bounded structured outputs — typed classifications with probabilities — rather than the open-ended text generation of LLMs. The core use case is programmatic routing: pass in a customer support ticket or a website domain, get back a probability distribution over categories, use it to trigger actions without a human in the loop. The models currently lead the Jev Decision Index, the benchmark established by Typesafe AI's own Jev model that created this product category weeks ago. The technical architecture is genuinely interesting. Both models use frozen Qwen backbones (Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash) and perform a prefill-only forward pass, then score valid schema choices in parallel via a specialized two-stage attention routing head. This is non-autoregressive — no token-by-token text generation, which is why latency drops dramatically. Median latency is 209ms for Clef and 38.8ms for Clef-flash, versus 524ms for Jev. The training uses rank-256 LoRA adapters jointly optimized with the routing head, label-smoothed cross-entropy, Brier loss for calibration, and a custom RL method called RLCD (Reinforcement Learning for Calibrated Decisions). Benchmark results are strong but require context. Clef tops Jev on BFCL case-exact accuracy (98.47 vs 95.75), API-Bank accuracy (91.93 vs 88.19), and BANKING77 macro-F1 (94.20 vs 79.74). Jev wins on When2Call accuracy (80.97 vs 72.37) and BRIGHT nDCG@10 (47.52 vs 45.91). On Typesafe's own eval suite, Clef beats Jev in 3 of 4 workflows but trails on agent trace observability. This is an honest competitive picture, not a sweep — though Cloudflare is clearly positioning Clef as the category leader. The differentiators beyond raw accuracy matter. Clef includes a vision encoder for image classification, which Jev lacks. The context window is 64k tokens versus Jev's 32k. And critically, Clef is API-compatible with Jev, meaning existing integrations can swap in with minimal code changes. The models are Apache 2.0 licensed on Hugging Face, so anyone can run them locally — but the hosted version on Workers AI leverages Cloudflare's edge GPU network for low-latency inference. The strategic play is the RL fine-tuning product. Cloudflare is not just releasing models — it is building a platform loop. You use Clef on Workers AI, your usage data feeds fine-tuning (with consent), the fine-tuned model redeploys to Cloudflare's edge. The initial offering is hands-on with a forward-deployed engineering team; self-serve comes later. This is the classic infrastructure vendor playbook: give away the model weights to build adoption, capture value through the hosted platform and fine-tuning services. The internal dogfooding case is illustrative. Cloudflare's Threat Intelligence team uses Clef with Browser Run to classify website domains — fetching, rendering, and classifying in 2.2 seconds versus 4.7 seconds for their fastest general LLM (gpt-oss-120b), while returning more classification categories. This is the pitch: decision models are not replacing LLMs, they are replacing the expensive, slow, non-deterministic classification step that agents currently delegate to LLMs. The open-source release under Apache 2.0 is genuinely generative — it expands the decision-model category beyond Typesafe's proprietary offering. But the platform bundling is where the extraction risk lives. If Cloudflare's edge hosting and fine-tuning become the default way to run decision models, the open weights are a loss leader for infrastructure lock-in. The question is whether the open-source community builds enough independent tooling to keep the category competitive.