Imagine you have a seasoned sommelier. You can ask them to write a wine review — that's text generation. Or you can hand them four glasses and say 'rank these, percentages please.' A good sommelier can do both without retraining; the knowledge is already there, you just need to ask the right question. LLM2Jev is the realization that LLMs are already that sommelier for structured decision-making — you just need a clean protocol for reading their confidence levels off the existing next-token distribution. The core claim: general-purpose LLMs already function as 'Jev-style' decision models — systems that return calibrated probability distributions over a fixed set of labeled options — without any fine-tuning. The key insight is architectural: instead of generating free-form text and parsing it, LLM2Jev reads the next-token probabilities assigned to bracketed numeric identifiers (like [1], [2], [3]) that correspond to predefined choices. This turns the model's existing softmax output into a structured decision interface with zero additional training. The framework has two operating modes. The training-free recipe simply prompts the LLM, extracts logits over the option tokens, and normalizes. The fine-tuning pathway introduces a tree-factorized listwise loss that directly optimizes candidate selection, combined with KL divergence anchoring to prevent the fine-tuned model from forgetting how to hold a conversation. The anchoring matters: without it, you get a model that picks options well but generates degraded text, which defeats the purpose of using a general-purpose LLM in the first place. The empirical results on Qwen3.5-4B and Qwen3-0.6B tell a clean story with a clear inflection point. The 4B model, with zero training, matches purpose-built Jev-style models that were specifically fine-tuned on the same Qwen backbone. It also beats the naive approach of reading probabilities from letter tokens (A/B/C/D) rather than bracketed identifiers, handles arbitrary numbers of options without architectural changes, and natively processes multimodal inputs including images. The method is not just competitive — it demonstrates that the capability was latent in the base model all along. Fine-tuning follows a clear diminishing-returns curve. The smaller 0.6B model benefits substantially — it needs the optimization signal to compensate for weaker representations. Specific task types like many-option intent routing also see meaningful improvement. But for the 4B model on most tasks, fine-tuning adds marginal gains at best. LoRA emerges as the strongest fine-tuning strategy on capable models, balancing task-specific improvement with the KL anchor's protection of general capabilities. The broader implication is architectural, not incremental. If LLMs are already Jev-style decision models, the engineering question shifts from 'how do we build a decision model?' to 'how do we extract the decision model that's already there?' This reframes a substantial amount of work on structured output, tool use, and agent routing. The bracketed-identifier protocol is simple enough that it could become a de facto standard for LLM-as-classifier pipelines, particularly in production systems where retraining is expensive and base-model updates are frequent. What keeps this from being a complete answer is the narrow model family tested. Qwen 3 and 3.5 are strong models, but the claim 'modern LLMs are inherently effective decision models' rests on two checkpoints from one model family. The mechanism — reading logits over structured tokens — should generalize, but calibration quality across architectures (Llama, Mistral, Gemma) is an open empirical question. The paper also doesn't address adversarial robustness or calibration under distribution shift, which are the failure modes that matter most in production deployment.