Imagine you're a paralegal at a law firm. Your job isn't to argue the case — it's to find every relevant document in the filing cabinets and hand the lawyer a ranked stack with sticky-note justifications on each one. You don't write the brief. You retrieve the evidence. That division of labor is exactly what T-Search formalizes for multi-step question answering: a retrieval agent that runs bounded search rounds, returns ranked evidence chunks with justifications, and explicitly leaves answer generation to a downstream model. The committed claim: an open-weight 35B-parameter agent (Qwen3.6-35B-A3B, mixture-of-experts with 3B active) trained with adversarially filtered synthetic tasks and a two-stage regime — round-sliced supervised fine-tuning followed by GSPO on a recall reward — that outperforms larger open models on hard multi-step retrieval. Averaged over seven English and Russian benchmarks with gold evidence annotations, T-Search reaches 56.0 Recall@10 on a single rollout (+14.4 points over its base model) and 61.3 with three fused rollouts. The architectural bet is interesting and deliberate. By decoupling retrieval from generation, T-Search becomes modular: you can swap the search backend or the downstream generator without retraining the retriever. This is the opposite of monolithic RAG pipelines where retrieval and generation are entangled. The training pipeline is also worth noting — adversarial filtering of synthetic data means the training signal is specifically tuned to hard, multi-hop queries, not the easy single-hop stuff that inflates benchmarks. GSPO (a preference optimization variant) on recall reward directly optimizes for evidence coverage rather than surface fluency. On the ladder, T-Search claims superiority over "larger open models" on this benchmark suite, but the abstract doesn't name specific competing systems or provide head-to-head numbers beyond the base-model delta. The 14.4-point gain over its own base is real and meaningful, but we don't get a named SOTA comparison — a significant gap. The rollout-fusion trick (three rollouts → 61.3) is a useful practical lever, though it triples inference cost. The integrity picture is mixed-positive. Seven benchmarks with gold evidence annotations is a decent validation spread, and including both English and Russian is unusual and welcome. The release of model weights, harness, live demo, and three benchmarks (including TRuST, described as the first native-Russian hard-search benchmark) is genuinely strong open-science practice. But the benchmarks appear to be team-constructed, not community-standard, and there's no mention of pre-registration or independent replication. The most interesting artifact here may not be T-Search itself but TRuST — the first hard-search benchmark in Russian. Hard multi-step retrieval benchmarks are sparse even in English; a native-Russian one fills a real gap. If the community adopts TRuST, this paper's contribution to evaluation infrastructure could outlast its contribution to model architecture. The obvious experiment not run: scaling T-Search to a dense 70B+ model rather than the 35B MoE (3B active), or testing it against closed-model retrieval agents like Perplexity's or Google's Deep Research. The honest read is (a) and (c) — compute constraints limited the model scale, and testing against closed systems is both expensive and strategically saved for a follow-up that would attract more attention.