Imagine you run a hospital switchboard. Every incoming call could be cardiology, surgery, or pharmacy — and the best operator for each is different. A dumb switchboard routes everyone to the same generalist. A smart one reads the caller's chart, checks which department they need, and connects them to the right specialist — dynamically, per call. ARBOR does this for low-rank adaptation: instead of applying one fixed LoRA update to every medical question, it reads the question text, its specialty tag, its operation type (diagnosis vs. treatment vs. pharmacology), and their interaction, then selects a subset of rank-one 'atoms' from a shared basis. A learned scalar then controls how much the adapter residual contributes. The committed claim: conditional rank allocation guided by clinical taxonomy tags beats fixed-rank LoRA and MoE-LoRA on four medical QA benchmarks, and the advantage grows with specialty diversity. The paper reports 69.69% mean accuracy across CMB, CMExam, MedQA, and MedMCQA on Qwen3-8B, exceeding LoRA r=16 by 1.26 pp and MoELoRA by 1.30 pp over five seeds. Critically, when training expands from one to seven specialties, the gap over LoRA r=16 widens from 0.08 to 1.94 pp. That scaling behavior is the real signal: the method's advantage is negligible in the single-specialty regime and emerges precisely when routing matters. Architecturally, ARBOR sits in the conditional/gated LoRA family — think MoE-LoRA, AdaLoRA, and recent routing-based PEFT work. The key structural choice is an additive gate that combines four signal sources (question representation, specialty tag, operation tag, and their pairwise interaction) to produce atom-selection weights. This is conceptually cleaner than top-k routing in MoE-LoRA because it decomposes routing into interpretable clinical dimensions. The paper includes an illustrative proof showing conditional selection avoids an approximation floor that fixed-rank updates hit under orthogonal, equiprobable subtasks — but the authors are admirably careful to note this doesn't directly bound performance on real medical corpora. Integrity is mixed but honest. The benchmarks (CMB, CMExam, MedQA, MedMCQA) are community-standard medical QA datasets, not cherry-picked. Five-seed experiments with reported means is better than most PEFT papers deliver. Tag perturbation and atom-masking ablations support the claim that clinical routing is doing real work. The adjusted Rand index of 0.62 between atom clusters and specialty labels is a nice interpretability check. However, the baselines are limited to LoRA r=16 and MoELoRA — no comparison to AdaLoRA, DoRA, or full fine-tuning. And the 1.26 pp improvement, while consistent, is modest in absolute terms. The honest gap: clinical safety and deployment are explicitly untested. The paper says so in its final sentence, which is refreshing. But the more revealing gap is the absence of scaling experiments on larger models — Qwen3-8B is the only backbone tested. Whether conditional routing still helps at 70B parameters, where the base model may already have sufficient capacity to internalize specialty distinctions, is the obvious follow-up the authors did not run. My read: they ran out of compute. Multi-GPU experiments on 70B with five seeds across four benchmarks and seven specialties is expensive, and this is a nine-author academic group, not a lab with 1,000 H100s. The milestone to watch is whether conditional PEFT methods like this can demonstrate gains on clinical deployment metrics — not just benchmark accuracy, but calibration under distribution shift, safety on adversarial medical queries, and performance in real clinical decision support. The paper's own calibration analysis is a good start, but the gap between 'beats LoRA on MedQA by 1.3 pp' and 'changes how clinicians use LLMs' remains large. Bottom line: a well-executed, honest PEFT paper that introduces a clean mechanism and demonstrates it on the right benchmarks. The scaling behavior across specialty count is the strongest evidence that conditional routing isn't just noise. But the effect sizes are modest, the model scale is limited, and clinical validation is explicitly absent. Worth reading if you work on medical NLP or parameter-efficient fine-tuning — not because the result is transformative, but because the mechanism design is instructive and the experimental methodology is above average for the subfield.