Imagine you're researching a used car. The standard approach is like using a search engine where you can only retype queries — you find a promising listing, but when you go back to refine your search, you've lost the tab. Every follow-up starts from scratch. Now imagine a spreadsheet where every listing you've ever opened stays in a live workspace, you can write formulas that cross-reference price against mileage against seller history, and you choose which cells to inspect next based on what the formulas surface. That's the structural shift this paper proposes: from query-as-action to program-cell-as-action. The committed claim: search agents that can execute local computations over retrieved candidates — not just reformulate queries — achieve meaningfully higher task success (4.0 and 7.56 percentage points on InfoSeek-Eval and BrowseComp-Plus) while using roughly 30% fewer tokens at the final step. The paper's diagnostic insight is sharp: trajectory analysis reveals that supporting passages are often retrieved but never shown to the agent because the fixed search interface doesn't surface them. A same-page oracle experiment confirms that when the right evidence is presented, agents issue fewer subsequent queries. The bottleneck isn't retrieval — it's evidence routing. PSA (Programmatic Search Agent) introduces three structural elements: a persistent candidate workspace that survives across turns, composable primitives (filter, rank, extract, compare) that the agent chains into executable program cells, and selective evidence presentation that lets the agent decide what it actually reads. The runtime resolves data dependencies within each cell while the agent adapts strategy across cells. This is closer to a notebook-style execution environment than a traditional tool-use wrapper. The experimental design is more careful than typical agent papers. Three agent interfaces — Query-based, Tool-based, and PSA — share the same underlying search substrate. Critically, the Tool-based Agent also shares PSA's primitives and persistent workspace, isolating the contribution of programmatic composition and selective evidence presentation. Five policy backbones are tested without task-specific training. This controlled comparison is the paper's strongest integrity feature. The gains are real but bounded. On InfoSeek-Eval, PSA improves macro-averaged task success by 4.0 percentage points over the Query-based Agent. On BrowseComp-Plus — a harder benchmark — the gap widens to 7.56 points. Token reductions at the final step average 28.3% and 33.9% respectively. These are meaningful efficiency gains, not just accuracy bumps. But the paper doesn't report wall-clock latency for the program-cell execution overhead, which matters for deployment. The architectural lineage here runs through ReAct-style agents, tool-augmented LLMs (Toolformer, Gorilla), and retrieval-augmented generation pipelines. PSA's contribution is making the post-retrieval processing stage itself an agentic action space rather than a fixed pipeline step. The persistent workspace echoes scratchpad and chain-of-thought approaches, but the key move is treating evidence selection as a first-class decision the agent controls. What's missing is the scaling story. Five backbones are tested, but all appear to be prompted without fine-tuning. The obvious next experiment — fine-tuning a backbone specifically for program-cell generation — would test whether PSA's gains compound or plateau when the policy is optimized for the interface. The authors likely ran out of compute budget or are reserving this for a follow-up, given the "code will be released subject to approval" caveat that signals corporate constraints on the research.