Imagine you're a building security guard reviewing footage, but half the cameras are off and none of the clips have timestamps. You can still spot someone carrying a crowbar, but you can't distinguish the maintenance worker from the burglar. That's the state of runtime monitoring for tool-using LLM agents: the logic works, but the trace data is too sparse to let it do its job. This paper asks a precise, useful question: can off-the-shelf formal verification — specifically metric first-order temporal logic (MFOTL) running through the MonPoly monitor — catch prompt-injection and tool-misuse attacks in LLM agents, without ever running the agent, just by replaying recorded action traces? The answer is a qualified yes. Five generic temporal obligations (things like 'don't send email without prior user approval') flag 71.8% of STAC attack chains and 70.1% of successful AgentDojo injection attacks. That's meaningful detection from generic, reusable rules rather than application-specific guardrails. The catch is the 29.3% benign-run false-positive rate. The authors are admirably clear-eyed about why: the benchmark corpora almost never record user approvals and never record timestamps. Without that history, a temporal obligation like 'only send email if the user approved it within the last 5 minutes' collapses into a crude 'flag any email send.' The logic is sound; the data starves it. This is the paper's central insight, and it's more engineering diagnosis than algorithmic breakthrough. The most interesting finding is the provenance attack. When traces do carry relational context — say, the source of a piece of data — provenance-aware policies discriminate much better between benign and malicious runs. But an attacker can plant a single line in the data that fools the naive provenance check, defeating it on 94–99% of runs it would otherwise catch. The fix is simple: bind the provenance tag to the lookup operation that produced it, not to the data itself. This closes the evasion completely, with zero cost to detection or false-positive rates. It's a clean result. Methodologically, the paper is offline-only and benchmark-replay-only. No agent was actually run. The authors fed existing recorded trajectories from AgentDojo, STAC, and R-Judge through MonPoly, an existing MFOTL monitor, with no modifications. This is both a strength (reproducible, cheap, no confounds from agent behavior) and a limitation (you're measuring what the benchmarks chose to record, not what a real deployment would expose). The authors know this and make it their main argument: current benchmarks are inadequate for evaluating formal monitors because they omit the fields those monitors need. The paper's concrete deliverable is a twelve-field enforcement-ready trace schema — a proposal for what agent traces should record so that temporal monitors can actually work. Think of it as a minimum-viable black box recorder for LLM agents. Whether the community adopts it is an open question, but the argument for it is well-supported by the data. Taken together, this is a careful diagnostic paper rather than a breakthrough. It demonstrates that formal methods have real value for agent safety — but only if the observability infrastructure catches up. The gap between 'logic works' and 'traces expose enough' is the actual bottleneck, and this paper quantifies it precisely.