Imagine you're a sous chef and the head chef says: "Sear the steak, then once the asparagus arrives from the walk-in, blanch it, and keep the hollandaise warm the whole time." That's three tasks — one immediate, one event-triggered, one persistent — and you have to hold all three in your head while the kitchen evolves around you. Current autonomous driving planners are like a sous chef who can only remember the last five seconds of instructions. doPlan is the first dataset that writes down the full multi-stage ticket. The committed claim: doPlan is, to the authors' knowledge, the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context for autonomous driving. Built on top of Motional's nuPlan platform, it contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned driving context across 50.9 hours of unique driving data. Annotation windows range from 30 seconds to 508.8 seconds — an order of magnitude beyond the typical 5-second planning horizon used by most language-conditioned driving models. The ladder result is damning for existing models. The authors evaluate four language-conditioned driving architectures and find that "sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction." Among 2,161 examples containing a matched future maneuver, the median time to the first relevant maneuver is 24.6 seconds after the evaluation point, and only 9.8% of those maneuvers fall within the standard 5-second prediction window. This is not a marginal miss — it's a structural mismatch between what passengers actually say and what planners can see. Architecturally, doPlan itself is a dataset and benchmark, not a model. The contribution sits in the annotation ontology: five categories of passenger intent (immediate, deferred, event-conditioned, persistent, and multi-stage) mapped onto real sensor logs with human-written natural language. The evaluation pipeline feeds these instructions to existing language-conditioned planning models and measures whether the model's planned trajectory aligns with the intended maneuver direction. The four tested models represent the current crop of language-to-trajectory architectures — the paper exposes their shared blind spot rather than proposing a fix. The integrity picture is solid for a dataset paper. Annotations are human-written (not LLM-generated), the underlying driving data comes from nuPlan (a well-known community resource), the dataset and annotation interface are publicly released on GitHub, and the evaluation is transparent about what the models fail at. The weakness is that the four evaluated models are not named with full SOTA context — we know they fail, but the paper positions itself more as a diagnostic than a horse race. No pre-registration, but that's standard for dataset contributions. The milestone framing is clear: current planners operate on a ~5-second horizon; 90.2% of real passenger intent requires reasoning beyond that window. The next concrete threshold is a planner that can retain and ground unresolved goals across at least 25 seconds (the median maneuver delay), ideally across the full annotation windows of 30–509 seconds. No existing model demonstrated this capability. The field needs to move from reactive trajectory prediction to persistent goal tracking — and doPlan provides the first public benchmark to measure progress. The obvious experiment not run: training or fine-tuning a model on doPlan's multi-stage annotations and showing that persistent-context architectures (memory-augmented transformers, hierarchical planners, or explicit goal stacks) actually close the gap. The authors built the diagnostic but stopped short of the treatment. Honest read: this is likely scoped as a dataset paper for a venue submission, with the model paper coming next. The infrastructure is in place — the training loop is the sequel.