Imagine you're a security guard watching 20 monitors at once. You can't stare at all of them in high resolution simultaneously, so you glance at each one every few seconds. That works fine when nothing moves fast — a parked car, an empty hallway. But when someone sprints through a frame, you miss it entirely. Streaming Video Large Language Models face exactly this tradeoff: they have a fixed context budget and must decide how much temporal history to keep, at what spatial resolution, and at what frame rate. FastBench is designed to catch them sprinting. The committed claim: current streaming VLMs cannot perceive high-dynamic events in real-world video streams, and no existing benchmark tests for this because they all focus on slow-moving, low-dynamic scenarios where 1-2 FPS sampling is adequate. FastBench provides 306 human-verified QA pairs across eight domains and six capability types, specifically filtered to exclude anything answerable at 2 FPS. The benchmark's construction pipeline is notably rigorous — it combines automated QA generation from high-FPS clips, filtering of questions answerable at low frame rates, trajectory-based verification using SAM3 and CoTracker3, and three rounds of human inspection. The results are blunt. Gemini-3.5-Flash, the strongest model tested, hits 50.7% — barely above chance on many question types. Qwen3-VL-8B improves from 32.9% at 2 FPS to 44.6% at 24 FPS, a substantial 12 percentage-point lift, but gains saturate as the model is forced to compress older history to accommodate denser recent frames. This is the fundamental tension the paper surfaces: you can look more often, but you pay for it by forgetting more of what you saw. The authors also introduce ProactiveFrame, a training-free baseline that dynamically adjusts incoming frame rates using text tokens as triggers. A dual-tier sliding window retains recent high-FPS observations while downsampling older ones into sparse history. ProactiveFrame beats uniform sparse sampling by 5.4 and 1.5 percentage points on two model variants, but the gap between ProactiveFrame and an oracle-guided approach (which knows exactly when fast events occur) remains large. The models simply cannot determine from the stream alone when finer temporal perception is needed. Architecturally, this work sits at the intersection of vision-language model evaluation and streaming inference. The benchmark targets any VLM that processes video as a token stream under bounded context — this includes transformer-based multimodal models like Gemini and Qwen-VL. The key constraint being exploited is the fixed context window: more frames in means either lower resolution or shorter memory, and the paper shows neither tradeoff is sufficient for fast events. The integrity profile is mixed but honest. The benchmark construction is more rigorous than most — trajectory-grounded verification with SAM3/CoTracker3 plus three human inspection rounds is above average for a benchmark paper. However, the evaluation covers a relatively small set of models, the 306 QA pairs are modest in scale, and there's no independent replication yet. The filtering step (removing questions answerable at 2 FPS) is clever but introduces a selection bias toward the hardest possible cases, which the authors acknowledge implicitly through their temporal scope taxonomy. The milestone this paper points toward is clear: streaming VLMs need to break the 70-80% accuracy range on high-dynamic perception before they can be trusted for real-time applications like autonomous driving, sports analytics, or surveillance. Right now, the best model is at 50.7%. The gap between ProactiveFrame and oracle-guided focusing quantifies exactly how much room remains for adaptive attention mechanisms — and it's a lot. The obvious next experiment the authors didn't run is end-to-end fine-tuning of a VLM on high-dynamic video with a temporally adaptive sampling objective, rather than the training-free heuristic they propose. Most likely reason: compute budget. Training a streaming VLM from scratch or fine-tuning a large one on high-FPS video is expensive, and this paper's contribution is the benchmark, not the model.