Imagine you're a detective reconstructing a crime scene from photographs. You can tell roughly where the furniture was, but if someone asks you the exact distance between the couch and the window to the centimeter, your confidence collapses. That's exactly what happens when you ask multimodal AI models to read numbers off scientific plots — they're decent sketch artists but terrible accountants. PlotGround addresses a foundational gap in how we evaluate AI's ability to extract quantitative data from scientific figures. The committed claim: existing benchmarks for plot digitization are either synthetic (drawn by the benchmark authors, not real scientists) or narrow (covering few chart types), so we have no honest picture of how well models actually recover values from real published figures. PlotGround is an automated pipeline that pairs real bioRxiv figures with their author-released source data, generating questions with ground-truth answers that come from the data itself, not from human annotation. The resulting benchmark, PlotGround-1k, contains 1,119 questions drawn from 1,066 bioRxiv preprints, verified by humans. The ladder results are revealing. Sixteen multimodal models were tested. The best hits 87.5% accuracy at a ±5% relative-error tolerance — meaning it gets within 5% of the true plotted value roughly seven out of eight times. But tighten that tolerance to ±2% and every single model drops 11-24 percentage points. This is the key finding: current vision-language models are doing approximate visual reading, not precise quantitative recovery. They're reading the room, not the ruler. The architecture story here is multimodal vision-language models (think GPT-4o, Gemini, Claude with vision) being used as chart readers. The method doesn't introduce a new model — it introduces a new evaluation regime. The pipeline itself uses figure-table matching, panel identification, and question generation to automate benchmark construction. The most structurally interesting comparison: when you give a coding agent the source table instead of the figure, accuracy jumps from 90.0% to 97.4% while cutting cost by 72%. The figure is the bottleneck, not the reasoning. Integrity is a genuine strength here. The ground truth comes from author-released source data, not annotator guesses. The benchmark covers real scientific figures across the diversity of bioRxiv, not cherry-picked clean charts. Human verification adds a second layer. The paired figure-vs-table comparison is an unusually clean ablation — same values, different input modality, same agent. That said, the 1,066-paper sample is bioRxiv-only, which skews toward life sciences chart conventions. We don't know if these results generalize to physics plots, engineering schematics, or economics charts. The milestone question is straightforward: can models close the gap between ±5% and ±2% tolerance without specialized fine-tuning? Right now the delta is 11-24 percentage points across all sixteen models. Getting that delta below 5 points — meaning models are nearly as accurate at ±2% as at ±5% — would signal that vision models have crossed from approximate reading to genuine digitization. That's probably 1-2 model generations away, given current trajectory. The obvious successor experiment is fine-tuning a vision model specifically on figure-to-data extraction using PlotGround's pipeline as a training data generator. The authors built the data factory but didn't train a specialized model on its output. Most likely read: they're saving it for the next paper. The pipeline is the contribution; the model is the sequel.