Imagine you're a copy editor, but instead of fixing typos in a document, you need to change the text on a moving storefront sign in a movie — while the camera pans, the lighting shifts, and pedestrians walk by. You can't reshoot. You can't freeze the frame. The new text has to look like it was always there, across every single frame. That's video scene text editing, and until now, there was no standardized way to measure whether anyone was doing it well. ViTeX-Bench is a benchmark, not a model — though the authors do release a reference model alongside it. The core contribution is a dataset of 387 real-world 720p videos with text-region masks and editing instructions, split into 230 training pairs (with pipeline-generated ground truth) and 157 frozen evaluation videos. The three-axis evaluation protocol measures text correctness (does the OCR read the right characters?), visual/temporal quality (does it look good and stay consistent across frames?), and edit locality (did the edit leave the rest of the scene alone?). Thirteen metrics total, one primary per axis, with Pareto analysis of trade-offs. The claim is specific and honest: this is the first benchmark suite that provides paired real-video data and a protocol designed to directly measure whether the requested text remains correct over time. Previous resources offered limited paired data, and general video-editing metrics don't capture character-level text accuracy across frames. The authors evaluated eight baselines spanning four editing families — and the headline result is that no method simultaneously achieves accurate text, temporal stability, and scene preservation. Their reference editor, ViTeX-Edit-14B, uses motion-aligned glyph-video conditioning and achieves a CharAcc of 0.688 — the highest among video-native editors tested. It also achieves the lowest text-crop Warp Error among raw editor outputs, meaning its text is the most temporally stable. But 0.688 character accuracy means roughly one in three characters is wrong. The bar is not cleared; the bar is measured. This is the distinction that matters. The integrity setup is relatively strong for a benchmark paper. The evaluation split is frozen (157 videos), OCR calibration is performed, human evaluation supplements automated metrics, and annotation-sensitivity analysis examines how mask quality affects scores. The authors acknowledge limitations in their pipeline-generated pairs for training (they're reviewed but not perfect). Code and data are released via a project page. Notably, the paper is accepted to the NeurIPS 2026 Datasets and Evaluations Track — peer review specifically geared toward benchmark rigor. Architecturally, the reference editor sits in the conditional video diffusion family — a 14B-parameter model fine-tuned on the paired training split with glyph-video conditioning that aligns text renderings with motion. The four baseline families tested span image-based editors applied frame-by-frame, video diffusion editors, instruction-following editors, and hybrid approaches. The Pareto analysis is the most useful part: it shows that methods optimizing for text accuracy tend to sacrifice temporal consistency, and vice versa. The field fight here is whether video text editing should be treated as a solved sub-problem of general video editing or as a distinct task requiring dedicated benchmarks. ViTeX-Bench argues forcefully for the latter — general video metrics miss character-level correctness entirely. The milestone to watch is CharAcc breaking 0.90 on this benchmark while maintaining competitive temporal quality scores. That would signal readiness for real-world applications like sign translation in video localization.