ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu
- Published
- Sep 30, 2026 — 17:59 UTC
Problem
This work addresses a significant gap in the capabilities of video scene text editing, particularly focusing on high visual quality, temporal consistency, and edit locality. The authors highlight that existing methods struggle to maintain these attributes simultaneously, which is critical for producing realistic and coherent video edits. The paper is a preprint and has not yet undergone peer review.
Method
The authors present a new dataset comprising 387 real-world 720p videos, which include text-region masks and editing instructions. This dataset is split into 230 videos for training and 157 for evaluation. The evaluation protocol is structured around three axes, utilizing 13 distinct metrics to assess performance. A primary metric is designated for each axis, and a Pareto comparison is employed to analyze trade-offs among the metrics. The core technical contribution is the ViTeX-Edit-14B editor, which is fine-tuned on a paired training split. This model incorporates motion-aligned glyph-video conditioning to enhance the editing process.
Results
The ViTeX-Edit-14B editor achieves a Character Accuracy (CharAcc) of 0.688, which is reported as the highest mean among the evaluated video-native editors. Additionally, it demonstrates the lowest Text-crop Warp among raw editor outputs, indicating a potential advantage in maintaining text integrity during edits. However, the available text does not report quantitative results for other metrics or comparisons against specific baselines beyond these highlights.
Limitations
The authors acknowledge the inherent difficulty in achieving a balance between accurate text representation, temporal stability, and scene preservation. This limitation suggests that while the proposed method shows promise, there are still challenges to overcome in producing edits that are both visually appealing and contextually coherent. The paper does not discuss other potential limitations, such as computational efficiency or scalability of the approach.
Why it matters
The introduction of ViTeX-Bench and the ViTeX-Edit-14B editor has significant implications for downstream work in video editing and computer vision. By establishing a benchmark for high-fidelity video scene text editing, this research paves the way for future advancements in the field, potentially leading to more sophisticated editing tools that can handle complex video scenarios with improved quality and consistency.
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
