Skip to content

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

Conference: NeurIPS 2026 โ€” Evaluations and Datasets Track
arXiv: 2609.40356
Code: https://github.com/taco-group/ViTeX-Bench
Area: Video Generation / Video Scene Text Editing / Datasets & Evaluation
Keywords: scene text editing, glyph-video conditioning, character correctness, temporal consistency, edit locality

TL;DR

ViTeX-Bench uses paired training data from real videos, a frozen evaluation split, and 13 metrics across three axes to distinguish text correctness, motion stability, and background preservation, while providing a motion-aligned glyph-conditioned reference editor, ViTeX-Edit-14B, with the highest mean character accuracy among the evaluated video-native editors rather than the best performance on every metric.

Background & Motivation

Video generation models can produce coherent scenes, but accurately replacing one string in a scene with another demands more than a plausible appearance. Characters must be correct in every frame, remain attached to the original moving surface, and leave surrounding objects, illumination, and camera dynamics unchanged. Image text editors such as FLUX-Text can render correct characters in individual frames but lack temporal coupling; first-frame editing followed by AnyV2V propagation can introduce progressive text drift; video editors such as VACE, VideoPainter, and Kling model motion but do not necessarily control exact character sequences.

Existing video-quality and instruction-following metrics do not directly establish whether the requested string remains visible throughout a clip. An editor can preserve the source text and obtain strong stability and locality scores, or generate polished imagery with incorrect characters. Paired edits of real videos are also expensive to create: perspective, occlusion, and motion change the text region over time, manual framewise editing is costly, and the instruction does not determine one uniquely correct font, color, or lighting interaction.

The paper therefore prioritizes a reusable resource and measurement protocol, then demonstrates the training data with a reference editor. Core Idea: construct training references through a model-assisted, human-reviewed pipeline and measure character correctness, visual and temporal quality, and edit locality separately, preventing stable but incorrect outputs from being mistaken for successful edits.

Method

Overall Architecture

Each task supplies a source video, per-frame text-region masks, a source string, and a target string; the output should change text only in the specified region. ViTeX-Dataset contains 387 real source videos: 230 include reviewed synthetic edited references for training, while 157 form a permanently frozen evaluation split without edited references. Standard clips have 1280ร—720 resolution, 120 frames, and 24 fps, lasting 5 seconds; training references are not unique ground-truth renderings.

The system comprises four parts: Paired Data Construction, Motion-Aligned Glyph Conditioning, Shared Composite Control, and Three-Axis Evaluation. The first two support reference-editor training and inference; Shared Composite Control is optional deterministic post-processing; Three-Axis Evaluation independently reads the source, masks, strings, and outputs without participating in optimization. Dashed edges indicate training supervision or measurement dependencies, whereas solid edges represent editing-inference and post-processing data flow.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Source, masks<br/>source and target strings"] --> B["Paired Data<br/>Construction"]
    A --> C["Motion-Aligned<br/>Glyph Conditioning"]
    B -.->|230 paired references| T["Supervised fine-tuning"]
    T -.->|Trained editor| C
    C --> R["Raw edited output"]
    R --> D["Shared Composite<br/>Control"]
    R -.->|Raw measurement| E["Three-Axis Evaluation"]
    D -.->|Separate Composite measurement| E
    A -.->|Frozen evaluation inputs| E

Key Designs

1. Paired Data Construction: create learnable references through first-frame editing and video propagation

The authors curate real scene-text clips from Panda-70M and InternVid and construct four assets. An annotator prompts SAM 3 on the first frame, propagates masks, corrects drift, and dilates them three times with a 25ร—25 elliptical kernel to leave editing margins around character boundaries. Qwen3-VL-32B-Instruct reads the first-frame text and proposes a similar-length replacement, followed by human review. removal-1.3B removes source glyphs and their associated shadows and highlights to produce a clean background; Gemini 3 Pro Image rewrites the first frame, and SAM 3 isolates the target-text patch. The reference is therefore not simply a line of rendered type pasted into a scene: an image editor first supplies text with scene-compatible appearance.

For static text, Strategy A alpha-composites the same patch into every background frame; for moving text, Strategy B uses a PISCO inserter to propagate the first-frame reference. Static clips run both strategies and retain the better output, while dynamic clips use B exclusively. Consequently, the final 56 A and 174 B videos count construction strategies, not static and dynamic clips. A separate coverage audit reports approximately 25% static and 75% dynamic videos. PISCO is additionally fine-tuned on auxiliary scene-text videos disjoint from these 387 clips, using text-removed videos as inputs, original videos as reconstruction targets, and amodal-completion supervision in the text region.

The training split provides source videos, edited references, masks, and string pairs; the evaluation split provides editing inputs only. Review does not eliminate synthetic errors: reference edits have OCR exact-match accuracy 0.585 and CharAcc 0.790, while human transcription of calibration crops reaches CharAcc 0.917 and 97% are judged readable. This indicates that OCR can underestimate the readability of generated text, but also that the references should not be treated as error-free, unique ground truth.

2. Motion-Aligned Glyph Conditioning: deliver target-character structure along the source-text trajectory

ViTeX-Edit-14B builds on Wan2.1-VACE-14B. Its existing conditions are target text encoded by frozen uMT5-XXL and source video plus masks handled by the Video Condition Unit (VCU); the new third stream is a target-text glyph video rather than a linguistic description. For Latin text, Qwen3-VL selects a close source-font match from a font library; non-Latin text uses script-specific default fonts. The target string is rendered white on black, EasyOCR detects the first-frame text quadrilateral, CoTracker3 tracks it through the remaining 119 frames, and per-frame homographies make the target glyphs follow the source position, scale, and perspective.

A frozen Wan VAE encodes the glyph video, and a patch embedding with stride (1,2,2) converts the latent into tokens. Cross-attention pooling with 64 learnable queries produces a fixed-length glyph-token bundle; added condition cross-attention lets every VACE block read it through a residual path. Zero-initialized output projections preserve the original backbone output at initialization. The rationale is task-specific: language conditioning may describe the target string without specifying its strokes or framewise placement, whereas glyph video provides both character structure and a motion reference.

This is not a training-free plug-in. Qualitative pilots using glyph video only at inference, or routing it through the existing VCU during fine-tuning, showed poor character correctness and motivated the dedicated branch. However, controlled component ablations were not completed, so the aggregate improvement cannot be assigned precisely to font selection, tracking, 64-query pooling, or the additional attention layers individually.

3. Shared Composite Control: separate regional synthesis from source-background restoration

Full-frame video editors reconstruct background that should remain unchanged, whereas local crop editors such as TextCtrl and RS-STE directly copy many source pixels. Locality comparisons can therefore reflect interface differences. Composite applies the same post-processing to all eight baselines and the reference editor: an annulus outside the mask supplies prediction/source LAB means and standard deviations for color correction, after which a 4-pixel feathered boundary blends the edited region onto the source frame before common encoding.

The wrapper is training-free and requires no GPU; it is not a learned module exclusive to the reference editor. Before encoding, exterior pixels come directly from the source except in a narrow boundary band. Large locality gains therefore primarily demonstrate the effect of source restoration, not improved learned background preservation. Raw and Composite outputs must be reported separately rather than mixing post-processed scores into raw-editor rankings. Baseline Composite OCR was checked only on a sample; complete text and text-crop temporal re-scoring was performed only for the ViTeX-Edit-14B output pair.

4. Three-Axis Evaluation: inspect correct text, stable imagery, and localized changes independently

The text axis uses PP-OCRv5, followed by NFKC normalization, case folding, and whitespace/punctuation removal. Source readability is gated first: a source-frame OCR similarity of at least 0.5 to the source string places that frame in the detectable set. Five evaluation clips have no such frames, leaving 152 supported videos for SeqAcc and CharAcc; TTS additionally requires at least one consecutive detectable-frame pair. Scores are calculated over frames within each video and then averaged over supported videos, not pooled into one accuracy across all frames.

The central distance is substring edit distance: the minimum number of edits needed to turn the target into any contiguous substring of the OCR candidate, with extra candidate prefixes and suffixes free. The following defines character similarity and complete target-substring accuracy; \(r\) is the reference, \(c\) the candidate, and \(\mathcal D\) the source-detectable frame set.

\[ \mathrm{Sim}(r,c)=1-\frac{d_{\mathrm{sub}}(r,c)}{\max(|r|,1)},\qquad \mathrm{SeqAcc}=\operatorname*{mean}_{t\in\mathcal D}\mathbf 1[d_{\mathrm{sub}}(s_{\mathrm{tgt}},\hat s_t)=0]. \]

CharAcc averages target-string similarity over detectable frames, granting partial character credit. SeqAcc requires the complete target as a substring but does not establish the absence of extra text. TTS measures the fraction of adjacent detectable pairs whose decoded strings are identical; consistently incorrect text can score highly, so TTS must be read with the other two measures.

The visual and temporal axis has six metrics: Flicker, Warp, and MUSIQ at both full-frame and text-crop scopes. Flicker is the mean absolute difference between adjacent output frames; Warp aligns adjacent output frames using RAFT flow from the source video before measuring error, rather than estimating flow from edited outputs; MUSIQ is a no-reference perceptual-quality estimate. The text crop is a fixed bounding box enclosing the spatial union of all per-frame masks with a 16-pixel margin, shared across the entire clip to avoid crop-window jitter. Full-frame MUSIQ uses all frames, crop MUSIQ uses source-detectable frames, and Flicker/Warp use all adjacent frames. For large text trajectories the fixed crop may still include substantial background; smoothing or nearly constant outputs can also reduce temporal errors.

The locality axis contains PSNR, SSIM, LPIPS, and DreamSim. For measurement, predicted pixels inside the mask are replaced with source pixels, leaving only exterior prediction differences for comparison with the source. The following locality-only composite is a measurement construct, not the actual Composite output wrapper described above.

\[ \hat f_t^{\mathrm{loc}}=(1-m_t)\odot\hat f_t+m_t\odot f_t. \]

Higher PSNR and SSIM are better; lower LPIPS and DreamSim are better. The three primaries are SeqAcc, text-crop Warp, and DreamSim-loc; the remaining ten metrics offer diagnostic detail, without a weighted aggregate. Pareto membership means that no other method is at least as good on all three and strictly better on one. It is not a certificate of editing success: VACE with zero SeqAcc can remain on the front through stability and locality.

A Worked Example

Consider a 5-second whiteboard clip explicitly labeled as research material, replacing the neutral text โ€œLABโ€ with โ€œMAPโ€ while the camera moves slowly. This is an explanatory example, not a paper evaluation sample or an additional result.

Per-frame masks identify the original text region, and tracking follows the first-frame quadrilateral. Rendered โ€œMAPโ€ glyphs change position and perspective with the whiteboard in a 120-frame glyph video. The reference editor consumes glyphs, target text, source video, and masks to produce raw outputs; optional Composite blends only the synthesized region back into the original clip, and the two outputs are evaluated separately.

If OCR reads โ€œAMAPBโ€ in a frame, substring distance is still zero, giving that frame a complete-match score and exposing the lack of an extra-character penalty. If every frame consistently reads โ€œLAB,โ€ TTS can remain high although the requested edit failed. Correct characters drifting off the whiteboard instead expose a different failure through source-flow-aligned text-crop Warp. Frames excluded because source OCR is unreadable do not count toward correctness, so the protocol cannot establish correctness on every frame of the clip.

Loss & Training

Paired references supervise Flow-Matching fine-tuning. The main DiT trunk, uMT5-XXL, and Wan VAE remain frozen, while the VACE branch, glyph encoder, and condition cross-attention are updated. OCR metrics are evaluation measures here, not training losses, and the paper does not supply a special character-reward formula that needs reconstruction.

Stage 1 trains at 720p and 49 frames for 5 epochs, with learning rate \(5\times10^{-5}\), AdamW weight decay 0.01, and 10-fold dataset repetition. Stage 2 continues at 720p and 121 frames for 2 epochs with learning rate \(1\times10^{-5}\). Both use effective batch size 64 on 8 H100 80GB GPUs, gradient accumulation, checkpointing, and CPU offload. Approximately 22 and 50 hours respectively amount to 576 GPU-hours for this reference-editor fine-tuning alone, excluding upstream pretraining, paired-data generation, and auxiliary PISCO fine-tuning. Inference uses one 50-step generation pass.

Key Experimental Results

Main Results

The following selection from Table 2 covers raw outputs from all four baseline families and the reference editor. Higher SeqAcc/CharAcc and lower text-crop Warp/DreamSim-loc are better; correctness has 152-video support, while other metrics use their respective support sets.

Method / Editing family SeqAcc CharAcc TTS Text-crop Warp DreamSim-loc
Source video, reference only 0.000 0.317 0.760 1.27 0.000
AnyText2 / Per-frame 0.280 0.633 0.382 3.95 0.043
TextCtrl / Per-frame 0.475 0.734 0.511 2.09 0.004
FLUX-Text / Per-frame 0.528 0.737 0.326 13.01 0.012
RS-STE / Per-frame 0.354 0.626 0.534 1.81 0.007
TextCtrl + AnyV2V / First-frame propagation 0.057 0.308 0.257 3.97 0.073
Wan2.1-VACE-14B / Video inpainting 0.000 0.298 0.689 1.56 0.007
VideoPainter / Video inpainting 0.364 0.619 0.606 Unranked 0.024
Kling Video 3.0 Omni / Instruction-guided editing 0.000 0.208 0.641 2.90 0.061
ViTeX-Edit-14B / Reference editor 0.341 0.688 0.648 1.53 0.024

VideoPainter is adapted from native 8 fps and 49 frames to 24 fps and 120 frames through repeated padding, linear interpolation, and spatial resizing. These operations mechanically alter adjacent-frame residuals, so the paper excludes its Flicker and Warp from temporal rankings and its outputs from the three-primary Pareto comparison. The unedited Source video is also excluded from editor rankings.

ViTeX-Edit-14B exceeds VideoPainter in CharAcc by 0.069 but trails it in SeqAcc by 0.023. Table 3 gives 95% video-bootstrap intervals of CharAcc 0.688 [0.644, 0.732] versus 0.619 [0.559, 0.676], and SeqAcc 0.341 [0.266, 0.413] versus 0.364 [0.300, 0.434]. These intervals overlap, and the paper does not establish pairwise significance. โ€œHighestโ€ refers to mean character accuracy among the evaluated video-native editors, not to the stronger per-frame TextCtrl and FLUX-Text results.

Ablation Study

The paper does not provide controlled glyph-component ablations. The following shared post-processing analysis from Table 2 and Appendix H is not a substitute for network-design ablation. The raw and Composite ViTeX-Edit-14B outputs were fully re-scored.

Metric Raw Composite Note
SeqAcc 0.341 0.345 Complete target-substring matching changes only slightly
CharAcc 0.688 0.689 Character accuracy remains close
PSNR-loc, dB 29.08 42.95 Source restoration supplies the principal locality gain
DreamSim-loc 0.024 0.002 Exterior perceptual differences decline substantially
Text-crop Warp 1.53 1.56 Post-processing does not reduce this temporal error

The same Composite wrapper yields approximately 43 dB PSNR-loc across all eight baselines, showing that the locality benefit spans models. Baseline post-processing OCR was checked only on a sample; the reported \(|\Delta|\leq0.04\) does not establish unchanged correctness over the full evaluation split and cannot support a complete Composite ranking.

Measurement validity is also calibrated through Table 10's study of three non-author raters and 70 outputs. Human scores range from 1 to 3; correlations describe agreement between metrics and human judgments on this sample, not significance tests between editors.

Axis Inter-rater agreement, ordinal Krippendorff \(\alpha\) Automatic primary Spearman \(\rho\) with human ratings
Text correctness 0.87 SeqAcc +0.71
Temporal quality 0.80 Text-crop Warp -0.40
Edit locality 0.37 DreamSim-loc -0.53

Key Findings

  • FLUX-Text obtains SeqAcc 0.528 but text-crop Warp 13.01: framewise character accuracy does not imply temporally stable text editing.
  • Source video has TTS 0.760 and SeqAcc 0; Kling has full-frame MUSIQ 72.23 and SeqAcc 0. Stability, visual quality, and successful replacement are distinct propositions.
  • The raw three-primary front contains FLUX-Text, TextCtrl, RS-STE, ViTeX-Edit-14B, and VACE. VACE's membership with zero SeqAcc reiterates that the front is not a pass threshold.
  • Across 1,000 resamples of the 152 supported videos, mean Kendall \(\tau=0.936\) and the leading method is retained 95% of the time. This supports in-domain ranking stability, not broader scene coverage.
  • On the seven-clip non-Latin slice, AnyText2 achieves SeqAcc 0.168 and CharAcc 0.295, while the other eight editors all have CharAcc below 0.19. The sample is too small for per-language conclusions.

Highlights & Insights

  • Keeping an unedited Source reference is particularly informative. It directly demonstrates that high stability and locality can coexist with zero target matching, forcing evaluation to verify semantic success.
  • Fixed mask-union crops and source flow address different confounds: the former removes crop-window jitter, while the latter checks output motion against the source trajectory. Both still require joint interpretation with character correctness.
  • Glyph conditioning becomes a moving video reference rather than a static stroke hint. Coupling the requested characters with their expected surface motion may inform other video-editing tasks requiring precise local structure.
  • A shared Composite wrapper controls for background-preservation mechanisms. It exposes the benefit of copying source pixels instead of attributing the entire gain to the learned editor.

Limitations & Future Work

  • Coverage is small and favors readable, localized Latin text: the evaluation split has Latin 150, Chinese 4, Japanese 1, and Cyrillic 2 clips. Dense layouts, curved surfaces, severe occlusion, and extreme motion remain underrepresented.
  • The reference editor depends on its font library, first-frame OCR detection, quadrilateral tracking, and projective model. Non-Latin defaults may mismatch source style, and tracking failures can corrupt the condition; controlled ablations and mismatch tests should isolate these dependencies.
  • Source-detectability gating reduces unreadable-input confounds but excludes the hardest frames and five clips from correctness support. Substring distance also leaves surrounding extra text unpenalized; strict whole-string accuracy and occlusion-stratified results would complement the protocol.
  • Human transcription uses one author, and the independent mask audit covers only 12 clips with the same propagation pipeline. Locality inter-rater agreement is only 0.37, requiring caution when interpreting absolute perceptual scores.
  • Model-assisted references may retain errors, and the initial release lacks per-record target rejection, resampling, and first-frame retry histories. Expanded provenance and independent review would be more informative than synthetic scale alone.
  • The dataset is CC-BY-NC 4.0 for non-commercial research, with upstream licensing obligations still applying. Apache-2.0 code and model licensing do not make the data commercially usable; examples should use explicitly labeled research material, not misleading or evidentiary alteration.
  • vs STRIVE: Both address video scene-text replacement. STRIVE emphasizes still-image edits, propagation, and region-centric evaluation; this work adds paired training resources, a frozen real-video split, and joint character, temporal, and exterior-preservation diagnostics.
  • vs GlyphMastero: Image glyph encoding supplies stroke priors, whereas this work expands it into motion-aligned glyph video injected into VACE. The temporal adaptation is useful, but controlled experiments have not separated individual component contributions.
  • vs EditBoard / FiVE / IVEBench / VEFX-Bench: General editing suites evaluate instructions, quality, and preservation; ViTeX-Bench makes framewise correctness of a specified string operational. The protocols complement rather than fully replace one another.
  • vs per-frame text editors: TextCtrl and FLUX-Text are stronger at exact characters, while video-native editors more readily preserve temporal structure. Combining them is a research direction, but it requires re-evaluating all three axes rather than optimizing one favorable score.

Rating

  • Novelty: 4/5 โ€” Systematizes real-video text-replacement resources and character-level three-axis evaluation; the reference architecture largely adapts existing glyph-conditioning ideas.
  • Experimental Thoroughness: 4/5 โ€” Eight baselines, confidence intervals, shared post-processing, and human calibration provide useful coverage, with remaining gaps in scale, non-Latin coverage, and controlled ablation.
  • Writing Quality: 4/5 โ€” Metric support, adaptation exceptions, and success boundaries are clearly specified; reproduction costs and annotation provenance still need fuller accounting.
  • Value: 4/5 โ€” Offers a reusable task protocol, particularly for detecting stable but incorrect edits, with substantial room for real-world generalization.