From Script to Shot: A Benchmark for Grounding Screenplays in Movies¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/jungucho92/script2shot
Area: Video Understanding / Segmentation
Keywords: screenplay-to-shot alignment / video grounding / movie narrative understanding / multimodal retrieval / scene segmentation
TL;DR¶
Introduces "From Script to Shot", the first benchmark spanning 50 feature films over nine decades with 55K+ human-verified scene-to-shot alignments, uncovering the dialogue anchoring effect in multimodal grounding and demonstrating strong downstream utility in zero-shot scene segmentation and guided text-to-video generation.
Background & Motivation¶
Multimodal video-language understanding has advanced rapidly in recent years, yet prevailing progress remains heavily driven by short video datasets such as MSR-VTT, VATEX, HowTo100M, and WebVid. These datasets are predominantly comprised of second-scale clips paired with brief, literal descriptive captions, offering virtually no coverage of long-range temporal dependencies or the hierarchical narrative structures inherent to complete storytelling. Feature films naturally manifest this long-form narrative structure: a film is organized into scenes defined by continuity of time, location, or action, which are visually realized on screen through a continuous sequence of hundreds or thousands of cinematic shots. The screenplay serves as the foundational narrative blueprint, uniquely combining spoken dialogue with scene headings, character tags, stage directions, and camera movements.
However, existing attempts to align screenplays with video have suffered from substantial structural limitations. Early alignment approaches, such as Cour et al.'s Movie/Script, rely almost exclusively on matching external subtitle transcripts with screenplay dialogue lines, reducing the task to surface-level lexical overlap and completely failing on non-dialogue scenes. Conversely, modern vision-language models and video-text retrieval encoders (e.g., CLIP, CLIP4Clip, DiffusionRet) as well as temporal grounding networks (such as TAN and MATR) are trained on homogeneous, single-action clip-caption pairs. When directly transferred to two-hour full-length movies characterized by complex heterogeneous stage directions and non-dialogue storytelling, these models experience catastrophic domain shifts and scale mismatches. A systematic, full-length benchmark that isolates the specific contributions of dialogue matching versus visual grounding has been sorely missing.
Addressing the ambiguous, one-to-many mapping between screenplay scenes and filmed shots, as well as real-world editorial post-production modifications (e.g., deleted or reordered scenes), the authors curate and rigorously validate 50 feature films from MovieNet spanning nine decades and 16 genres. The core idea is to establish the first benchmark of 55K+ human-verified shot-level scene alignments across feature films, introduce a standardized evaluation protocol that disentangles dialogue matching from visual grounding, and uncover the dialogue anchoring effect that enables sparse text matches to steer non-dialogue visual grounding.
Method¶
Overall Architecture¶
The "From Script to Shot" benchmark formalizes screenplay-to-shot grounding as a sequence-level monotonic assignment problem. The inputs consist of an ordered sequence of screenplay scenes \(S = \{s_1, \dots, s_M\}\) and an ordered sequence of extracted video shots \(V = \{v_1, \dots, v_N\}\). Each screenplay scene is structured into typed elements including scene headings (\(\text{H}\)), narrative stage directions (\(\text{N}\)), character names (\(\text{C}\)), and dialogue (\(\text{D}\)). Each video shot is represented by three keyframe images and any temporally overlapping subtitle text lines.
The benchmarking protocol operates under a unified two-stage pipeline. In the first stage, any candidate model (sparse lexical matching, contrastive vision-language encoders, temporal grounding networks, or video-text retrieval models) computes a cross-modal similarity matrix \(S \in \mathbb{R}^{N \times M}\) between all shots and scenes under its respective modality configuration. In the second stage, a shared assignment algorithm decodes this similarity matrix into a globally consistent, monotonic alignment mapping \(\phi: \{1, \dots, N\} \to \{1, \dots, M\}\) using Drop-DTW paired with a Gaussian length prior. This strict two-stage decoupling ensures that all empirical variations accurately reflect representation and similarity modeling quality rather than disparate search heuristics.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Screenplay Text + Movie Video Shots"] --> B["Structured Parsing & Temporal Projection"]
B --> C["Decoupled Cross-Modal Similarity"]
C --> D["Temporally Constrained Sequence Alignment"]
D --> E["Diagnostic Evaluation & Downstream Applications"]
Key Designs¶
1. Structured Parsing & Temporal Projection: Constructing Standardized Multimodal Inputs
Raw screenplays feature irregular formatting and lack standardized scene-level divisions, while unaligned video files lack frame-rate metadata. To resolve these formatting and timing discrepancies, the benchmark applies a screenplay parser to detect scene headings, splitting each script into distinct scene units \(s_j = \{(t_k, x_k)\}_{k=1}^{L_j}\) where \(t_k \in \{\text{H}, \text{N}, \text{C}, \text{D}\}\) classifies headings, directions, character tags, and spoken lines. On the video side, because MovieNet omits native frame-rate metadata, the authors leverage a subset of known shot-level subtitle matches to back-estimate the precise subtitle timecode offset and video frame rate for each individual film. All subtitle transcripts are then projected onto discrete shot intervals, yielding a unified representation for each shot \(v_i = (f_i, u_i)\) with keyframe images \(f_i\) and associated subtitle text \(u_i \in \{u, \emptyset\}\). This preprocessing isolates dialogue and visual channels cleanly across the entire movie.
2. Decoupled Cross-Modal Similarity: Benchmarking Fourteen Approaches Across Four Paradigms
To systematically examine where existing vision-language models succeed and fail, the benchmark establishes a standardized similarity evaluation across four distinct paradigms: - Dialogue-based Lexical Matching: Treats scene dialogues as documents and shot subtitles as queries using BM25 and TF-IDF sparse vector representations, alongside reproducing the classic token-level DTW and Jaccard overlap approach M/S; - Contrastive Visual Encoders: Ignores all subtitles and encodes keyframes against scene narrative descriptions and text chunks using CLIP, SigLIP, InternVideo2, FG-CLIP 2, and Qwen3-VL-Embedding via cosine similarity; - Temporal Grounding & Video-Text Retrieval: Evaluates TAN, MATR, T-MASS, CLIP4Clip, DiffusionRet, and DiscoVLA to test whether models trained for moment localization in short clips transfer to feature-length narratives; - Adaptive Multimodal Fusion: When a shot contains subtitle text (\(u_i \neq \emptyset\)), the row-normalized BM25 similarity and visual similarity are combined via a convex combination; when subtitles are absent (\(u_i = \emptyset\)), the system falls back entirely onto visual similarity:
The fusion parameter \(\alpha\) is tuned via 5-fold movie-level cross-validation, ensuring that visual semantics complement lexical matching without over-relying on either modality.
3. Temporally Constrained Sequence Alignment: Robust Dynamic Programming with Drop-DTW
During film post-production and editing, screenplay scenes are frequently reordered, truncated, or dropped entirely (approximately 20% of screenplay scenes have no realized shots in the final cut). Applying standard dynamic time warping (DTW) under these conditions causes catastrophic boundary drift due to forced false-positive assignments. The second stage addresses this by deploying Drop-DTW, which natively supports dropping unmatched outlier scenes without penalty. Furthermore, to prevent alignment wandering over two-hour sequences, the decoder incorporates a Gaussian positional length prior proportional to cumulative scene text length. This enforces monotonic temporal progression and yields a robust shot-to-scene assignment mapping \(\phi\).
Loss & Training¶
The benchmark primarily evaluates zero-shot and off-the-shelf transfer capabilities of existing foundation models and retrieval architectures; no task-specific end-to-end backpropagation is performed on the movie test set. The fusion weight \(\alpha\) is selected strictly through movie-level 5-fold cross-validation to prevent test-set leakage. Evaluation metrics center on Scene Intersection over Union:
reported separately as \(\text{mSIoU}\) (all matched scenes), \(\text{IoUd}\) (dialogue scenes), and \(\text{IoUnd}\) (non-dialogue scenes), complemented by boundary localization accuracy (\(\text{BndF1}\)), Recall at rank \(k\) (\(\text{R@1}, \text{R@5}\)), and median rank (\(\text{MedR}\)).
Key Experimental Results¶
Main Results¶
Fourteen representative approaches and multimodal fusion variants were evaluated across all 50 feature films under the unified decoding pipeline:
| Paradigm | Method | mSIoU | BndF1 | IoUd | IoUnd | R@1 | R@5 | MedR |
|---|---|---|---|---|---|---|---|---|
| Dialogue | BM25 | 0.399 | 0.543 | 0.505 | 0.190 | 0.389 | 0.573 | 5.82 |
| Dialogue | TF-IDF | 0.396 | 0.546 | 0.507 | 0.175 | 0.346 | 0.576 | 5.65 |
| Dialogue | M/S (Cour et al.) | 0.324 | 0.480 | 0.412 | 0.156 | 0.446 | 0.545 | 6.10 |
| Visual | Qwen3-VL-Embedding | 0.500 | 0.629 | 0.521 | 0.456 | 0.257 | 0.536 | 6.40 |
| Visual | FG-CLIP 2 | 0.317 | 0.497 | 0.329 | 0.289 | 0.174 | 0.417 | 10.70 |
| Visual | CLIP (ViT-B/32) | 0.252 | 0.434 | 0.253 | 0.252 | 0.144 | 0.375 | 11.88 |
| Visual | InternVideo2 | 0.222 | 0.429 | 0.227 | 0.205 | 0.128 | 0.306 | 19.52 |
| Visual | SigLIP | 0.023 | 0.189 | 0.020 | 0.035 | 0.048 | 0.167 | 30.58 |
| Temporal Grounding | T-MASS | 0.130 | 0.300 | 0.123 | 0.141 | 0.088 | 0.261 | 21.02 |
| Temporal Grounding | MATR | 0.016 | 0.205 | 0.014 | 0.028 | 0.023 | 0.078 | 66.49 |
| Temporal Grounding | TAN | 0.009 | 0.181 | 0.006 | 0.015 | 0.003 | 0.026 | 67.32 |
| Video-Text Retrieval | CLIP4Clip | 0.012 | 0.213 | 0.010 | 0.017 | 0.006 | 0.043 | 58.57 |
| Video-Text Retrieval | DiffusionRet | 0.011 | 0.217 | 0.008 | 0.016 | 0.014 | 0.059 | 67.30 |
| Video-Text Retrieval | DiscoVLA | 0.013 | 0.221 | 0.011 | 0.026 | 0.013 | 0.059 | 69.80 |
| Multimodal Fusion | Qwen3-VL-Emb. + BM25 (\(\alpha=0.7\)) | 0.599 | 0.695 | 0.654 | 0.496 | 0.448 | 0.686 | 2.70 |
| Multimodal Fusion | CLIP + BM25 (\(\alpha=0.5\)) | 0.531 | 0.633 | 0.597 | 0.403 | 0.416 | 0.594 | 4.50 |
| Multimodal Fusion | FG-CLIP 2 + BM25 (\(\alpha=0.5\)) | 0.526 | 0.638 | 0.598 | 0.378 | 0.426 | 0.606 | 4.30 |
Ablation Study¶
The cross-validation study explores sensitivity to the fusion parameter \(\alpha\), while downstream zero-shot transfer evaluates boundary detection against human scene segmentation labels on MovieNet-SSeg:
| Evaluation / Task | Setting or Model | Key Metric 1 | Key Metric 2 | Note |
|---|---|---|---|---|
| 5-Fold CV: CLIP + BM25 | Selected \(\alpha=0.5\) (5/5 folds) | mSIoU = 0.531 | BndF1 = 0.633 | Consistent across folds (0.452โ0.615) |
| 5-Fold CV: Qwen3-VL + BM25 | Selected \(\alpha=0.7\) (4/5 folds) | mSIoU = 0.599 | BndF1 = 0.695 | Stronger visual encoder justifies higher visual weight |
| Zero-Shot Movie Scene Segmentation | Ours (Alignment-derived boundaries) | F1 = 25.61% | mIoU = 32.24% | Matches classical unsupervised DP baseline (32.00%) |
| Boundary Agreement (44 films) | Tolerance \(\pm 3\) shots F1 | F1 = 0.752 | Cohen's \(\kappa\) = 0.425 | 77.5% of human ambiguous boundaries fall within \(\pm 2\) shots |
Key Findings¶
- The Dialogue Anchoring Effect: While CLIP alone achieves only \(0.252\) IoUnd on non-dialogue scenes, fusing it with BM25 boosts IoUnd to \(0.403\) (+60% relative gain) despite using identical visual representations. Because BM25 anchors dialogue scenes to their exact temporal locations, the monotonicity constraints in Drop-DTW propagate these anchors to adjacent non-dialogue scenes, dramatically shrinking their visual search space.
- Catastrophic Failure of Short-Clip Models: Methods tailored for short-clip retrieval or moment localization (CLIP4Clip, DiffusionRet, DiscoVLA, TAN, MATR) fail catastrophically on full-length movie alignment (\(\text{mSIoU} \le 0.016\), \(\text{MedR} > 50\)). These architectures cannot handle long-form context and struggle with the nuanced literary language of screenplays.
- Persistent Visual Bottleneck in Narrative Grounding: Even with state-of-the-art vision-language embeddings (Qwen3-VL-Embedding) fused with BM25, a significant performance gap remains between dialogue and non-dialogue scenes (\(0.654\) vs. \(0.496\)). This identifies stage direction and visual narrative encoding as the primary remaining bottleneck for long-form multimodal understanding.
Highlights & Insights¶
- Curated 50 Feature Films Spanning Nine Decades: Provides 55K+ rigorously verified shot-level ground-truth annotations across 16 genres, overcoming the historical lack of benchmark data for full-length narrative alignment.
- Discovered and Quantified the Dialogue Anchoring Effect: Demonstrates how discrete, high-precision semantic signals (dialogue text) provide temporal structural constraints that significantly empower continuous, low-confidence visual representations.
- Validated Concrete Downstream Utility: Shows that alignment boundaries serve as effective zero-shot pseudo-labels for movie scene segmentation, while aligned scene-shot keyframe pairs allow LLMs to rewrite abstract script text into cinematographically grounded prompts for commercial video generation tools (Kling, Seedance, Veo).
Limitations & Future Work¶
- Domain and Language Concentration: All 50 feature films originate from MovieNet and represent Western, English-language cinema; generalizability to non-linear arthouse films or non-English cinema remains unverified.
- Lack of Joint Unmatched Scene Detection: Approximately 20% of screenplay scenes are deleted in post-production and unfilmed; current evaluation isolates matched scenes rather than jointly classifying deleted scenes.
- Future Directions: Developing specialized multimodal representation learning for screenplay syntax (e.g., stage directions and camera movements) and leveraging aligned film data for controllable, script-guided long-video generation.
Related Work & Insights¶
- vs Movie/Script (Cour et al., ECCV 2008): Movie/Script relied strictly on word-level DTW dialogue matching and failed on silent or action scenes; this work establishes a 50-film benchmark integrating visual encoders and diagnosing both channels.
- vs MovieNet (Huang et al., ECCV 2020): MovieNet provided raw video and screenplay files without alignment; this work performs frame-rate recovery, screenplay parsing, and 55K-shot manual validation to build a complete grounding testbed.
- vs Video-Text Retrieval (CLIP4Clip, UCoFiA): Conventional VTR models assume short clips and single captions; this work demonstrates their fundamental failure on long-form, heterogeneous screenplay alignment.
Rating¶
- Novelty: โญโญโญโญโญ Establishes the first systematic full-length screenplay-to-shot grounding benchmark with extensive dual-annotator verification.
- Experimental Thoroughness: โญโญโญโญโญ Evaluates 14 diverse approaches across four paradigms, complete with 5-fold cross-validation and two downstream tasks.
- Writing Quality: โญโญโญโญโญ Rigorous task formulation, clear mathematical grounding, and thorough empirical diagnostics.
- Value: โญโญโญโญโญ Provides foundational data and insights for narrative video understanding, long-form multimodal grounding, and AI-assisted filmmaking.