Natural Language Camera Movement Understanding¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://1yuwen.github.io/ACaM-Project-Page/
Area: Video Generation
Keywords: Camera movement understanding, Vision-language models, Cinematographic taxonomy, Video benchmark, Instruction tuning
TL;DR¶
Addressing the core weaknesses of vision-language models (VLMs) in confusing rotation with translation, swapping directions, and mistaking foreground motion for camera trajectories, this paper introduces a 17-class cinematographic taxonomy, the dual-domain ACaM benchmark, and targeted geometric augmentations to fine-tune specialized VLMs that substantially outperform commercial frontier models.
Background & Motivation¶
Camera movement control is a foundational capability for modern text-to-video (T2V) generative models seeking to transition toward professional filmmaking, where creators increasingly use precise natural language cinematographic prompts (such as "a slow pan" or "dolly in") to direct dynamic scenes. Scaling controllable generative systems requires massive volumes of high-quality paired video-text data with accurate camera descriptions, and evaluating these systems depends heavily on automated, reliable verification of whether generated clips faithfully carry out camera movement instructions. However, classical geometric pose estimation frameworks (such as SLAM or SfM) produce coordinate-level poses that are computationally intensive, demand calibrated camera intrinsics, and fail to comprehend semantic, object-centric movements like tracking or arcing. Vision-language models (VLMs) offer an intuitive and scalable cross-modal alternative, yet current models exhibit surprising perceptual failures when confronted with camera dynamics.
A closer look at these failures uncovers the core tension: contemporary VLMs lack an intrinsic understanding of 3D physical space and cinematographic mechanics. Specifically, existing models suffer from five pervasive failure modes: they are largely insensitive to subtle inter-frame shifts (frequently misclassifying gentle camera motion as "static"); they conflate physical translation with camera body rotation (e.g., confusing pans with lateral trucks); they suffer from left-right directional confusion; they fail to separate optical zoom from physical dolly motion because they overlook perspective parallax; and they routinely misinterpret salient foreground object motion as global camera movement (such as classifying an approaching race car as a "push in"). Furthermore, existing benchmarks either focus solely on real-world clips without assessing generative models, or submerge camera dynamics within general action recognition rather than isolating atomic cinematographic primitives.
To resolve these challenges, this paper establishes natural language camera movement understanding as an independent research problem. Core idea: develop a two-level cinematographic taxonomy covering 17 atomic camera movements alongside the dual-domain ACaM benchmark (real and synthetic), and leverage targeted geometric motion augmentation to eliminate long-tail imbalance when fine-tuning dedicated VLMs.
Method¶
Overall Architecture¶
The framework establishes an end-to-end pipeline spanning taxonomy definition and failure mode diagnosis, real and synthetic benchmark construction (ACaM), multi-source instruction-tuning data curation with targeted geometric augmentation, and parameter-efficient model fine-tuning. Grounded in a 17-class cinematographic taxonomy, the evaluation pipeline integrates refined real-world film clips with closed-loop Veo 3.1 video generation. On the training side, the system counteracts the extreme dominance of "push in" and the scarcity of roll and zoom in raw video pools by applying spatial dynamic cropping, temporal reversal, horizontal flipping, and affine rotation augmentations, yielding a balanced 27K sample instruction-tuning corpus for Qwen3-VL.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-Source Video Pool<br/>45K raw video clips"] --> B["17-Class Cinematographic Taxonomy<br/>Translation/Rotation/Zoom/Static/Object-centric"]
B --> C["Dual-Domain ACaM Benchmark<br/>Real-world curation + Veo 3.1 closed-loop generation"]
C --> D["Targeted Physical Motion Augmentation<br/>Dynamic crop/Temporal reverse/Horizontal flip/Affine roll"]
D --> E["Balanced Instruction-Tuning Dataset<br/>27K filtered MCQ video pairs"]
E --> F["VLM Instruction Tuning<br/>High-rank LoRA on Qwen3-VL"]
F --> G["Camera Movement Understanding Model<br/>Precise natural language camera motion reasoning"]
Key Designs¶
1. Two-level cinematographic taxonomy and five VLM failure modes: establishing geometric and semantic boundaries
To eliminate ambiguity between camera movement and general object activity, the authors ground the task in cinematographic literature by structuring camera dynamics into 5 high-level dimensions and 17 mutually exclusive atomic classes. The first-level dimensions comprise physical translation (Dolly In/Out, Truck Left/Right, Boom Up/Down), angular rotation (Pan Left/Right, Tilt Up/Down, Roll Clockwise/Counterclockwise), focal length change (Zoom In/Out), static shots, and object-centric movements (Tracking, Arc). Systematic evaluations reveal five structural blind spots in general VLMs: insensitivity to subtle inter-frame shifts, confusion between translation and rotation, directional left-right reversals, inability to differentiate optical zoom from physical dolly motion, and reliance on salient foreground dynamics over global background parallax.
2. Dual-domain ACaM benchmark construction: bridging real-world cinematography and generative video evaluation
To support both training data filtering and generative model benchmarking, the ACaM benchmark introduces a dual-track evaluation suite. For real-world videos, the authors unify and filter subsets from CameraBench, ShotBench, CineTechBench, MotionBench, and FavorBench, pruning compound movements and questions tainted by foreground object motion. Cinematography-trained graduate students manually audit the set and collect additional YouTube clips to enrich underrepresented classes, resulting in 1,423 clips and 1,464 QA pairs. For synthetic videos, Gemini-3-Pro converts real-world examples into structured generation prompts for Veo 3.1 to synthesize 4-second videos. A quality-control loop either accepts faithful videos, reclassifies shifted movements, or rewrites ambiguous prompts for a second generation pass, establishing the first camera movement benchmark with calibrated human baseline comparisons.
3. Targeted physical motion augmentation: mitigating multi-source imbalance and perceptual shortcuts
The 45.3K raw video dataset collected across five repositories (CameraBench, ShotBench, SpatialVID, MultiCamVideo, and GenDoP) exhibits extreme long-tail imbalance, where forward translation dominates while zoom, roll, and lateral translations are scarce. Directly fine-tuning on this distribution causes models to learn shortcut priors. The authors formulate four physically valid data augmentation operators: synthesizing optical zoom by progressively cropping static high-resolution frames and resizing them back to standard resolution (~1K clips); applying temporal reversal to convert Dolly In clips into geometrically sound Dolly Out shots (~2K clips); performing horizontal flipping on Truck movements to balance left and right lateral shifts (~1.7K clips); and applying continuous affine rotations to static shots to produce realistic rolling dynamics. This strategy reshapes the corpus into a clean, balanced 27K video-MCQ instruction-tuning dataset.
4. 3D motion-aware VLM instruction tuning: mapping spatio-temporal dynamics to cinematographic reasoning
To endow general multimodal backbones with 3D camera mechanics perception, the authors perform instruction tuning on Qwen3-VL (4B and 8B). Unlike spatial VLMs that focus primarily on the static 3D layout of scene objects, this formulation requires the model to track background optical flow and perspective scaling ratios across sampled video frames (8 to 32 frames) to answer multiple-choice questions. Empirical results demonstrate that configuring LoRA with higher rank capacity (\(r=256\)) effectively expands the visual backbone's capacity to represent subtle geometric deformations and motion parallax, enabling the compact 8B model to surpass commercial closed-source frontier models.
Loss & Training¶
The instruction tuning is optimized using standard autoregressive next-token cross-entropy loss: $$ \mathcal{L}{\text{SFT}} = - \sum) $$ where }^{T} \log P(y_t \mid y_{<t}, \mathbf{V}, \mathbf{X}_{\text{prompt}\(\mathbf{V}\) represents the sampled video frame sequence, \(\mathbf{X}_{\text{prompt}}\) denotes the user instruction prompt encoding the multiple-choice question format, and \(y_t\) represents the target answer tokens. The models are optimized with LoRA adaptors using a cosine learning rate decay schedule over the refined 27K balanced video-MCQ corpus.
Key Experimental Results¶
Main Results¶
Performance on ACaM Real-World Videos in the Multiple-Choice QA (MCQ) evaluation setting (Random guess baseline is ~24.37%):
| Model | Static | Rot. | Trans. | Zoom | Arc | Track | Overall | Avg |
|---|---|---|---|---|---|---|---|---|
| Random Guess | 24.50 | 24.28 | 24.00 | 24.54 | 24.78 | 25.00 | 24.27 | 24.37 |
| Human Performance | 99.00 | 93.07 | 90.39 | 90.83 | 97.10 | 96.97 | 93.44 | 93.14 |
| Mega-SaM (Visual Geometry) | 89.55 | 74.94 | 54.16 | โ | โ | โ | โ | โ |
| ViPE (Visual Geometry) | 61.81 | 49.71 | 73.17 | โ | โ | โ | โ | โ |
| Spatial-MLLM (Spatial VLM) | 59.00 | 16.98 | 31.72 | 41.67 | 24.29 | 77.27 | 34.22 | 29.68 |
| Qwen3-VL-8B (General VLM) | 85.00 | 65.84 | 39.50 | 61.67 | 79.71 | 74.24 | 58.27 | 55.25 |
| Qwen3-VL-32B (General VLM) | 92.00 | 64.11 | 47.93 | 64.17 | 57.97 | 86.36 | 61.95 | 58.66 |
| GPT-5 (Proprietary VLM) | 61.50 | 66.58 | 54.55 | 61.67 | 68.12 | 80.30 | 61.20 | 61.54 |
| Gemini-3.1-Pro (Proprietary VLM) | 83.92 | 63.34 | 64.00 | 62.39 | 76.81 | 95.38 | 68.44 | 67.73 |
| CameraModel-7B (Specialized VLM) | 55.50 | 68.81 | 47.93 | 45.83 | 91.30 | 98.48 | 58.88 | 60.56 |
| ShotVL-7B (Specialized VLM) | 91.00 | 60.40 | 51.07 | 32.50 | 68.12 | 98.48 | 60.52 | 57.09 |
| Ours SFT Qwen3-VL-4B | 76.00 | 71.53 | 64.13 | 76.67 | 78.26 | 93.94 | 70.83 | 72.27 |
| Ours SFT Qwen3-VL-8B | 85.50 | 70.05 | 64.13 | 84.17 | 84.06 | 96.97 | 72.75 | 74.28 |
Performance on ACaM Synthetic Videos (generated via Veo 3.1) in the Multiple-Choice QA (MCQ) evaluation setting:
| Model | Static | Rot. | Trans. | Zoom | Arc | Track | Overall | Avg |
|---|---|---|---|---|---|---|---|---|
| Human Performance | 98.88 | 97.72 | 98.04 | 96.15 | 100.00 | 97.62 | 98.05 | 97.61 |
| Qwen3-VL-8B (General VLM) | 90.50 | 62.39 | 45.10 | 80.77 | 55.56 | 83.33 | 61.92 | 60.91 |
| Qwen3-VL-32B (General VLM) | 91.62 | 68.95 | 55.77 | 82.69 | 51.85 | 67.86 | 67.01 | 64.17 |
| Gemini-3.1-Pro (Proprietary VLM) | 92.07 | 62.00 | 74.07 | 69.23 | 55.56 | 85.37 | 73.01 | 69.58 |
| ShotVL-7B (Specialized VLM) | 96.09 | 69.80 | 63.62 | 63.46 | 29.63 | 98.81 | 71.33 | 64.49 |
| Ours SFT Qwen3-VL-4B | 82.68 | 74.36 | 66.67 | 78.85 | 68.52 | 84.52 | 73.28 | 74.40 |
| Ours SFT Qwen3-VL-8B | 94.41 | 72.36 | 74.51 | 82.69 | 68.52 | 91.67 | 78.20 | 77.44 |
Ablation Study¶
Ablation of balanced sampling, targeted augmentation, and LoRA rank capacity across Qwen3-VL models:
| Sampling | Augmentation | LoRA Rank | Model Base | Real Avg. (%) | Syn Avg. (%) | Note |
|---|---|---|---|---|---|---|
| โ | โ | \(r=64\) | 4B | 48.42 | 58.56 | Direct tuning on raw unbalanced 45K clips |
| โ | โ | \(r=64\) | 4B | 63.78 | 72.07 | Frequency-balanced sampling prevents class bias |
| โ | โ | \(r=64\) | 4B | 70.88 | 75.03 | Targeted physical augmentation covers rare classes |
| โ | โ | \(r=128\) | 4B | 71.88 | 75.20 | Moderate adaptation capacity gain |
| โ | โ | \(r=256\) | 4B | 72.27 | 74.40 | Performance plateau for the 4B parameter scale |
| โ | โ | \(r=256\) | 8B | 74.28 | 77.44 | Optimal scaling combining full strategy on 8B model |
Key Findings¶
- Data balancing and targeted augmentation are primary drivers of performance: Class-balanced re-sampling alone boosts 4B synthetic video average accuracy from 58.56% to 72.07%. Incorporating geometric augmentations such as dynamic cropping and temporal reversal yields an additional leap to 75.03% (a 16.47% net absolute gain), verifying that resolving training distribution shifts is far more impactful than increasing model parameter count.
- Spatial VLMs fail on dynamic observer kinematics: Spatial-MLLM and G2VLM, despite being augmented with 3D coordinate and point cloud objectives, achieve near-random accuracy on camera motion (5%โ34%). They are tuned to comprehend static spatial relations among objects rather than the dynamic temporal relationship between an observer's trajectory and visual optical flow.
- Frontier commercial models exhibit severe blind spots in translation and zoom: Even Gemini-3.1-Pro lags behind human perception by over 26% on real-world translation (64.00% vs 90.39%) and zoom (62.39% vs 90.83%). The fine-tuned 8B model outperforms Gemini-3.1-Pro by 6.55% on real videos (74.28% vs 67.73%) and 7.86% on synthetic videos (77.44% vs 69.58%) in macro-averaged accuracy.
Highlights & Insights¶
- Formulates camera movement understanding as a standalone research primitive: Elevates camera movement understanding from an overlooked sub-task of generic video QA into a clearly defined benchmark with a 17-class taxonomy, serving as foundational infrastructure for generative video filtering and evaluation.
- Geometrically principled synthetic data augmentation: Addresses extreme class scarcity without manual annotation by converting static frames to zooms via crop-resizing, and converting push-in to pull-out via temporal reversal, maintaining pure physical parallax consistency.
- Diagnoses the "localized tracker bias" in general VLMs: Empirically uncovers that current vision-language models default to tracking dominant foreground objects instead of computing global background parallax, providing concrete guidance for incorporating optical flow or epipolar constraints into future multimodal architectures.
Limitations & Future Work¶
- Evaluation focuses exclusively on atomic camera movements: Both ACaM and the training corpus focus on single, dominant camera actions per clip, omitting complex compound movements common in professional films (such as a simultaneous dolly-pan or Hitchcock vertigo zoom).
- Video generator bias limits synthetic distribution: Using Veo 3.1 to synthesize rare motions encountered persistent generative biases (e.g., strong biases toward counterclockwise rolls or falling back to static shots on zoom prompts), restricting synthetic diversity for certain rare categories.
- Future directions: Extending the benchmark to temporal camera movement grounding (identifying start and end timestamps of movements) and exploring visual encoders that explicitly separate camera ego-motion fields from object dynamic fields.
Related Work & Insights¶
- vs CameraBench & ShotBench: CameraBench includes over 50 primitives but lacks synthetic video evaluation and provides limited training scale; ShotBench focuses on film QA without isolating pure atomic movements. ACaM provides 17 unambiguous atomic categories across both real and generative domains with human baseline parity.
- vs Visual Geometry Pose Estimation (Mega-SaM / ViPE): While SLAM and SfM methods estimate fine-grained numerical camera trajectories, they cannot process semantic object-centric motions (such as tracking or orbiting) and struggle with zero-parallax zooms. The proposed VLM framework demonstrates superior semantic flexibility and natural language accessibility.
Rating¶
- Novelty: โญโญโญโญ [Establishes a principled standalone task formulation, complete taxonomy, and dual-domain benchmark]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive comparisons across geometry baselines, spatial VLMs, frontier commercial models, and fine-tuned architectures]
- Writing Quality: โญโญโญโญโญ [Clear structural diagnosis of VLM failure modes with comprehensive qualitative and quantitative ablations]
- Value: โญโญโญโญโญ [Directly applicable to data curation and automated quality benchmarking for modern controllable video generation]