CrossView: Can Vision-Language Models Reason Across Cameras?¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://utaustin-swarmlab.github.io/CrossView
Area: LLM Reasoning
Keywords: Multi-Camera Reasoning, Vision-Language Models, Multi-View Video QA, Spatio-Temporal Scene Graph, Context Scaling
TL;DR¶
To address the pervasive "single-camera assumption" in multimodal evaluation, this paper presents CrossViewโthe first video QA benchmark spanning autonomous driving, surveillance, ego-exo interaction, and robotics across 2 to 8 concurrent feedsโrevealing that leading VLMs exhibit severe performance drops on cross-view spatial-temporal stitching, instance counting, and viewpoint selection.
Background & Motivation¶
Modern large vision-language models (LVLMs) and long-video understanding benchmarks (e.g., Video-MME, LongVideoBench) have reached near-saturation on single-camera video question answering, often achieving accuracy competitive with human annotators. However, real-world embodied systems and intelligent infrastructure rarely operate through a single perspective: autonomous vehicles fuse surround-view camera rigs to navigate occluded intersections, collaborative robots synchronize wrist-mounted and overhead views for fine manipulation, and wide-area security infrastructures track subjects across disjoint, non-overlapping camera feeds. Despite this physical reality, visual models continue to be evaluated almost exclusively on single-stream visual inputs.
The transition from single-camera to multi-camera video understanding introduces two qualitative challenges rather than a mere volume expansion: Context Scaling, where ingesting \(N\) concurrent high-resolution streams balloons visual token counts by an order of magnitude, readily exceeding the effective context window of existing models; and Cross-View Spatial Reasoning, which requires true spatio-temporal stitching across views. Models must distinguish overlapping from disjoint fields of view, identify the most informative viewpoint for a specific sub-task (camera identification), deduplicate and aggregate evidence across cameras (such as unique object instance counting), and establish causal temporal ordering across disjoint vantage points.
Existing multi-camera datasets remain confined to narrow, task-specific formulations: driving datasets such as nuScenes-QA evaluate localized object perception tied to single perspective crops, while human activity datasets like Ego-Exo4D emphasize viewpoint correspondence and procedural keystep recognition. None benchmark whether a model can directly digest heterogeneous, unconstrained multi-stream video to perform joint spatio-temporal semantic reasoning. Core idea: CrossView introduces a comprehensive multi-camera video question-answering benchmark comprising 6,000 algorithmic questions across four diverse domains, leveraging a two-stage Spatio-Temporal Scene Graph engine to systematically stress-test and quantify the multi-camera reasoning gap in contemporary VLMs.
Method¶
Overall Architecture¶
CrossView is constructed via a principled two-stage pipeline. The first stage builds a Spatio-Temporal Scene Graph (STSG) that consolidates multimodal 3D coordinates, LiDAR point clouds, bounding box trajectories, and dense activity descriptions across diverse video collections into a unified relational graph. The second stage deploys a programmatic Grounding-Target question-generation engine that queries topological and temporal constraints from the STSG to sample valid question-answer pairs alongside calibrated negative distractors, followed by LLM-assisted syntactic verbalization to guarantee factual consistency and eliminate hallucinations.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-Source Heterogeneous Video Feeds<br/>nuScenes / MEVA / Ego-Exo4D / AgiBot"] --> B["Stage 1: Spatio-Temporal Scene Graph (STSG)<br/>LiDAR/BBox grounding + VLM activity annotation + relational snapshots"]
B --> C["Stage 2: Programmatic Grounding-Target QA Engine<br/>Spatio-temporal predicate querying + topological sampling + 3 distractor types"]
C --> D["Multi-Camera Reasoning Task Suite<br/>Temporal ordering / Spatial relations / Counting / Camera-ID / Summarization"]
D --> E["Multi-View Visual Aggregation Evaluation<br/>Uniform independent sampling vs. Stitched composite tiling"]
Key Designs¶
1. Spatio-Temporal Scene Graph (STSG): Unified Multimodal Foundation To insulate the benchmark from generative hallucinations, question generation must be anchored on a physically verified, topologically coherent representation. The STSG unifies data from nuScenes (autonomous driving), MEVA (wide-area surveillance), Ego-Exo4D (first- and third-person collaboration), and AgiBot (dual-arm robotic manipulation). For spatial localization, nuScenes incorporates 3D LiDAR point clouds to localize objects in world coordinates, while other datasets utilize verified 3D or 2D bounding-box annotations. For semantic activity metadata lacking native annotations (nuScenes and AgiBot), the pipeline extracts cropped object tubes and queries InternVL-3.5 38B to generate object-centric activity descriptions, whereas Ego-Exo4D and MEVA supply native action labels. GPT-5.2 parses these per-frame objects and actions into timestamp-level scene captions. Pairwise directional predicates (such as behind, near, left, right), orientation quaternion deltas, and Euclidean metric distances are computed across all concurrent views. Consecutive timestamps sharing uniform object activity are grouped into event intervals \([t_{start}, t_{end}]\), yielding snapshot graphs \(G_t = (N_t, E_t)\) linked sequentially into a continuous STSG \(G\).
2. Programmatic Grounding-Target QA Engine: Constrained Multi-View Synthesis The question synthesis engine enforces that every task intrinsically requires cross-camera evidence integration. By querying the STSG with structured anchors, the engine identifies grounding events \(E_g\) and target events \(E_t\) governed by explicit temporal relations: Before (\(E_t.end < E_g.start\)), After (\(E_t.start > E_g.end\)), During (temporal overlap exceeding 50% of the shorter event duration), and In-between (the target event falls strictly within the temporal gap separating two non-overlapping anchor events). This logic extends to spatio-temporal cross-referencing: Category I queries temporal events triggered by a spatial configuration visible in another camera, while Category II queries spatial configurations bounded by temporal markers across cameras. Hard distractors are sampled programmatically across three categories: spatial distractors (incorrect direction at the correct timestamp), temporal distractors (correct event at an incorrect timestamp), and existential distractors (plausible objects entirely absent from the scene).
3. Comprehensive Multi-View Task Suite: Disentangling Cross-Camera Capabilities CrossView structures 6,000 multiple-choice and open-ended questions evenly across six diagnostic tasks (1,000 per task, spanning 2 to 8 concurrent synchronized cameras): - Temporal Reasoning & Chronological Event Ordering: Reconstructing the global chronological sequence of 3 to 5 scrambled, non-overlapping events occurring across distinct camera feeds. - Cross-View Spatial Reasoning: Evaluating metric proximity and relative orientations in non-ego coordinate frames using orientation quaternions, forcing models to reason allocentrically across viewpoint transformations. - Cross-Camera Instance Counting: Requiring models to track global Unique Identifiers (UIDs) across viewpoints to accurately count target classes under persistent occlusion, camera transit, and appearance changes without double-counting. - Optimal Viewpoint Selection (Best-Camera Identification): Identifying which camera perspective maintains the most persistent, unobstructed visual evidence of a designated target activity. - Global Scene Summarization: Generating holistic narratives of multi-actor dynamics across all concurrent camera perspectives, evaluated via ROUGE metrics.
4. Visual Input Aggregation Strategies: Uniform vs. Stitched Framing To investigate how input structure impacts attention across cameras, CrossView compares two canonical frame aggregation strategies: Uniform Sampling, where \(N\) frames are sampled independently from each camera and passed as interleaved image sequences (16 frames per camera for Qwen, GPT, and Gemma; 4 to 8 for InternVL due to context limits); and Stitched Sampling, where synchronized frames across all cameras at timestamp \(t\) are tiled into a single composite surround-view image, providing explicit spatial layout cues within each visual token grid.
Key Experimental Results¶
Main Results¶
Eleven representative VLMs across four major families (GPT-5.2, Qwen2.5/Qwen3, InternVL2/2.5/3.5, and Gemma-3) were evaluated under uniform frame sampling. The table below details multiple-choice accuracy (%) on the vehicle-centric (nuScenes) and robot-centric (AgiBot) benchmarks.
| Family | Model | nuScenes Counting | nuScenes Event Ordering | nuScenes Spatial | nuScenes Temporal | AgiBot Temporal | AgiBot Event Ordering |
|---|---|---|---|---|---|---|---|
| Qwen | Qwen2.5-3B | 32.5 | 36.8 | 30.1 | 32.2 | 52.0 | 50.8 |
| Qwen | Qwen2.5-7B | 27.5 | 24.4 | 45.9 | 21.6 | 58.8 | 54.4 |
| Qwen | Qwen3-4B | 46.6 | 46.4 | 28.9 | 33.8 | 66.4 | 49.2 |
| Qwen | Qwen3-8B | 43.2 | 41.2 | 33.1 | 36.0 | 69.6 | 49.2 |
| InternVL | InternVL2-8B | 28.8 | 36.8 | 37.3 | 29.2 | 42.0 | 40.4 |
| InternVL | InternVL2.5-8B | 37.1 | 28.4 | 32.1 | 33.4 | 60.0 | 46.4 |
| InternVL | InternVL3.5-4B | 40.2 | 30.0 | 26.1 | 38.4 | 52.8 | 42.0 |
| InternVL | InternVL3.5-14B | 46.2 | 36.4 | 43.3 | 44.8 | 54.8 | 47.6 |
| GPT | GPT-5.2 | 47.8 | 29.2 | 36.5 | 25.4 | 66.8 | 66.0 |
| Gemma | Gemma-3-4B | 44.7 | 41.6 | 18.6 | 38.2 | 50.8 | 36.8 |
| Gemma | Gemma-3-12B | 43.8 | 43.2 | 26.4 | 35.6 | 60.4 | 41.2 |
Evaluations on human-centric subsets (Ego-Exo4D and MEVA) corroborate this bottleneck: on Ego-Exo4D best-camera identification, GPT-5.2 obtains only 34.6% accuracy, while open-source models hover between 20.2% and 33.2% (barely exceeding the 25% random guessing floor). In wide-area MEVA scenarios, accuracy drops severely across all models, with GPT-5.2 scoring 23.3% on temporal reasoning and 27.8% on instance counting.
Ablation Study¶
To verify that CrossView questions strictly require multi-camera synthesis rather than exploiting single-view visual shortcuts, extensive ablation studies were conducted using Qwen2.5-7B-Instruct.
1. Multi-Camera vs. Single-Camera Input Ablation (Qwen2.5-7B)
| Dataset | Input Setting | Counting (%) | Temporal (%) | Event Ordering (%) | Spatial (%) |
|---|---|---|---|---|---|
| Ego-Exo4D | Full Multi-Camera | โ | 45.2 | 49.2 | โ |
| Ego-Exo4D | Single Best Exocentric (Best Exo) | โ | 40.7 (-4.5) | 47.3 (-1.9) | โ |
| Ego-Exo4D | Single Egocentric (Ego) | โ | 40.4 (-4.8) | 48.8 (-0.4) | โ |
| nuScenes | Full 6-Camera Multi-View | 27.5 | 21.6 | 24.4 | 45.9 |
| nuScenes | Single Front-Facing Camera | 18.7 (-8.8) | 21.6 (ยฑ0.0) | 32.0 (+7.6) | 41.5 (-4.4) |
2. Uniform vs. Stitched Frame Sampling Strategy (Qwen2.5-7B)
| Dataset | Frame Strategy | Counting (%) | Temporal (%) | Event Ordering (%) | Spatial (%) |
|---|---|---|---|---|---|
| nuScenes | Uniform Independent Sampling | 27.5 | 21.6 | 24.4 | 45.9 |
| nuScenes | Stitched Tiled Composite Image | 38.0 (+10.5) | 26.2 (+4.6) | 39.6 (+15.2) | 46.3 (+0.4) |
| AgiBot | Uniform Independent Sampling | โ | 58.8 | 54.4 | โ |
| AgiBot | Stitched Tiled Composite Image | โ | 58.4 (-0.4) | 48.4 (-6.0) | โ |
| MEVA | Uniform Independent Sampling | 11.1 | 49.2 | 83.2 | 17.6 |
| MEVA | Stitched Tiled Composite Image | 13.3 (+2.2) | 52.2 (+3.0) | 70.9 (-12.3) | 34.2 (+16.6) |
Key Findings¶
- Model Scale Does Not Overcome Multi-Camera Blind Spots: Proprietary flagship models such as GPT-5.2 achieve top-tier performance on single-camera leaderboards but collapse on multi-camera integration, achieving only 25.4% on nuScenes temporal reasoning and 29.2% on event ordering. The multi-camera gap persists across model scales, confirming that general visual instruction tuning does not impart multi-view geometric awareness.
- Cross-View Synthesis Tasks Are Inherently Hardest: Tasks requiring joint visual integrationโcross-camera object counting and optimal viewpoint selectionโconsistently register the lowest performance across all models (mostly between 20% and 35%), demonstrating that current architectures treat multi-image inputs as disjoint token bags rather than a unified 3D environment.
- Scene Complexity and Camera Density Dictate Difficulty: Controlled indoor manipulation with two static views (AgiBot) yields the highest accuracy (up to 69.6%), whereas wide-area outdoor surveillance with dense, overlapping cameras (MEVA) degrades accuracy to near-random levels and suppresses summarization ROUGE-L scores below 12.0.
- Stitched Framing Mitigates Attention Dilution in Surround Settings: Tiling concurrent camera views into a single composite frame dramatically boosts performance on contiguous setups: counting on nuScenes increases by +10.5% and event ordering by +15.2%, while spatial reasoning on MEVA jumps by +16.6%. However, on low-clutter, disjoint setups like AgiBot, stitching slightly degrades performance due to reduced per-camera resolution and visual clutter.
Highlights & Insights¶
- Zero-Hallucination Question Synthesis via STSG: By anchoring questions to ground-truth LiDAR points and verified spatial-temporal predicates, the STSG framework synthesizes challenging multi-hop questions with guaranteed physical consistency, establishing a rigorous template for multi-sensor benchmark generation.
- Empirical Validation of Multi-View Necessity: Single-camera degradation tests show consistent accuracy drops (e.g., -8.8% on nuScenes counting), validating that CrossView genuinely evaluates cross-view fusion rather than single-view shortcuts.
- Tiling as a Practical Alternative to Sparse Multi-Image Attention: The significant gains observed with stitched frame inputs reveal that pre-merging multi-camera views into shared visual coordinate grids enables models to leverage standard 2D spatial attention, offering concrete architectural guidance for vision-language models deployed in autonomous systems.
Limitations & Future Work¶
- Reliance on Fixed and Pre-Calibrated Rigs: Scenarios predominantly feature static surveillance poles or rigidly mounted vehicle rigs, leaving uncalibrated dynamic multi-agent camera networks (e.g., swarms of drones or mixed mobile agents) unexamined.
- Sparse Frame Budgets Constrained by Context Windows: Models are evaluated under uniform budgets of 4 to 16 frames per camera to respect token limits, leaving dense continuous video stream understanding across cameras as an open challenge.
- Future Directions: Developing dedicated multi-camera position embeddings and cross-view geometry adapters that explicitly ingest extrinsic camera matrices into transformer backbones without sacrificing token resolution.
Related Work & Insights¶
- vs nuScenes-QA / NuPlanQA: While prior autonomous driving QA datasets draw frames from multi-camera rigs, questions are grounded primarily in single views or pre-rendered BEV maps; CrossView explicitly tests spatio-temporal reasoning directly over raw, multi-view concurrent feeds.
- vs Ego-Exo4D / EgoExoBench: Existing first/third-person datasets emphasize pairwise visual alignment and keystep correspondence; CrossView extends to arbitrary networks of 2 to 8 heterogeneous cameras with complex global spatial queries and instance deduplication.
- vs 3D-LLM / BEVFormer: 3D LLMs depend on heavy offline 3D point cloud reconstruction or explicit BEV tokenizers; CrossView establishes a pure 2D vision-language baseline, challenging models to perform implicit multi-view spatial reasoning without external 3D neural representations.
Rating¶
- Novelty: โญโญโญโญโญ Pioneering multi-camera video QA benchmark directly addressing the long-overlooked "single-camera assumption" in multimodal reasoning.
- Experimental Thoroughness: โญโญโญโญโญ Rigorous evaluation across 11 frontier models, 4 distinct real-world domains, and insightful ablations on camera restrictions and frame aggregation modes.
- Writing Quality: โญโญโญโญโญ Exceptionally clear narrative, disciplined methodology, self-consistent empirical metrics, and transparent error analyses.
- Value: โญโญโญโญโญ Crucial benchmark guiding the architectural evolution of multimodal foundation models for autonomous driving, robotics, and physical-world perception.