Skip to content

OneCanvas: 3D Scene Understanding via Panoramic Reprojection

Conference: NeurIPS2026 Spotlight (acceptance metadata supplied with the task)
arXiv: 2606.19253
Paper: Project page
Code: https://github.com/baranowskibrt/onecanvas
Area: VLM Reasoning
Keywords: panoramic feature reprojection, spatial reasoning, metric position embedding, continuous visual tokens, spatial curriculum learning

TL;DR

OneCanvas lifts multi-view image patch features into 3D and places them as separate continuous tokens in a shared panoramic coordinate system, using native RoPE for angles and frame order and an additive embedding for metric position, then combines synthetic spatial pretraining with real-scene QA adaptation to score 65.3, 71.3, and 72.1 on SQA3D, VSI-Bench, and SPBench, respectively.

Background & Motivation

A vision-language model (VLM) may recognize objects in a video without reliably determining their separation or which object lies to the left from a specified position. Ordinary video input fragments a room into local views, while camera motion places the same object at different image locations; the model must both associate entities across frames and recover a shared reference frame. Point-cloud encoders, depth position encodings, and geometry-feature fusion provide one remedy. VLM-3R and SpaceMind demonstrate its effectiveness, but require alignment and training of additional modules. Another route scales spatial QA data, as in SenseNova-SI and Cambrian-S, at the cost of substantial high-quality supervision.

The deeper issue is that supplying geometry does not ensure the model uses it. Typical door heights, typical bathroom sizes, and frequent answers for a question template can all become shortcuts; exploiting those priors is easier than reading coordinates on biased data. OneCanvas therefore changes both input organization and supervision: it gives observations across frames a common spatial address, then establishes geometric reading with a curriculum whose answers follow from placement geometry, rather than expecting real-scene QA to teach this skill implicitly.

The panorama does not render the room into a new RGB image. It provides shared angular addresses for existing visual features. The vision encoder still processes familiar perspective images, and reprojection happens afterward, retaining collinear, occluded, and repeated observations. Core idea: organize existing patch tokens by continuous panoramic coordinates, restore metric depth through a separate content pathway, and first teach the VLM to read this representation with a synthetic curriculum that removes class-based shortcuts.

Method

Overall Architecture

Inference takes multi-view RGB, depth, camera intrinsics and poses, and a text question, and produces the usual VLM-generated answer. The representation follows 3D Lifting โ†’ Continuous Panoramic Tokens โ†’ Additive Metric Embedding, after which Qwen3-VL's existing language-model attention reads it. Two-Stage Spatial Training establishes this reading capability during training; it is not an additional module executed at inference.

The canvas origin and orientation can be selected for the task. SQA3D uses the queried person's position and heading, SPBench single-image questions use the queried camera pose, and other cases generally use the centroid of camera positions. Changing the origin gives the same world point a different angular address and local metric coordinate, allowing the model to read the scene in the reference frame required by the question instead of first inferring the viewpoint transformation from language.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["RGB, geometry, and question"] --> B["3D Lifting"]
    B --> C["Continuous Panoramic Tokens"]
    C --> D["Additive Metric Embedding"]
    D --> E["Native VLM โ†’ answer"]
    D -. "same representation, training only" .-> F["Two-Stage Spatial Training"]
    G["Synthetic curriculum supervision<br/>Real-scene QA supervision"] -.-> F
    F -. "learned inference weights" .-> E

Key Designs

1. 3D Lifting: convert image locations into shared world positions

Qwen3-VL's vision encoder remains frozen and independently encodes the original perspective frames. For each feature-grid patch, the method reads the corresponding depth, scales the camera intrinsics to feature-map resolution, backprojects the patch, and applies the camera-to-world pose. The key transformation is:

\[ \mathbf{p}^{\text{world}}_{u,v}=T_k\begin{pmatrix}(u-c_x)z/f_x\\(v-c_y)z/f_y\\z\end{pmatrix}. \]

Here \(T_k\) is a rigid transformation, and the camera follows OpenCV's right, down, and forward convention. The result is not a fused point-cloud feature representation: it is a set of patches retaining their world positions, original visual features, and source-frame indices. Those indices have a separate downstream role, ensuring that aligning observations of the same surface does not erase when it first appeared.

This organization delegates geometric alignment to an explicit input transformation rather than adding a heavyweight 3D encoder that must learn alignment with the language model. The corresponding cost is that depth and pose errors directly affect token positions; the method is not geometry-free RGB processing.

2. Continuous Panoramic Tokens: change addresses without rendering or merging observations

Given canvas origin \(\mathbf{c}\) and orientation \(R\), the method transforms a world point into canvas-local coordinates and computes longitude and latitude:

\[ \mathbf{q}_i=R^\top(\mathbf{p}_i-\mathbf{c}),\qquad \theta_i=\operatorname{atan2}(q_x,q_z),\qquad \phi_i=\arctan\left(\frac{-q_y}{\sqrt{q_x^2+q_z^2}}\right). \]

The negative sign makes latitude positive upward. Each patch remains a separate token at a continuous angular position rather than being written into a discrete panorama pixel. Even points sharing a direction are not averaged, reduced to the nearest surface, or overwritten by the last observation. Different surfaces along a ray and observations from overlapping frames remain available to attention. There is no z-buffer visibility test and no second encoding pass over an RGB panorama.

The algorithm of Qwen3-VL's native three-axis RoPE is unchanged; only the position-ID contents change. \(W\) receives longitude, \(H\) receives latitude, and \(T\) still receives the source-frame index, with each axis linearly rescaled to \([0,100]\). Thus โ€œ3D-RoPEโ€ does not mean placing three Cartesian metric coordinates into RoPE: the spatial axes carry angles, and the temporal axis carries frame order. This preserves the intended semantics of the pretrained position mechanism while making cross-view relationships readable in a common angular plane.

3. Additive Metric Embedding: preserve the depth lost in angular projection through the content pathway

Angles alone cannot distinguish nearby and distant objects along the same direction or reliably encode distances in metres. OneCanvas applies sine and cosine encodings to the three axes of the canvas-local offset: each axis uses 16 log-spaced frequencies in \([0.1,100]\) rad/m, with raw coordinates retained. It additionally encodes the total radius and horizontal radius with 8 frequencies each in \([0.1,10]\) rad/m, then concatenates the raw unit-ray direction. The radial quantities expose distance and horizontal-scale information directly instead of requiring the model to synthesize them entirely from per-axis components.

The resulting 136-channel encoding passes through a two-layer MLP with hidden width 64 and is added to the visual feature through a learned scalar gate. The mechanism described in the paper can be summarized as:

\[ \widetilde{\mathbf{f}}_i=\mathbf{f}_i+g\,\operatorname{MLP}(\operatorname{PE}_{\text{metric}}(\mathbf{q}_i)). \]

This equation summarizes the mechanism; it is not an additional numbered equation from the paper. Metric position enters through feature content, while angles and time enter through RoPE, and these pathways should not be conflated. The separation preserves native positional encoding while allowing attention to read actual scale and distinguish observations with matching angular positions but different depths.

4. Two-Stage Spatial Training: learn geometric reading before adapting to real-scene answers

Stage 1 does not transfer a real room onto an empty canvas. The default curriculum procedurally places synthetic boxes on an otherwise empty canvas, with every patch of a box carrying the same feature sampled from a pool of real visual activations. Questions reference boxes through inline copies of their features rather than class names such as โ€œdoorโ€ or โ€œchair.โ€ The activation pool comes from training-source scenes in ScanNet, ScanNet++, and ARKitScenes; box features are sampled independently and do not carry the current synthetic scene's identity. An inline copy uses the MRoPE ID of its text slot rather than its counterpart's angular canvas ID, and also receives the additive 3D position embedding.

Object identity is therefore established by content matching, while position, size, and separation follow from sampled geometry. Most samples add 6โ€“12 unreferenced distractor boxes so that the only object on the canvas or a feature outlier cannot reveal the answer; floor-area tasks omit distractors. Unlike a curriculum populated by real class-labelled objects, the default scheme deliberately disconnects categories from typical scales instead of prioritizing visual realism.

The curriculum has six task families: metric measurement, egocentric direction, multi-turn navigation, observability, counting, and multi-target 3D bounding-box readout. Each family contributes equally to each minibatch, and tasks within a family have equal weights, with answer distributions controlled. Measurement uses surface-to-surface distances and non-rectangular floor areas; navigation covers 1โ€“4 turns; appearance-order tasks explicitly control starting frames on the \(T\) axis; line-of-sight tasks instead use synthetic occlusion geometry and are not temporal. These tasks require scale, direction, frame order, and object binding rather than a single template prior.

Stage 1 trains rank-256 language-model LoRA and the metric embedding. The LoRA update is then merged into the base and frozen, and stage 2 trains a fresh rank-64 LoRA. The lower rank does not prove that forgetting cannot occur; it is the authors' way of limiting how far real-scene QA adaptation can move away from geometric reading. The vision encoder and token embeddings remain frozen in both stages.

A Worked Example

Consider a question asking which piece of furniture lies to the left when standing at a specified position and facing a specified direction. The method first encodes each perspective frame and places furniture-related patches in world coordinates using depth and poses. Centering and orienting the canvas on the queried person's pose makes panoramic longitude reflect direction relative to that person. Tokens observing the same furniture in different frames remain separate, with their frame information retained.

If a more distant piece of furniture lies in the same direction, its continuous angular address may be identical or nearby, but its additive metric embedding differs; input construction does not force either observation to be discarded. The language model combines the question with all tokens to generate an answer. Synthetic-box training previously taught it to read direction and position; boxes are not regenerated during inference on this real scene. This is an explanatory example, not a separately reported successful case in the paper.

Loss & Training

Training uses curriculum answers and downstream QA supervision, with the paper emphasizing that no auxiliary losses or dedicated 3D encoder are introduced. โ€œRegression targetโ€ in the appendix describes numerical labels and does not justify inventing a separate MSE objective. The backbone is Qwen3-VL-8B, with LoRA targeting the language model's attention and MLP projections.

Stage 1 uses LoRA rank 256, scaling coefficient 512, dropout 0.05, effective batch 16, learning rate \(2\times10^{-5}\), cosine scheduling, and 3% warmup. The metric embedding's learning rate is 50 times the base rate, and its norm is rescaled on every forward pass to half the mean visual-feature norm, so this stage mainly learns direction rather than magnitude. Each sample has at most 200 placed patches, with at least 8 per box. Stage 2 initializes the gate at the same magnitude, then lets it learn freely.

Stage 2 uses LoRA rank 64, scaling coefficient 128, dropout 0.05, 10,000 steps, and effective batch 32, with the same learning rate and schedule. Sampling weights for VLM-3R-VSIBench, SQA3D, and ViCA are approximately 70/15/15; these are not dataset-size proportions. Training resolution is \(320\times240\), with randomized canvas position and yaw during training. Evaluation uses \(320\times240\) for SQA3D and \(640\times480\) for VSI-Bench and SPBench. Scene videos generally use 32 uniformly sampled frames, while SPBench's predicted-geometry experiment uses the views actually supplied with each question.

Key Experimental Results

Main Results

Main results prioritize dataset-provided ground-truth depth and poses; a small subset with corrupted ground-truth tracks uses reconstruction estimates. The DA3 column below instead replaces depth, poses, and intrinsics entirely with RGB predictions at inference using the same checkpoint, with no retraining, extra views, or ground-truth geometry. โ€œZero-shotโ€ on SPBench means no SPBench training data, not an absence of other spatial QA training.

Dataset and metric Ours, ground-truth-geometry protocol Ours, DA3-predicted geometry Strongest prior result Gain under ground-truth protocol
SQA3D, EM@1 65.3 64.9 Ross3D, 63.0 +2.3
VSI-Bench, eight-task average 71.3 70.0 SpaceMind, 69.6 +1.7
SPBench, zero-shot Overall 72.1 71.3 SpaceMind, 67.3 +4.8

SQA3D EM@R1 is 68.4, while per-question-type results use strict EM@1. VSI-Bench route planning scores 60.8, substantially above SenseNova-SI's 48.5; absolute distance scores 62.3, relative distance 73.1, and room size 76.8. Not every subtask leads: relative direction at 88.0 trails SpaceMind's 88.4, and counting at 68.9 trails SpaceMind's 73.3. SPBench multi-view and single-image averages are 81.5 and 62.8; the Overall lead does not imply leadership on every numerical-question subset.

Ablation Study

The following component ablation is on VSI-Bench. Single-stage variants use rank-256 LoRA to control adaptation capacity, whereas the full model uses a merged rank-256 stage-1 update plus a rank-64 stage-2 adapter. These comparisons should not be understood as switching one component with entirely identical optimization histories.

Config Average Absolute distance Room size Relative direction Route planning
Full model 71.3 62.3 76.8 88.0 60.8
Full model without metric 3D embedding 70.4 58.1 71.3 88.6 66.0
Panorama + metric 3D embedding, no first stage 68.5 56.6 74.1 91.6 45.4
Panorama only, no first stage 65.3 54.8 56.9 87.4 45.3
Multi-view backbone, no proposed components 65.5 53.2 68.6 77.4 43.9

The curriculum improves route planning from 45.4 โ†’ 60.8, but removing the metric embedding raises that score to 66.0; one cannot claim that every component monotonically improves every task. More consistent evidence for the metric embedding comes from absolute distance falling from 62.3 โ†’ 58.1 and room size from 76.8 โ†’ 71.3 when it is removed from the full model.

The appendix also directly tests whether continuous tokens and RGB panoramas are interchangeable. Each comparison must be read within its own protocol:

Analysis protocol Config VSI-Bench Evidence boundary
Same checkpoint, inference substitution only Continuous tokens, 9,600 tokens 71.1 Appendix Table 8 baseline; not replaced with the main-table 71.3
Same checkpoint, inference substitution only \(181\times362\) grid, 7,144 tokens 68.2 Cell-wise mean pooling, without retraining
Same checkpoint, inference substitution only \(45\times90\) grid, 1,563 tokens 65.8 Stronger compression, not an upper bound for a retrained grid model
Matched 2,000-step short training Feature canvas 60.8 Matched to the next row, not full-training performance
Matched 2,000-step short training z-buffer RGB panorama 52.1 Changes both visibility retention and vision-tower input

The continuous baseline in Table 8 is 71.1 rather than the main-table 71.3, and its appearance-order score is 62.6 rather than 65.7. The cache does not explicitly explain these differences, so the original values are retained separately. The finest grid reduces appearance order to 39.8, indicating that losing multiple frame observations within a cell is a major risk. Selecting an original patch per cell instead scores 68.6 overall and does not eliminate that problem.

Key Findings

  • Reference-frame choice follows question semantics: with the same SQA3D checkpoint, agent-pose centering gives 65.3 EM@1, scene-center placement 61.2, and outside-scene placement 59.8. Which improves from 56.7 at the scene center to 72.6 at the agent pose. This does not involve another model or extra training.
  • The geometry-substitution conclusion is scoped: DA3 reduces the three aggregate scores by 0.4, 1.3, and 0.8 points, supporting robustness to this particular feed-forward estimator rather than immunity to arbitrary reconstruction failure.
  • Seed results do not replace the main results: three full-training replications give 70.4 ยฑ 0.6 on VSI-Bench, 65.1 ยฑ 0.9 on SQA3D, and 73.2 ยฑ 0.3 on SPBench. These quantities capture seed variation under a fixed recipe only.
  • The compute claim requires specifying competitors: 8 A6000 GPUs for approximately 64.8 hours yield 518 raw GPU-hours, analytically normalized to 259 A100-equivalent GPU-hours. SpaceMind uses 4,800 and SenseNova-SI 27,648, whereas VLM-3R uses 240, slightly less than OneCanvas. Normalization uses peak BF16 throughput, not measured cross-hardware speedups.

Highlights & Insights

  • A panorama is an address space, not necessarily a rendered image. Reassigning positions without re-encoding pixels avoids asking a perspective-pretrained vision tower to read distorted panoramas and retains surfaces observed by cameras even when occluded from the current canvas viewpoint.
  • Angles, metric position, and observation time have distinct roles. Native RoPE expresses direction and frame order, while a small metric encoding in feature content expresses actual scale; this division explains the ablations more precisely than simply saying that the method โ€œadds 3D positional encoding.โ€
  • The curriculum shares the representation's reading pathway. Synthetic boxes do not aim for realistic category appearance; real patch distributions establish object binding while geometry determines labels. This approach could inform other spatial tasks dominated by category priors.

Limitations & Future Work

  • Depth, intrinsics, and poses remain necessary. DA3 inference substitution narrows the gap between ground-truth geometry and RGB-based deployment, but does not cover all failure modes, including severe occlusion and reconstruction collapse. Further evaluation should trace how such errors propagate to angular addresses and metric embeddings.
  • The global canvas does not fuse repeated observations, preserving token cost as well as information. Fine structure and scalability in large or outdoor scenes remain acknowledged limitations; rasterization cannot directly replace the continuous representation without a cost.
  • Longitude is periodic, but RoPE position IDs are not periodically continuous. In the appendix's 15ยฐ distance probe, with objects fixed and only seam placement changed, error rises from 0.14 m to 1.23 m. Rotating the canvas can avoid a region of interest, and wrap-aware encoding is a possible direction, but the paper does not fully solve this issue.
  • Curriculum generators are hand-authored, so adding a new spatial skill requires new generation rules. ERNIE's 2,000-step short experiment without the curriculum supports representation portability only (40.2 โ†’ 46.3), not full-recipe performance for a non-Qwen model.
  • vs VLM-3R / SpaceMind: these methods explicitly fuse geometry-encoder outputs; OneCanvas uses geometry primarily for token addresses and a lightweight metric embedding while retaining the original visual encoding path. It simplifies geometric alignment architecture rather than removing geometry inputs.
  • vs PanoGrounder / LiftProj: rendering or stitching panorama pixels entails visibility handling and resampling; OneCanvas retains original perspective-frame features and co-located or occluded observations. The 2,000-step comparison changes both visibility and visual input, so its 8.7-point gain cannot be attributed to either difference alone.
  • vs SenseNova-SI / Cambrian-S: scaling spatial QA supervision improves capability; OneCanvas instead uses a shared representation to support procedural geometric tasks with controlled labels. The approaches are compatible, but additional real data should be checked for reintroducing category and answer priors.

Rating

  • Novelty: 4/5. Continuous panoramic addresses, content-side metric embedding, and a curriculum suppressing category shortcuts form a coherent unified approach.
  • Experimental Thoroughness: 4/5. Three benchmarks, component and origin ablations, RGB-predicted geometry, and appendix controls are included, but some cross-backbone comparisons use short training only.
  • Writing Quality: 4/5. Version 2 clearly explains the representation and comparison boundaries, while a few appendix/main-table numerical differences still require explanation.
  • Value: 4/5. The method offers a relatively lightweight spatial adaptation route for existing VLMs, with practical value still constrained by geometry reliability and scene scale.