SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding¶
Conference: NeurIPS2026
arXiv: 2609.33518
Area: 3D Vision
Keywords: scene states, visual bottleneck, spatial references, role constraints, unified scene understanding
TL;DR¶
SceneScaffold allocates a fixed budget of 100 visual tokens to details, entities, spatial references, relations, and a global summary, organizing scene evidence before language reasoning and improving ScanRefer / Multi3DRefer mIoU from 43.3 / 42.7 to 47.0 / 47.9 over 3D-LLaVA.
Background & Motivation¶
Unified 3D large multimodal models use one language model for object grounding, question answering, and object description, but point clouds contain far more superpoint evidence than the language model can receive as visual tokens. Methods such as 3D-LLaVA compress inputs through selection or aggregation. The resulting object features may indicate which objects exist without preserving scene boundaries, occupied regions, or how objects are located relative to those references. In rooms containing multiple objects of the same category, correct category recognition is insufficient for selecting the intended instance.
The paper argues that the problem is not only an insufficient token budget, but also an inappropriate division of responsibilities between vision and language. The visual side retains semantically salient object fragments, while the language side must recover missing spatial organization from a flat sequence while performing task reasoning. Question answering needs global layout, whereas grounding needs fine-grained references; improving object-token salience alone does not ensure that both survive compression. SceneScaffold therefore lets different states read different evidence sources and stabilizes their responsibilities through training objectives, without increasing the final token count.
Here, active means that the visual connector actively constructs a representation, not that an agent moves, selects viewpoints, or explores an environment in real time. A scene-state is a latent state computed in one forward pass over a static scene, not a world model maintained across time. Core Idea: construct an entityโspatial-referenceโrelation scaffold on the visual side before passing a fixed-length sequence with explicit role identities to the language model.
Method¶
Overall Architecture¶
The input is a point cloud, from which a frozen 3D U-Net extracts superpoint features and their 3D centroids. The connector preserves existing object details while separately constructing object-core evidence and geometric anchors. Entity and scene-frame readers aggregate their respective evidence, and the relation reader then reads the two constructed state types. A global summary is formed, and the representation is projected and serialized into 100 visual tokens for different tasks handled by the same LLM.
The default allocation, ordered as residual details / entities / scene frames / relations / global summary, is 75 / 8 / 8 / 8 / 1. The 128 entity-evidence superpoints are candidates consumed by the reader, not 128 additional LLM input tokens. The number of frame-evidence anchors also varies with anchor validity, but the output always contains 8 frame states. Residual details enter the final sequence directly, without passing through the relation reader.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
X["Point-cloud encoding<br/>Features and centroids"] --> E["Entity and Detail<br/>Preservation"]
X --> F["Geometric Scene Frames"]
E -->|Entity states| R["EntityโFrame<br/>Relation Reading"]
F -->|Frame states| R
E -->|75 details and 8 entities| Z["Shared State Interface"]
F -->|8 frames| Z
R -->|8 relations| Z
Z --> O["LLM text output<br/>or SEG mask decoding"]
T["Training-only supervision<br/>Semantics and geometry"] -.-> E
T -.-> F
T -.-> R
Key Designs¶
1. Entity and Detail Preservation: preserve object-core information separately from local discriminative cues
The entity branch uses a shared lightweight MLP to score each superpoint without conditioning on user language, and selects 128 object-core evidence tokens through hard TopK. Eight learnable entity queries read this evidence through cross-attention to construct scene-conditioned entity states. The query parameters are learned, but their outputs depend on the scene. The states are not defined as one detected object per slot, nor are the 8 slots required to enumerate all instances.
Meanwhile, 75 tokens from the original visual connector remain as a residual object buffer, preserving small-object evidence and appearance differences. This division avoids averaging all details into a few entity states: the entity branch carries reusable object semantics, while the residual branch supplies fine-grained evidence. During training, the residual buffer receives dropout at a rate of 0.2, reducing the chance that the LLM bypasses the state branches entirely. This dropout is disabled during inference.
2. Geometric Scene Frames: provide spatial references that salience-based selection might discard
Scene frames do not select the most object-like superpoints. Instead, they construct anchors from the horizontal coordinates of their centroids. The method first obtains the scene's x and y extents and averages features within left, right, front, and back boundary bands, each with a default width of 0.15 of the corresponding axis extent. It then divides the horizontal plane into a 3ร3 grid and aggregates features in occupied cells. Valid anchors require at least 4 superpoints. Boundary anchors are prioritized, with the remaining budget assigned to region anchors; when the budget is insufficient, regions with more superpoints are retained first.
Eight frame queries read these anchors rather than the entity branch's TopK evidence. Consequently, boundary and layout evidence has an independent representation pathway even when its semantic salience is low. Left, right, front, and back are defined by scene-coordinate extents, not by reconstructing an observer frame for each question. Boundary anchors are not independently detected physical walls either. These coarse horizontal references help represent indoor layouts, but are not a complete 3D topology.
Entity, frame, and relation readers each use one-layer, 4-head cross-attention with independent queries and reader parameters. The appendix specifies the reading mechanism as:
Queries read evidence as keys and values, while the residual connection and normalization preserve query role identity. Attention dropout is 0.0 and the temperature is 1.0.
3. EntityโFrame Relation Reading: derive relations from organized objects and references
Relation states do not directly read the raw point cloud or all superpoints. Once entity and frame states are formed, the relation reader reads their concatenation:
The default 8 relation queries consume 8 entity states and 8 frame states. Relation aggregation thus has structural access to both object semantics and spatial references, rather than inferring location from another collection of salient fragments. The appendix's control baseline that constructs relations directly from primitive evidence performs worse, supporting this induction pathway empirically. It does not prove that each relation slot learns an individually nameable relation.
Relation supervision is also deliberately coarse. Selected entity evidence is weighted by normalized scores to compute one aggregate centroid. The nearest horizontal boundary and the grid cell containing this centroid provide boundary and region labels. Pooled relation states predict these two labels. These are automatically derived geometric references, not ground-truth pairwise annotations of support, contact, or occlusion, and they do not guarantee recovery of all inter-object relations.
4. Shared State Interface: preserve structure in language inputs and mask queries
Entity, frame, and relation states are pooled separately, concatenated, and mapped into 1 global summary token. This is a compact summary of the three state types, not an independent branch that rereads the point cloud. Residual, entity, frame, relation, and summary blocks are then concatenated in a fixed order, mapped into the LLM embedding space by the same cross-modal projector, and augmented with a learnable role embedding for each block. Role embeddings are initialized to zero; the method does not introduce five independent projectors.
Question answering and captioning use this shared interface directly. Grounding retains standard [SEG] query generation and mask decoding, while adding a projected concatenation of pooled entity, frame, and relation states to the segmentation query as a residual condition. Its learnable scale is initialized to 0.0. This condition gives the decoder access to the same spatial summaries as the LLM, but state tokens neither replace dense visual features nor directly produce point-level masks.
A Worked Example¶
Consider โthe chair near the room boundaryโ as an illustrative input, not an additional experimental example from the paper. Chair evidence may enter the 128 entity candidates and then be aggregated into 8 entity states. The geometric branch independently reads four boundary bands and valid grid anchors to produce 8 frame states. The relation reader aggregates objectโreference information from these 16 states; it neither moves an agent nor searches an explicit chairโwall relation graph.
The final 75 detail tokens preserve appearance differences between candidate chairs, while 24 role states and 1 summary contribute organized information. After the user expression reaches the LLM, a [SEG] query is generated and adjusted using the pooled-state condition before the standard mask-decoding pathway localizes the target. This example illustrates data flow, not a guarantee that the aggregate entity centroid corresponds to the queried chair. The auxiliary geometry labels themselves are independent of that user expression.
Loss & Training¶
Task supervision includes a language modeling loss and grounding losses consisting of BCE, Dice, and decoder auxiliary losses. Role-preserving supervision constrains entity semantics, frame coverage, and relation geometry. A projected pooled entity state is aligned by cosine similarity with an evidence target weighted by softmax-normalized entity scores. Two classification heads on pooled relation states predict the boundary and region labels described above.
Frame coverage does not require every slot to read all anchors uniformly. For each valid group, attention to its anchors is summed within each state slot, followed by a maximum across slots: coverage is sufficient when at least one slot attends strongly to the group. The paper defines:
Boundary directions and valid grid cells form the group set. The loss penalizes ignored groups and encourages slots to share spatial coverage rather than focusing only on one high-response location. The overall objective retains the paper's compact formulation:
The appendix reports weights of 0.05 for entity, frame, relation-boundary, and relation-region terms; the relation term contains two cross-entropies. Hard TopK does not relax the discrete index selection, and gradients propagate through selected feature paths. Auxiliary semantic and geometry objectives are used only during training; inference requires no manually supplied boundary, region, or relation labels.
The model uses LLaVA-1.5-7B and follows 3D-LLaVA's unified instruction-tuning protocol. The point-cloud encoder and LLM main body are frozen; state-construction modules, cross-modal projections, segmentation-query projections, and LoRA parameters are trained. Training uses 8 NVIDIA A800 GPUs, AdamW, a cosine learning-rate schedule, an initial learning rate of \(2\times10^{-4}\), batch size 2, 1 epoch, and 8 gradient-accumulation steps.
Key Experimental Results¶
Main Results¶
The following comparison includes only point-cloud-only generalist models, not image-using PC+I methods or task specialists in a common-protocol ranking. Grounding uses mIoU; ScanQA reports validation CIDEr, SQA3D reports test EM / EM-R, and Scan2Cap reports validation [email protected], evaluating captions under an IoU 0.5 localization-matching threshold. Higher is better within each column; absolute magnitudes across different metrics are not comparable.
| Point-cloud generalist | ScanRefer mIoU | Multi3DRefer mIoU | ScanQA C | SQA3D EM / EM-R | Scan2Cap [email protected] |
|---|---|---|---|---|---|
| 3D-LLaVA | 43.3 | 42.7 | 92.6 | 54.5 / 56.6 | 78.8 |
| NDTokenizer3D | Not reported | 46.0 | 98.6 | 54.4 / 57.1 | 79.0 |
| SceneScaffold | 47.0 | 47.9 | 95.8 | 55.5 / 58.3 | 79.1 |
Relative to 3D-LLaVA, the grounding metrics improve by 3.7 / 5.2 points and ScanQA CIDEr improves by 3.2, but BLEU-4 decreases from 17.1 to 16.8. SceneScaffold also does not surpass NDTokenizer3D on every question-answering metric: the latter scores 98.6 ScanQA CIDEr versus 95.8 here. The Scan2Cap CIDEr improvement is only 0.3 and should not be described as comparable to the larger grounding gains.
Ablation Study¶
State ablations keep the final budget at 100 tokens. Removing 8 role slots reallocates them to detail tokens; removing the summary reallocates only 1 slot. These comparisons therefore test representation organization rather than increased input length.
| Config | ScanRefer mIoU | Multi3DRefer mIoU | ScanQA C | SQA3D EM | Scan2Cap [email protected] |
|---|---|---|---|---|---|
| Full model | 47.0 | 47.9 | 95.8 | 55.5 | 79.1 |
| Without entity states | 45.0 | 46.2 | 91.1 | 52.7 | 77.5 |
| Without frame states | 45.9 | 46.2 | 95.3 | 55.3 | 77.9 |
| Without relation states | 45.1 | 46.4 | 92.6 | 56.4 | 75.5 |
| Without global summary | 45.8 | 46.3 | 93.0 | 54.8 | 76.1 |
Removing entity states reduces ScanQA CIDEr by 4.7; removing relation states reduces Scan2Cap CIDEr by 3.6. However, the relation-free variant scores 56.4 SQA3D EM, above the full model's 55.5, while EM-R decreases from 58.3 to 57.8. This mixed result must remain explicit: โthe full model is more balancedโ does not mean every component improves every metric.
Same-budget construction controls also support the importance of evidence sources. The two grounding scores are 45.7 / 46.5 for generic query states, 45.9 / 46.4 for shared-evidence role states, 45.8 / 46.7 when object evidence replaces geometric frame evidence, and 45.7 / 46.6 when relations are induced directly from primitive evidence. All are below the full model's 47.0 / 47.9. Removing only the relation geometry loss while retaining the architecture gives 46.0 / 46.8 on ScanRefer / Multi3DRefer, demonstrating that loss removal and state removal are different interventions.
Key Findings¶
The Rel.+Amb. subset intersects two criteria: explicit spatial or scene-reference words in the text, and multiple same-category candidates in the scene. It is not a manually annotated test of pure relational reasoning. The following table retains subset sizes and comparisons on the same subsets; these scores must not be cross-subtracted from full-split scores.
| Dataset and metric | Subset / full sample count | 3D-LLaVA subset | Ours subset | Subset gain |
|---|---|---|---|---|
| ScanRefer mIoU | 6329 / 9508 | 37.6 | 40.6 | 3.0 |
| Multi3DRefer mIoU | 6235 / 11120 | 40.7 | 43.5 | 2.8 |
| ScanQA C | 2272 / 4675 | 84.9 | 88.6 | 3.7 |
| Scan2Cap [email protected] | 1631 / 2007 | 76.0 | 77.9 | 1.9 |
- All subsets improve, but grounding gains of 3.0 / 2.8 are smaller than full-split gains of 3.7 / 5.2. The conclusion's summary of larger gains on difficult subsets cannot be generalized unconditionally to every task.
- Replacing frame states in the trained full model with corresponding states from another scene reduces subset grounding from 40.6 / 43.5 to 38.4 / 41.3. This shows functional use of scene-dependent states, but the authors explicitly do not treat it as a strict causal explanation of LLM reasoning.
- Replacing all states across scenes yields four subset scores of 37.8 / 40.8 / 85.0 / 76.0, with residual details retained. Compared with retraining after component removal, this more directly tests whether the complete model relies on the states.
Highlights & Insights¶
- The budget constraint shifts from selecting which tokens survive to reserving capacity for distinct evidence types. Independent boundary and region pathways reduce the risk that semantic salience crowds out spatial references.
- Roles are jointly defined by evidence sources, reading order, and supervision, not merely by labeling homogeneous tokens. Shared-evidence and generic-query controls in the appendix make this argument stronger than module-removal experiments alone.
- Residual detail is a complement to structural states rather than an obsolete interface to eliminate. Separating local discrimination from layout references is a design hypothesis worth testing in other constrained visual interfaces.
Limitations & Future Work¶
- The authors acknowledge that experiments focus on static indoor, ScanNet-style environments. Outdoor scenes, dynamic settings, long-horizon interaction, and noisy real-time perception remain untested; the word active does not imply these capabilities.
- Horizontal boundary bands and a 3ร3 grid are only a coarse spatial scaffold, and one weighted entity center can mix several objects. Per-entity geometry supervision and vertical references are possible research directions, not improvements demonstrated here.
- No multi-seed results, error bars, or confidence intervals are reported, leaving the stability of small captioning gains unresolved. A fixed token budget also does not imply equal computation; the paper does not provide latency data sufficient to compare the added readers' overhead.
- The source gives conflicting code-release statements: the abstract says GitHub availability, whereas the experimental setting says supplementary code will become public upon publication. No verifiable repository URL appears in the provided text, so this note does not invent a link.
Related Work & Insights¶
- vs 3D-LLaVA: SceneScaffold retains the superpoint, instruction-tuning, and mask-decoding foundations while reorganizing an equal-length visual interface into role states. Grounding gains are substantial, but not every question-answering or captioning metric improves.
- vs Scenes as Tokens / NDTokenizer3D: The latter constructs holistic scene tokens with a multi-scale NDT tokenizer, while SceneScaffold emphasizes heterogeneous evidence roles and an entityโframe induction pathway. NDTokenizer3D still leads on ScanQA CIDEr.
- vs LSceneLLM: Task-relevant region selection adapts to current requirements. SceneScaffold's entity scoring is not language-conditioned; its goal is a shared scene foundation for multiple tasks, not reselecting all role evidence for each question.
- vs task-specific relation models such as 3DVG-Transformer: SceneScaffold embeds lightweight relational information in a shared LLM interface rather than serving only a grounding head. Its supervision is correspondingly coarse and is not equivalent to an explicit complete scene graph.
Rating¶
- Novelty: 4/5 โ Evidence pathways and role supervision jointly constrain a fixed-budget scene interface, with a clear connection between concept and implementation.
- Experimental Thoroughness: 4/5 โ Five benchmarks, construction controls, state interventions, and loss ablations provide broad evidence, but statistical significance and cross-environment validation are missing.
- Writing Quality: 3/5 โ The main method and appendix support implementation-level reading, while code-release statements and some gain summaries require greater precision.
- Value: 4/5 โ A reusable representation-organization approach for unified 3D models, supported primarily by grounding and spatial-ambiguity experiments.