PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation¶
Conference: ECCV 2026
Paper: CVF Open Access
Code: https://github.com/yyyyangyi/PUF
Area: 3D Vision
Keywords: scene graph generation, 3D scene graph, uncertainty estimation, Dirichlet distribution, plug-and-play
TL;DR¶
Addressing the failure of existing online 3D scene graph generation methods to account for observation truncation, 2D model ambiguity, and 3D spatial representation uncertainty under deterministic fusion pipelines, PUF introduces a training-free plug-and-play framework combining probabilistic node association, dual Dirichlet evidence accumulation, and class-conditional relation priors, boosting relationship recall on 3DSSG by 18.1 points while operating at a real-time latency of 15 ms per frame.
Background & Motivation¶
Online 3D scene graph generation aims to incrementally build a persistent, structured representation of an indoor environment—capturing 3D object bounding boxes, semantic classes, and pairwise spatial and semantic relationships—from an incoming stream of RGB-D video frames. This structured representation provides an indispensable foundation for high-level embodied AI tasks, such as visual navigation, robotic mobile manipulation, and spatial question answering. Offline 3D scene graph approaches typically operate on fully reconstructed 3D point clouds via dense global graph convolutional message passing; however, their heavy computational footprint prevents real-time interactive execution, and their dependence on dense 3D ground-truth graph annotations incurs prohibitive data acquisition costs. Recently, online 2D-to-3D lifting paradigms, exemplified by FROSS, have emerged by pairing fast 2D scene graph detectors with 3D Gaussian back-projections, enabling faster-than-real-time performance without requiring dense SLAM environmental mapping.
Nevertheless, existing online 2D-to-3D merging pipelines adhere strictly to deterministic, hard-decision heuristics, which systematically discard three critical sources of uncertainty. First, partial observation and sensor view truncation cause substantial noise, especially for inter-object relationships that require both participating entities to be simultaneously well-observed within the same video frame. Second, 2D scene graph models produce soft probability distributions over classes and predicates, which are routinely collapsed into hard \(\operatorname{argmax}\) one-hot predictions, discarding valuable epistemic uncertainty. Third, back-projecting noisy 2D depth patches into coarse 3D primitives (such as 3D Gaussians or voxel clusters) introduces inherent spatial and geometric ambiguity. When previous pipelines perform node association using rigid spatial distance or intersection thresholds followed by one-hot label overwriting, minor geometric inaccuracies prematurely veto correct associations, and no mechanism exists to redistribute relational evidence across plausible neighboring candidates.
To overcome these structural limitations, the core idea of this paper is to reformulate online 2D-to-3D scene graph fusion as a principled, training-free uncertainty-aware framework that models joint spatial-semantic likelihoods for probabilistic node association, employs Dirichlet evidence accumulation to softly distribute semantic and relational evidence across candidate nodes, and leverages a class-conditional Bayesian relation prior to complete sparsely observed edges.
Method¶
Overall Architecture¶
PUF processes a continuous stream of RGB-D frames \(\{(I_t, D_t, T_t)\}\) with known camera poses to incrementally maintain a global directed 3D scene graph \(\mathcal{G}=(\mathcal{V}, \mathcal{E})\) without retaining historical video frames. At each incoming time step, a real-time 2D scene graph model (comprising an RT-DETRv2-M detector and an EGTR relational transformer) extracts soft semantic class distributions \(\hat{p}_{obs}^c \in \Delta^{C-1}\) and pairwise relationship probability distributions \(\hat{p}_{obs,j}^r \in \Delta^{R-1}\) for detected 2D entities. Next, the 2D bounding boxes are lifted into 3D observation nodes using either 3D Gaussian primitives \(\Omega=(\mu, \Sigma)\) or discrete voxel occupancy sets \(\Psi\). Finally, the uncertainty-aware fusion stage computes joint spatial and semantic association likelihoods against existing global nodes alongside a birth probability, updates global Dirichlet evidence accumulators for semantics and predicates, and applies a Bayesian class-conditional relation prior to enhance poorly observed edges.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input RGB-D stream & camera poses"] --> B["2D SGG soft distributions & 3D lifting"]
B --> C["Probabilistic node association<br/>factored semantic-spatial likelihood & independent marginalization"]
C -->|Association prob beta <= 0.5 & above threshold| D["Dirichlet evidence accumulation<br/>semantic soft allocation & spatial update decoupling"]
C -->|Birth prob beta_birth > 0.5| E["Initialize new global nodes & edges"]
D --> F["Class-conditional relationship prior<br/>conjugate prior for sparse co-visibility completion"]
E --> F
F --> G["Output persistent uncertainty-aware 3D scene graph"]
Key Designs¶
1. Probabilistic node association: Factored joint likelihood and independent marginalization
Conventional pipelines rely on hard Euclidean distance or spatial overlap thresholds to decide whether a newly detected object corresponds to an existing 3D node, causing brittle tracking failures whenever depth noise induces slight spatial misalignments. PUF models node association via a continuous joint likelihood combining spatial overlap and semantic alignment: $\(L[obs, k] = L_{\text{sp}}(obs, k) \cdot L_{\text{se}}(obs, k)\)$ The spatial factor \(L_{\text{sp}}(obs, k)\) adapts to the underlying 3D representation: for the 3D Gaussian backend, it evaluates the continuous Bhattacharyya coefficient \(BC = 1 - H^2\) (where \(H\) is the Hellinger distance between two Gaussian distributions); for the voxel backend, it measures the containment score \(|\Psi_{obs} \cap \Psi_k| / |\Psi_{obs}|\) of the observation voxel set within the global node. The semantic factor \(L_{\text{se}}(obs, k)\) evaluates the Jensen-Shannon divergence (JSD) between the observation class distribution \(\hat{p}_{obs}^c\) and the normalized Dirichlet posterior mean \(\bar{\alpha}_k\) of global node \(k\): $\(L_{\text{se}}(obs, k) = \exp\left(-\frac{\text{JSD}(\hat{p}_{obs}^c, \bar{\alpha}_k)}{\sigma_{\text{se}}}\right)\)$ where \(\sigma_{\text{se}}\) serves as the semantic selectivity bandwidth. Rather than employing computationally prohibitive joint hypothesis enumeration as in classical JPDAF, PUF leverages the DETR non-maximum suppression property (which guarantees no duplicate detections in a single frame) to independently marginalize candidate likelihoods alongside an object birth parameter \(\lambda_{\text{birth}}\). When \(\beta_{\text{birth}} > 0.5\), a new 3D global node is spawned; otherwise, the observation evidence is softly distributed across all candidate nodes meeting \(\beta[obs, k] \ge \beta_{\text{min}}\).
2. Dirichlet evidence accumulation: Decoupling semantic soft allocation from spatial geometry update
To retain prediction uncertainty over sequential observations without collapsing distributions, PUF parameterizes semantic class distributions and directed edge relationship probabilities with Dirichlet distributions: \(\hat{p}_k^c \sim \operatorname{Dir}(\boldsymbol{\alpha}_k)\) and \(\hat{p}_{ij}^r \sim \operatorname{Dir}(\boldsymbol{\phi}_{ij})\), where \(\boldsymbol{\alpha}_k \in \mathbb{R}^C\) and \(\boldsymbol{\phi}_{ij} \in \mathbb{R}^R\) serve as non-negative evidence accumulators. When an observation is successfully associated with the global scene graph, evidence is softly accumulated across candidate nodes according to their association marginal weights: $\(\boldsymbol{\alpha}_k \leftarrow \boldsymbol{\alpha}_k + \beta[obs, k] \cdot \hat{p}_{obs}^c, \quad \forall k: \beta[obs, k] \ge \beta_{\text{min}}\)$ Edge relationship accumulators \(\boldsymbol{\phi}_{kj}\) and \(\boldsymbol{\phi}_{ik}\) are concurrently updated using identical association weights combined with current-frame 2D relation distributions \(\hat{p}_{obs,j}^r\) and \(\hat{p}_{i,obs}^r\). Crucially, PUF introduces a strict decoupling between semantic soft updating and spatial geometric updating: while semantic evidence can be validly divided among competing object candidates (e.g., an ambiguous object initially manifesting as 60% chair and 40% table), physical objects occupy unique, non-splittable spatial extents. Distributing spatial geometry across multiple candidate nodes would induce catastrophic geometric drift and ghosting. Therefore, 3D spatial properties (Gaussian centroid and covariance, or voxel union) are updated exclusively for the single best candidate \(k^* = \arg\max_k \beta[obs, k]\).
3. Class-conditional relationship prior: Conjugate prior for sparse co-visibility completion
Indoor scanning trajectories suffer from severe field-of-view constraints: in the 3DSSG benchmark, 25.3% of ground-truth object pairs have fewer than 10 joint frame co-observations, and 2.7% never appear simultaneously within any single camera view. To address this structural sparsity, PUF leverages Dirichlet-multinomial conjugacy by formulating an informative Bayesian relationship prior \(\boldsymbol{\phi}_{ij}^{\text{prior}}\) factorized into three distinct terms: $\(\boldsymbol{\phi}_{ij}^{\text{prior}} = P_{\text{cl}}(r \mid c_i, c_j) \cdot P_{\text{sp}}(\mu_i, \mu_j) \cdot P_{\text{ex}}(c_i, c_j)\)$ Here, \(P_{\text{cl}}(r \mid c_i, c_j)\) denotes the class-conditional predicate probability precomputed from training set annotations with Laplace smoothing (\(\varepsilon = 0.1\)); \(P_{\text{sp}}(\mu_i, \mu_j) = \exp(-\|\mu_i - \mu_j\|_2)\) computes an online exponential spatial decay based on the 3D Euclidean distance between object centroids; and \(P_{\text{ex}}(c_i, c_j)\) is an empirical existence gate indicating the training-set probability that class pair \((c_i, c_j)\) exhibits any relation, effectively suppressing false-positive relation hallucinations between physically unrelated entities. For observed edges, the final distribution is given by the Bayesian posterior \(\boldsymbol{\phi}_{ij}^{\text{post}} = \boldsymbol{\phi}_{ij}^{\text{prior}} + \boldsymbol{\phi}_{ij}\); for pairs that are never co-observed online, if \(\max(\boldsymbol{\phi}_{ij}^{\text{prior}}) > 0.5\), the prior autonomously completes the missing topological edge.
Loss & Training¶
PUF is a completely training-free inference algorithm that requires no gradient-based learning or architectural parameter optimization. The underlying 2D scene graph network weights are frozen: on 3DSSG, following FROSS, the model uses an RT-DETRv2-M detector paired with EGTR trained on 2D scene graph extractions from the 3DSSG training set; on ReplicaSSG, the framework evaluates zero-shot transfer using weights trained exclusively on Visual Genome without any in-domain adaptation. The statistical prior components \(P_{\text{cl}}\) and \(P_{\text{ex}}\) are acquired via a single lightweight offline counting pass across training annotations.
Key Experimental Results¶
Main Results¶
PUF was evaluated on the standard indoor 3DSSG benchmark and the zero-shot ReplicaSSG test benchmark. The evaluation employs official recall metrics, including relationship recall (Rel. Recall), object detection recall (Obj. Recall), predicate classification recall given detected objects (Pred. Recall), and mean recall (mRecall) across categories to assess performance on long-tail predicates.
Performance comparisons on the 3DSSG test set across diverse input modalities:
| Method | Input Modality | Rel. R@1 (%) ↑ | Obj. R@1 (%) ↑ | Pred. R@1 (%) ↑ | Obj. mR@1 (%) ↑ | Pred. mR@1 (%) ↑ | Latency (ms) ↓ |
|---|---|---|---|---|---|---|---|
| 3DSSG [32] | Point Cloud | 12.9 | 37.4 | 22.0 | 26.2 | 14.4 | - |
| VL-SAT [33] | Point Cloud | 23.5 | 53.7 | 28.9 | 42.1 | 25.3 | - |
| OCRL [10] | Point Cloud | 25.2 | 58.5 | 30.1 | 49.6 | 27.1 | - |
| SGFN [35] | RGB-D + SLAM | 22.0 | 51.6 | 27.5 | 37.7 | 24.0 | 245 |
| MonoSSG [34] | RGB-D + SLAM | 23.3 | 53.8 | 28.4 | 43.8 | 26.6 | 283 |
| SCRSSG [40] | RGB-D + SLAM | 25.7 | 61.1 | 27.6 | 60.5 | 27.8 | 350 |
| VGfM [8] | RGB | 19.6 | 50.0 | 20.4 | 34.8 | 11.0 | 379 |
| IMP [37] | RGB-D | 19.7 | 49.5 | 20.9 | 34.7 | 13.8 | - |
| Kim et al. [13] | RGB-D | 9.1 | 59.0 | 7.1 | 51.0 | 8.0 | 454 |
| FROSS [11] | RGB-D | 27.9 | 62.5 | 33.2 | 63.8 | 18.1 | 13 |
| PUF-Voxel (Ours) | RGB-D | 40.3 | 65.5 | 46.1 | 64.1 | 21.8 | 31 |
| PUF-Gaussian (Ours) | RGB-D | 46.0 | 69.7 | 51.4 | 65.8 | 28.2 | 15 |
Zero-shot transfer evaluation on the ReplicaSSG test set (evaluated strictly without relation prior):
| Method | Rel. R@1 (%) ↑ | Obj. R@1 (%) ↑ | Pred. R@1 (%) ↑ | Obj. mR@1 (%) ↑ | Pred. mR@1 (%) ↑ | Latency (ms) ↓ |
|---|---|---|---|---|---|---|
| FROSS [11] | 22.5 | 26.2 | 28.0 | 29.1 | 20.6 | 14 |
| PUF-Voxel (Ours) | 22.4 | 27.9 | 26.9 | 30.2 | 17.9 | 30 |
| PUF-Gaussian (Ours) | 25.3 | 31.0 | 35.6 | 33.7 | 26.2 | 16 |
Ablation Study¶
A component-wise ablation on the 3DSSG test set across both Gaussian and voxel backends demonstrates the orthogonal contributions of Dirichlet node accumulation (with probabilistic association), Dirichlet edge accumulation, and the relationship prior:
| Node Accum. | Edge Accum. | Rel. Prior | Gaussian Rel. R@1 (%) | Gaussian Obj. R@1 (%) | Gaussian Pred. R@1 (%) | Gaussian Latency (ms) | Voxel Rel. R@1 (%) | Voxel Obj. R@1 (%) | Voxel Pred. R@1 (%) | Voxel Latency (ms) |
|---|---|---|---|---|---|---|---|---|---|---|
| ✗ | ✗ | ✗ (FROSS) | 27.9 | 62.4 | 33.2 | 13.2 | 23.7 | 61.6 | 29.5 | 28.9 |
| ✓ | ✗ | ✗ | 31.3 | 69.0 | 36.6 | 14.6 | 28.6 | 64.9 | 33.4 | 31.0 |
| ✓ | ✓ | ✗ | 33.9 | 69.0 | 39.2 | 14.7 | 31.3 | 64.9 | 35.2 | 31.3 |
| ✗ | ✓ | ✗ | 27.9 | 62.9 | 33.7 | 13.5 | 24.2 | 61.8 | 30.6 | 29.2 |
| ✗ | ✓ | ✓ | 41.4 | 62.9 | 47.8 | 13.5 | 37.2 | 61.8 | 40.0 | 29.2 |
| ✓ | ✓ | ✓ (Full Model) | 46.0 | 69.7 | 51.5 | 14.8 | 40.3 | 65.5 | 46.1 | 31.3 |
Key Findings¶
- Clear and cumulative individual gains: Incorporating Dirichlet node modeling with probabilistic association alone lifts relationship recall from 27.9% to 31.3% (+3.4%) and object recall from 62.4% to 69.0% (+6.6%) on the Gaussian backend. Adding Dirichlet edge modeling pushes relationship recall to 33.9% (+2.6%), while enabling the class-conditional prior yields an overall relationship recall of 46.0% (+18.1% over FROSS).
- Representation-agnostic generalization: Identical monotonic improvements occur across both the compact 3D Gaussian representation and the discrete 3D voxel grid, establishing that PUF's probabilistic formulation operates independently of the geometric lifting primitive.
- Behavior across observation sparsity: Co-observation stratification reveals that the relationship prior delivers its largest marginal improvements in the sparsest bins (pairs observed <10 frames), smoothly giving way to empirical evidence as co-observations accumulate. Furthermore, PUF without prior ("w/o Prior") still consistently outperforms FROSS across all co-observation densities.
- Hyperparameter stability: Grid searches over semantic bandwidth \(\sigma_{\text{se}}\) and birth density \(\lambda_{\text{birth}}\) reveal that \((\sigma_{\text{se}}=0.3, \lambda_{\text{birth}}=0.4)\) provides optimal performance across both 3DSSG and ReplicaSSG benchmarks, demonstrating strong insensitivity to fine tuning.
Highlights & Insights¶
- Lossless end-to-end uncertainty propagation: Unlike prior online 3D SGG pipelines that prematurely collapse 2D continuous distributions via \(\operatorname{argmax}\) operators, PUF preserves probability distributions within Dirichlet accumulators, providing well-calibrated normalized predictive entropy for downstream robotic filtering and active re-querying.
- Decoupled semantic and geometric updates: PUF resolves the fundamental conflict between probabilistic evidence sharing and physical entity uniqueness by permitting semantic class distributions to blend across candidate tracks while strictly binding spatial updates to the argmax candidate, effectively eliminating multi-instance geometric ghosting.
- Instantaneous training-free deployment: Operating as an external probabilistic wrapper around any 2D scene graph model, PUF incurs negligible computational overhead (adding only 1–2 ms per frame to reach 15 ms latency), running an order of magnitude faster than conventional SLAM-based offline graphs.
Limitations & Future Work¶
- Reliance on empirical statistical co-occurrence: The class-conditional relationship prior depends heavily on training set co-occurrence statistics, which may introduce domain bias or become unavailable in unstructured open-vocabulary scenarios.
- Rigid static world assumption: Both the 3D lifting backends and the persistent Dirichlet accumulators assume stationary indoor entities, lacking explicit mechanisms to handle dynamic object displacements or temporal relationship mutations during physical human-object interactions.
- Future directions: Integrating zero-shot common-sense relational reasoning via open-vocabulary vision-language models could replace offline statistical priors, and introducing temporal decay factors into the Dirichlet accumulators would facilitate lifelong dynamic scene graph maintenance.
Related Work & Insights¶
- vs FROSS [11]: FROSS pioneered real-time online 3D SGG by coupling 2D SGG with 3D Gaussian lifting, but relied on rigid Euclidean association thresholds and deterministic label overwrites; PUF retains its fast lifting efficiency while introducing probabilistic data association and Dirichlet accumulation, nearly doubling relationship recall (46.0% vs 27.9%) at comparable speed (15 ms vs 13 ms).
- vs SGFN [35] & SCRSSG [40]: SGFN and SCRSSG rely on dense point cloud reconstruction via SLAM and global graph message passing, incurring heavy latencies of 245–350 ms per frame; PUF completely bypasses dense reconstruction, achieving substantially higher relational accuracy under purely incremental streaming conditions.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates the first comprehensive taxonomy of online 3D SGG uncertainties and proposes an elegant Dirichlet-based probabilistic fusion framework.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Exhaustive evaluations across 3DSSG and ReplicaSSG, featuring dual Gaussian/voxel backends, complete ablation matrices, and co-visibility density breakdowns.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical modeling, intuitive motivation, coherent structure, and high-quality figures.
- Value: ⭐⭐⭐⭐⭐ A plug-and-play, training-free, and real-time paradigm that significantly narrows the gap between 2D perception and embodied 3D understanding.