Skip to content

MessyKitchens: Contact-rich object-level 3D scene reconstruction

Conference: ECCV 2026
Paper: CVF Open Access
Project Page: https://messykitchens.github.io/
Area: 3D Vision
Keywords: 3D scene reconstruction, object-level reconstruction, physical contact constraints, Multi-Object Decoder, cluttered scene benchmark

TL;DR

Addressing the lack of physical plausibility and severe inter-object penetration in single-view object-level 3D reconstruction, this paper introduces the MessyKitchens benchmark with millimeter-level registration accuracy alongside a Multi-Object Decoder (MOD) that refines object poses via joint cross-attention, significantly enhancing contact fidelity and scene-level geometry.

Background & Motivation

Monocular 3D scene reconstruction has achieved dramatic progress in recent years. Empowered by massive foundation backbones and extensive datasets, state-of-the-art methods demonstrate exceptional capabilities in monocular depth estimation and scene-wide surface recovery. However, practical downstream applications—such as robotic manipulation, embodied navigation, physics simulation, and computer graphics animation—require far more than an unstructured global surface mesh. These domains inherently depend on decomposing complex, cluttered environments into discrete 3D object instances while faithfully capturing their fine-grained physical contacts, support relationships, and relative spatial poses. Existing object-level 3D reconstruction benchmarks (such as GraspNet-1B, HouseCat6D, and GraspClutter6D) are constrained by limited capture rigs or coarse multi-object registration workflows. As a result, their 3D ground truth meshes suffer from pronounced inter-object penetrations and unrealistic contact gaps, failing to provide a dependable benchmark for micro-contact accuracy and physical plausibility.

From a methodological standpoint, prevailing object-centric reconstruction pipelines either rely on constrained CAD retrieval or employ feed-forward single-object estimators on detected 2D bounding regions. For instance, the recent foundation model SAM 3D effectively infers 3D shapes and 7-DOF spatial poses from an image and corresponding object masks. Nevertheless, SAM 3D treats detected objects as mutually isolated tokens during inference, lacking end-to-end relational reasoning over spatial layouts, stacking patterns, and occlusions across multiple objects. This isolated prediction paradigm frequently causes overlapping predictions to interpenetrate each other or float unnaturally above supporting surfaces, compromising the physical validity and global structural coherence of reconstructed scenes.

To overcome the twin barriers of imprecise ground truth benchmarks and isolated single-object modeling, this paper advances object-level reconstruction on both fronts: first, by engineering a double-sided scanning rig with normal-consistent registration to curate MessyKitchens—a high-fidelity benchmark comprising 100 contact-rich real kitchen scenes; second, the core idea is to introduce a plug-and-play Multi-Object Decoder (MOD) that alternates multi-object self-attention with geometric cross-attention across shape and pose tokens, predicting scene-aware 7-DOF pose residuals to produce globally consistent and physically plausible multi-object 3D reconstructions.

Method

Overall Architecture

The framework accepts a single RGB image along with corresponding 2D instance segmentation masks as input. First, SAM 3D image encoders and mask-conditioning modules extract shape tokens and pose tokens for each of the \(N\) detected objects, yielding initial single-object voxel/mesh representations alongside initial 7-DOF poses \(p_i = (q_i, t_i, \sigma_i)\). Next, to enforce scene-level physical constraints and eliminate inter-object collisions without corrupting geometric fidelity, the aggregated shape tokens and pose tokens are passed to the Multi-Object Decoder (MOD). Through \(K\) transformer blocks comprising intra-object self-attention, multi-object pose self-attention, and pose-to-shape cross-attention, MOD decodes residual pose updates \(\tilde{p}_i = (\tilde{q}_i, \tilde{t}_i, \tilde{\sigma}_i)\) for all objects simultaneously. The refined poses are obtained by adding the residuals to the initial predictions, yielding an end-to-end physically coordinated 3D scene reconstruction.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Single RGB Image + 2D Instance Masks"] --> B["SAM 3D Backbone<br/>Single-object encoding & initial 7-DOF pose estimation"]
    B --> C["Aggregated Scene Token Sets<br/>Pose tokens Tp & Shape tokens Ts"]
    C --> D["Intra-Object Sequence Self-Attention<br/>Preserve individual object semantics & feature capacity"]
    D --> E["Cross-Object Global Pose Self-Attention<br/>Flattened tokens model inter-object spatial layouts"]
    E --> F["Geometry-Aware Multi-Object Cross-Attention<br/>Queries ground poses against 3D shape tokens"]
    F --> G["Linear Residual Pose Head<br/>Decode Δq, Δt, Δσ"]
    G --> H["Residual Summation & Scene Alignment<br/>Physically plausible multi-object 3D scene"]

Key Designs

1. Contact-Rich Benchmark Acquisition with Normal-Consistent Registration: Eliminating Penetration and Thin-Wall Ambiguities To address imprecise alignments and pervasive ground truth penetrations in earlier benchmarks, the data acquisition setup employs an optical scanning rig where kitchen items are placed on a transparent acrylic sheet. Utilizing double-sided reflective fiducial markers, complete scans are captured from above and beneath without disturbing object positions, delivering dense ground truth scans with point-to-mesh errors below 0.05 mm. Captured configurations span three systematic difficulty levels: Easy (4 well-separated objects), Medium (6 objects including flat-supported and stacked pairs), and Hard (8 objects with nested arrangements and maximal contact). For scene-level object registration, standard Euclidean distance minimization easily falls into local minima—especially on thin, concave tableware where top and bottom walls are within millimeters of each other. The proposed registration pipeline integrates a surface normal collinearity penalty into the optimization objective, penalizing opposing normals and preventing the reconstructed surface from settling between opposing walls. This reduces mean absolute registration depth error to 1.62 mm and lowers the penetration-to-contact area ratio to 0.14.

2. Multi-Object Pose-Shape Transformer Decoder: Resolving Floating and Inter-Object Penetration Because predicting objects independently ignores mutual clearance, MOD introduces \(K = 3\) stacked interaction blocks to reason over global context. Given aggregated pose tokens \(T^p \in \mathbb{R}^{N \times F_p \times C}\) and shape tokens \(T^s \in \mathbb{R}^{N \times F_s \times C}\) across \(N\) instances, each block first executes standard sequence self-attention on individual objects using \(N\) as the batch dimension to maintain individual feature expressiveness. The tokens are then flattened across instance and length dimensions into \(\hat{T}^p \in \mathbb{R}^{1 \times (N F_p) \times C}\) and fed to a multi-object self-attention layer \(\text{SA}_{\text{multi}}\), allowing every pose token to attend to all other objects in a single unified receptive field. Subsequently, a multi-object cross-attention layer \(\text{CA}_{\text{multi}}\) uses the updated pose tokens as queries and the flattened shape tokens \(T^s \in \mathbb{R}^{1 \times (N F_s) \times C}\) as keys and values. Conditioning pose refinement directly on 3D geometric boundaries enables the network to perceive bounding occupancy, contact interfaces, and supporting planes, adjusting 7-DOF parameters to avoid collision without altering the underlying mesh topology.

3. Symmetry-Robust Geometric Supervision and Double-Cover Quaternion Alignment: Stabilizing Rigid Optimization To supervise pose residual learning reliably, the training loss addresses dual geometric challenges: non-unique symmetric alignments and quaternion parameter ambiguities. To handle the double-cover property of quaternions where \(q\) and \(-q\) represent the same physical rotation, an inner-product squared loss \(\mathcal{L}_{\text{ip}} = 1 - \langle q, \hat{q} \rangle^2\) is applied, minimizing geodesic distance on the \(\mathrm{SO}(3)\) manifold without sign discontinuities. Furthermore, because kitchen items frequently exhibit continuous rotational symmetry around their central axis, point-to-point supervision introduces erroneous gradients; the model instead applies a bidirectional symmetric Chamfer Distance: $$ \mathcal{L}{\mathrm{CD}}(S, \hat{S}) = \frac{1}{|S|}\sum|}\min_{\hat{x} \in \hat{S}}|x - \hat{x2^2 + \frac{1}{|\hat{S}|}\sum|_2^2 $$ Dynamic nearest-neighbor assignment absorbs continuous symmetry ambiguity seamlessly. The overall multi-task training objective is defined as } \in \hat{S}}\min_{x \in S}|x - \hat{x\(\mathcal{L} = 0.1 \mathcal{L}_{\mathrm{CD}} + 100 \mathcal{L}_t + 100 \mathcal{L}_s + 10 \mathcal{L}_{\mathrm{ip}}\), steering fast convergence in translation, scale, and surface alignment.

Loss & Training

The Multi-Object Decoder is trained exclusively on MessyKitchens-synthetic, comprising 1,800 simulated scenes spanning Easy, Medium, and Hard configurations. Across these setups, 10,800 photorealistic images are rendered using Blender Cycles with camera azimuth \(\phi \in [0, 2\pi]\) and elevation \(\iota \in [\pi/4, \pi/2]\). Training is performed on 4 NVIDIA A100 (40GB) GPUs using an initial 10% linear warmup and a peak learning rate of \(5 \times 10^{-5}\) for 10 epochs (taking roughly 2 hours). With only ~81M parameters added to SAM 3D, MOD trains efficiently on token embeddings without requiring costly 3D volumetric backpropagation.

Key Experimental Results

Main Results

MOD is benchmarked on the real-world MessyKitchens test set alongside three unseen out-of-distribution (OOD) real datasets: GraspNet-1B, HouseCat6D, and GraspClutter6D in a zero-shot setting. Evaluations measure both object-level and scene-level Intersection-over-Union (IoU↑) and Chamfer Distance (CD↓), aligned via Sim(3) ICP across three initializations.

Dataset Evaluation Level PartCrafter MIDI SAM 3D MOD (Ours) Relative Gain over SAM 3D
MessyKitchens Object IoU↑ / CD↓ 0.071 / 0.495 0.186 / 0.285 0.409 / 0.064 0.445 / 0.061 +8.8% / -4.7%
Scene IoU↑ / CD↓ 0.133 / 0.228 0.238 / 0.165 0.431 / 0.054 0.472 / 0.050 +9.5% / -7.4%
GraspNet-1B Object IoU↑ / CD↓ 0.020 / 0.956 0.067 / 0.640 0.336 / 0.082 0.344 / 0.078 +2.4% / -4.9%
Scene IoU↑ / CD↓ 0.075 / 0.355 0.121 / 0.327 0.356 / 0.074 0.377 / 0.069 +5.9% / -6.8%
HouseCat6D Object IoU↑ / CD↓ 0.029 / 0.852 0.092 / 0.433 0.325 / 0.125 0.404 / 0.100 +24.3% / -20.0%
Scene IoU↑ / CD↓ 0.114 / 0.289 0.172 / 0.215 0.374 / 0.099 0.458 / 0.079 +22.5% / -20.2%
GraspClutter6D Object IoU↑ / CD↓ 0.067 / 0.674 0.086 / 0.500 0.328 / 0.105 0.340 / 0.103 +3.7% / -1.9%
Scene IoU↑ / CD↓ 0.227 / 0.227 0.213 / 0.201 0.487 / 0.059 0.496 / 0.058 +1.8% / -1.7%

Evaluating physical contact fidelity using the PhySIC contact graph protocol on MessyKitchens (Table 3):

Method Contact Precision (%) ↑ Contact Recall (%) ↑ Contact F1 Score (%) ↑
PartCrafter 27.2 17.4 21.2
MIDI 38.9 48.1 43.0
SAM 3D 66.0 48.3 55.8
MOD (Ours) 70.0 62.5 66.1

Ablation Study

1. Registration Strategy Ablation (Data Curation Accuracy)

| Registration Strategy | Mean Absolute Error \(\mu_{|\delta|}\) (mm) ↓ | Median Absolute Error \(\text{med}_{|\delta|}\) (mm) ↓ | Standard Deviation \(\sigma_\delta\) (mm) ↓ | Note | | :--- | :--- | :--- | :--- | :--- | | Rough Manual Alignment | 4.689 | 3.209 | 7.439 | Visual coarse initialization | | Distance Only | 2.892 | 2.141 | 4.822 | Prone to local minima on thin walls | | Distance + Normals (Ours) | 1.615 | 0.911 | 3.827 | Median error reduced by 57.5% vs. distance only |

2. MOD Attention Mechanism and Transformer Depth Ablation

Ablation Dimension Configuration Object IoU ↑ Object CD ↓ Scene IoU ↑ Scene CD ↓ Note
Feature Interaction Shape Only 0.438 0.063 0.463 0.054 Lacks global pose topology
Pose Only 0.432 0.065 0.457 0.056 Lacks 3D spatial occupancy bounds
Shape + Pose (S+P Ours) 0.445 0.061 0.472 0.050 Joint interaction achieves optimal performance
Number of Blocks \(K\) \(K = 1\) 0.441 0.061 0.467 0.051 Limited modeling capacity
\(K = 3\) (Default Ours) 0.445 0.061 0.472 0.050 Balances performance and efficiency
\(K = 6\) 0.403 0.065 0.427 0.059 Deeper blocks overfit and corrupt prior

Key Findings

  • Substantial Gains in Contact Accuracy: MOD raises the contact F1 metric from 55.8% to 66.1% (+10.3 percentage points), propelled primarily by a surge in recall from 48.3% to 62.5%. This verifies that contextual tokens enable the network to recover non-obvious support and nested contacts.
  • Robust Zero-Shot Out-of-Distribution Generalization: Although trained exclusively on synthetic kitchenware, MOD yields an impressive +24.3% object-level IoU improvement on HouseCat6D, confirming that learned physical support and spatial layout priors generalize effectively across diverse everyday object categories.
  • Compact Architecture Outperforms Deep Stacks: Increasing depth to \(K=6\) degrades object-level IoU to 0.403. Because SAM 3D already provides rich geometric representations, shallow non-linear adaptation (\(K=3\)) is optimal to avoid disrupting pre-trained features.

Highlights & Insights

  • Normal Consistency Unlocks High-Fidelity Ground Truth: Integrating normal alignment into ICP eliminates penetration artifacts across thin concave surfaces, establishing an ultra-clean 1.62 mm depth accuracy benchmark for physics-aware 3D vision.
  • Decoupled Pose Refinement Over Mesh Regeneration: Rather than attempting end-to-end multi-object regeneration from scratch, freezing object shapes while refining 7-DOF pose residuals via cross-attention delivers massive physical consistency gains with only 81M parameters.
  • Broad Transferability for Embodied Vision: The lightweight plug-in architecture can be directly retrofitted onto existing single-object detection and reconstruction backbones, preventing floating and collisions before executing robot manipulation.

Limitations & Future Work

  • Absence of Joint Non-Rigid Mesh Deformation: MOD adjusts rigid 7-DOF poses without deforming object meshes; severe initial geometric shape errors caused by heavy occlusions cannot be refined through contact pressure.
  • Dependency on Upstream 2D Segmentation: The framework relies on SAM 3 detection masks; missed detections or merged masks in extreme clutter prevent neglected instances from contributing to the global attention graph.
  • Future Directions: Integrating differentiable physics simulation into loss formulations and expanding rigid pose refinement to localized non-rigid mesh contact adaptations.
  • vs SAM 3D: SAM 3D provides strong single-object shape priors but lacks mutual spatial reasoning. MOD acts as a lightweight plug-in that directly resolves floating and collision artifacts.
  • vs PartCrafter & MIDI: Multi-instance diffusion models struggle significantly on real cluttered imagery (yielding IoU below 0.1). In contrast, MOD builds upon robust discriminative foundation priors, outperforming generative baselines by a wide margin.

Rating

  • Novelty: ⭐⭐⭐⭐ [Introduces a high-precision contact benchmark and a clean, decoupled pose refinement mechanism]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across 4 real benchmarks, with in-depth normal ablations, contact F1 scores, and OOD validation]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Rigorous mathematical formulation, clear qualitative illustrations, and complete metrics]
  • Value: ⭐⭐⭐⭐⭐ [Provides an essential physical ground truth standard and a strong baseline for robotic manipulation and 3D vision]