Geo-DPO: Aligning Semantic Intent with Geometry for 3D Affordance Segmentation¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Segmentation
Keywords: 3D affordance segmentation, point cloud, multimodal large language model, direct preference optimization, hierarchical geometry adapter
TL;DR¶
Addressing the severe geometric ambiguity and boundary overflow caused by the over-compression of the <SEG> token in 3D MLLM sequential affordance segmentation, Geo-DPO introduces a Hierarchical Geometry Adapter to inject multi-scale physical features into the semantic query and Contrastive Affordance Preference Optimization to penalize mask overflow via hybrid position and shape prototype rewards.
Background & Motivation¶
3D language-guided affordance segmentation aims to identify and localize actionable, functional regions of objects within 3D scenes based on natural language instructions. This capability is pivotal for embodied AI and robotic manipulation, as it directly translates high-level human intent into precise physical contact surfaces necessary for real-world interactions. Recently, 3D Multimodal Large Language Models (3D MLLMs) have demonstrated impressive common-sense reasoning on sequential affordance tasks, where an agent must deduce a logical series of functional parts to carry out complex, multi-step instructions (such as locating the handle of a watering can to lift it before localizing the spout to pour). To perform dense grounding, prevailing 3D MLLM frameworks append a special <SEG> token to the end of the generated text response, compressing the high-level semantic reasoning into a single vector that serves as a query for downstream mask decoders.
However, this paradigm suffers from a critical structural bottleneck: condensing rich, high-dimensional 3D spatial reasoning into a single 1D hidden state inevitably strips away fine-grained geometric structures and precise coordinate cues. This semantic over-compression induces severe "geometric ambiguity" in the output masks. Rather than isolating clean part boundaries, the predicted masks routinely bleed into physically distinct adjacent componentsโsuch as overflowing from a mug's handle onto its body. Existing frameworks trained under isolated point-wise objectives (such as Binary Cross Entropy and Dice loss) lack holistic structural awareness to penalize such boundary leakage, causing compounding geometric errors across multi-step execution.
To eliminate geometric ambiguity, the core insight is that semantic intent must be explicitly aligned with physical 3D geometries from both structural feature restoration and preference-based boundary regularization perspectives. Core idea: propose the Geo-DPO framework, which structurally restores fine-grained geometric awareness by injecting multi-scale 3D visual encoder features into the <SEG> token via a Hierarchical Geometry Adapter (HGA), and introduces Contrastive Affordance Preference Optimization (CAPO) with hybrid Soft-IoU position and prototype shape rewards to penalize mask overflow via contrastive ranking.
Method¶
Overall Architecture¶
The input to Geo-DPO comprises an object 3D point cloud \(X \in \mathbb{R}^{N \times 3}\) and a natural language instruction \(T\). The point cloud is first processed by a frozen pre-trained 3D visual encoder (Uni3D) to extract multi-scale dense geometric representations, which are projected into the embedding space of a 3D MLLM backbone (ShapeLLM-7B, adapted via LoRA). The MLLM performs sequential reasoning over the multimodal input and auto-regressively produces an output sequence ending with an initial semantic query \(h_{\text{raw}} \in \mathbb{R}^{1 \times D}\). To overcome semantic over-compression, the Hierarchical Geometry Adapter (HGA) queries low-level and intermediate feature maps from the 3D encoder in a cascaded fashion, refining \(h_{\text{raw}}\) into a geometry-grounded token \(h_{\text{geo}}\). This refined query interacts with high-resolution point features via dot-product decoding to yield the point-wise prediction mask \(P\). During training, the Contrastive Affordance Preference Optimization (CAPO) objective supervises the decoder alongside standard text generation and segmentation losses, regularizing boundary precision through in-batch preference ranking.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: 3D Point Cloud and Text Instruction"] --> B["3D Vision Encoder & MLLM Reasoning<br/>Auto-regressively generate h_raw"]
B --> C["Hierarchical Geometry Adapter HGA<br/>Cascaded cross-attention with multi-scale 3D features"]
C --> D["Geometry-Aware Query Decoding<br/>Dot-product generates point-wise mask P"]
D --> E["Contrastive Affordance Preference Optimization CAPO<br/>Hybrid position and shape rewards penalize overflow"]
E --> F["Output: Physically Aligned 3D Affordance Mask"]
Key Designs¶
1. Hierarchical Geometry Adapter: cascaded cross-attention for structural grounding
To resolve the loss of spatial coordinates within the compressed <SEG> token, the Hierarchical Geometry Adapter (HGA) extracts multi-scale representations directly from the intermediate stages of the frozen 3D visual encoder. Because early encoder layers excel at capturing sharp boundaries and local curvatures while intermediate layers retain part-level geometric context, HGA employs a cascaded cross-attention mechanism to progressively inject structural priors into the semantic query:
where \(F_{\text{low}}\) and \(F_{\text{mid}}\) denote shallow and intermediate 3D features linearly projected into the language hidden dimension \(D\). By attending first to fine-grained edges and subsequently to broader local part geometry, \(h_{\text{raw}}\) is transformed into an anchored physical token \(h_{\text{geo}}\). Finally, \(h_{\text{geo}}\) serves as the conditional query taking a dot product with point-wise dense features \(F_{\text{point}} \in \mathbb{R}^{N \times D}\) followed by a Sigmoid function \(\sigma(\cdot)\) to generate mask \(P = \sigma(h_{\text{geo}} F_{\text{point}}^\top)\).
2. Hybrid Preference Reward: disentangling spatial overlap and physical shape Because standard point-wise cross-entropy losses fail to penalize structural bleeding into adjacent functional parts, CAPO introduces a hybrid reward function \(R(P, G)\) comparing predicted mask \(P\) with ground truth \(G \in \{0, 1\}^N\). To assess macroscopic spatial overlap, the Position Reward \(R_{\text{pos}}\) is computed via Soft-IoU:
However, Soft-IoU alone is insensitive to geometric distortion when masks expand into geometrically distinct components. To counter this, a feature-level Shape Reward \(R_{\text{shape}}\) is formulated. Utilizing intermediate features \(F_{\text{mid}}\), mask-weighted pooling extracts the physical prototype vectors \(v_P, v_G \in \mathbb{R}^{C_2}\) for the predicted and ground-truth regions:
The shape reward is defined as their cosine similarity \(R_{\text{shape}}(P, G) = \frac{v_P^\top v_G}{\|v_P\|_2 \|v_G\|_2}\). When the predicted mask bleeds into adjacent surfaces with differing curvatures or structures, \(v_P\) diverges markedly from \(v_G\), drastically reducing the shape reward. The combined hybrid reward is \(R = R_{\text{pos}} + \alpha R_{\text{shape}}\) (where \(\alpha\) balances the two rewards).
3. Contrastive Preference Objective: in-batch negative sampling for boundary alignment Rather than directly applying complex reinforcement learning, CAPO frames mask refinement as a direct preference ranking problem. Using an In-Batch Negative Sampling strategy across batch size \(B\), for the \(i\)-th instruction-scene pair, the matched ground truth \(G_i\) serves as the positive preference target, while non-matching ground truths from other batch instances \(G_j (j \neq i)\) act as negative targets. The preference objective maximizes the reward margin:
where \(\beta\) is a temperature hyperparameter. By amplifying the reward margin between true physical affordances and negative candidate masks, this contrastive push-pull mechanism directly penalizes spatial leakage and enforces sharp boundary adherence.
Loss & Training¶
The entire Geo-DPO architecture is optimized end-to-end via a multi-task objective: $\(\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{txt}} + \mathcal{L}_{\text{seg}} + \lambda \mathcal{L}_{\text{CPO}}\)$ where \(\mathcal{L}_{\text{txt}}\) is the autoregressive cross-entropy loss for language reasoning, \(\mathcal{L}_{\text{seg}}\) is the composite point-wise segmentation loss combining Binary Cross Entropy (BCE) and Dice loss, and \(\mathcal{L}_{\text{CPO}}\) is the preference ranking objective scaled by \(\lambda = 0.2\). The model builds upon ShapeLLM-7B adapted via LoRA (rank 8), with the Uni3D vision backbone frozen. Training utilizes AdamW with an initial learning rate of \(2 \times 10^{-4}\), zero weight decay, a 3% warm-up ratio under a cosine annealing schedule, and FlashAttention acceleration.
Key Experimental Results¶
Main Results¶
On the comprehensive SeqAfford benchmark spanning single-step and sequential affordance tasks, Geo-DPO is evaluated against prior state-of-the-art methods across Seen, Unseen, and Sequential splits using mIoU (%), AUC (%), Similarity (SIM), and Mean Absolute Error (MAE).
| Task Setting | Method | mIoU (โ) | AUC (โ) | SIM (โ) | MAE (โ) |
|---|---|---|---|---|---|
| Seen | ReferTrans (NeurIPS 2021) | 11.4 | 77.2 | 0.449 | 0.135 |
| Seen | ReLA (CVPR 2023) | 12.1 | 76.3 | 0.480 | 0.130 |
| Seen | 3D-SPS (CVPR 2022) | 10.1 | 75.2 | 0.413 | 0.141 |
| Seen | IAGNet (ICCV 2023) | 14.2 | 81.7 | 0.510 | 0.117 |
| Seen | LASO (CVPR 2024) | 16.3 | 84.3 | 0.568 | 0.108 |
| Seen | SeqAfford (CVPR 2025) | 19.5 | 86.9 | 0.594 | 0.098 |
| Seen | Geo-DPO (Ours) | 20.9 | 88.1 | 0.610 | 0.091 |
| Unseen | ReferTrans | 9.1 | 67.4 | 0.427 | 0.151 |
| Unseen | ReLA | 9.3 | 68.2 | 0.423 | 0.147 |
| Unseen | 3D-SPS | 7.1 | 66.9 | 0.397 | 0.162 |
| Unseen | IAGNet | 11.7 | 73.6 | 0.438 | 0.143 |
| Unseen | LASO | 12.4 | 76.1 | 0.502 | 0.132 |
| Unseen | SeqAfford | 13.8 | 82.4 | 0.518 | 0.128 |
| Unseen | Geo-DPO (Ours) | 14.8 | 85.6 | 0.528 | 0.126 |
| Sequential | ReferTrans* | 10.8 | 74.5 | 0.425 | 0.142 |
| Sequential | ReLA* | 11.4 | 74.8 | 0.463 | 0.136 |
| Sequential | 3D-SPS* | 9.9 | 73.1 | 0.407 | 0.148 |
| Sequential | IAGNet* | 13.5 | 78.2 | 0.496 | 0.131 |
| Sequential | LASO* | 14.3 | 80.7 | 0.521 | 0.124 |
| Sequential | SeqAfford | 14.6 | 84.2 | 0.573 | 0.118 |
| Sequential | Geo-DPO (Ours) | 15.1 | 86.5 | 0.601 | 0.115 |
(Note: Baselines marked with * receive ground-truth sequential order to compensate for their inability to natively perform sequential affordance reasoning.)
Ablation Study¶
Ablations on the Sequential split systematically evaluate the core architectural and training contributions of Geo-DPO.
1. Core Component Contribution
| HGA Adapter | CAPO Objective | mIoU (โ) | AUC (โ) | SIM (โ) | MAE (โ) | Note |
|---|---|---|---|---|---|---|
| โ | โ | 14.5 | 84.0 | 0.572 | 0.117 | Baseline without geometric enhancement |
| โ | โ | 14.7 | 84.9 | 0.585 | 0.116 | Structural feature restoration alone |
| โ | โ | 14.8 | 85.4 | 0.591 | 0.117 | Preference boundary regularization alone |
| โ | โ | 15.1 | 86.5 | 0.601 | 0.115 | Full model combining both designs |
2. Multi-Scale Feature Fusion in HGA
| Fusion Strategy | mIoU (โ) | AUC (โ) | SIM (โ) | MAE (โ) | Note |
|---|---|---|---|---|---|
| Addition | 14.4 | 84.1 | 0.572 | 0.118 | Fails to exploit spatial correlation |
| Concatenation | 14.5 | 84.5 | 0.579 | 0.118 | Static channel aggregation |
| Cascaded Cross-Attention (Ours) | 15.1 | 86.5 | 0.601 | 0.115 | Adaptive context-aware structural querying |
3. Robustness Across Instruction Reasoning Steps
| Reasoning Steps | Baseline SeqAfford mIoU | Geo-DPO mIoU | Improvement |
|---|---|---|---|
| 1 Step (Single) | 19.5 | 20.9 | +1.4 |
| 2 Steps | 14.8 | 15.2 | +0.4 |
| \(\ge 3\) Steps | 14.5 | 15.1 | +0.6 |
Key Findings¶
- Complementary Dual Formulation: Introducing HGA alone (+0.2% mIoU) or CAPO alone (+0.3% mIoU) provides solid gains, but their combination achieves 15.1% mIoU, demonstrating that structural spatial feature recovery and reward-based boundary penalty reinforce each other.
- Error Accumulation Suppression: In multi-step sequential tasks (\(\ge 3\) steps), baseline SeqAfford drops by 0.3% mIoU due to compounding spatial confusion, whereas Geo-DPO exhibits only a 0.1% decline, highlighting that strict physical boundaries prevent error accumulation across long-horizon planning.
- Preference Scaling Balance: Sweeping \(\lambda\) reveals an optimum at 0.2; excessively high values (\(\lambda = 0.4\)) degrade mIoU to 14.8% by destabilizing base semantic token grounding.
Highlights & Insights¶
- Adapting DPO to 3D Physical Geometry: Rather than aligning with subjective human textual feedback, Geo-DPO cleverly treats the underlying 3D point cloud surfaces and geometric prototypes as objective preference ground truth, establishing an elegant unsupervised alignment paradigm.
- Feature Prototype Shape Reward: Using masked-pooling over intermediate 3D encoder features allows the model to detect subtle curvature discrepancies that point-wise IoU metrics miss, creating an effective critic against mask bleeding.
- Non-Invasive Architecture: The HGA module and CAPO loss integrate seamlessly without altering the core autoregressive language model pipeline, offering high transferability to general 3D vision-language grounding models.
Limitations & Future Work¶
- Autoregressive Inference Latency: Relying on a 7B MLLM incurs substantial inference latency, posing challenges for real-time high-frequency closed-loop robotic control. Compressing the model via knowledge distillation into a lightweight feedforward architecture is an important future direction.
- Restriction to Rigid Objects: CAPO currently evaluates static, rigid shapes; extending physical prototype matching to highly deformable items (e.g., cloth manipulation) or articulated mechanisms remains challenging.
- Generalization to Generic 3D Segmentation: The preference-based boundary regularization is directly applicable to broader 3D instance and semantic part segmentation benchmarks, which remains to be validated.
Related Work & Insights¶
- vs SeqAfford (CVPR 2025): SeqAfford introduced 3D MLLMs to sequential affordance grounding but suffered from severe mask overflow due to isolated point-wise supervision and
<SEG>token over-compression. Geo-DPO eliminates boundary confusion via HGA feature injection and CAPO preference ranking. - vs LASO (CVPR 2024) / 3D-SPS: Early grounding methods relied purely on shallow text-point cross-modal attention, lacking the high-level semantic reasoning necessary for multi-step affordance planning. Geo-DPO bridges high-level reasoning with low-level geometric fidelity.
Rating¶
- Novelty: โญโญโญโญโญ Formulates 3D affordance boundary alignment as direct preference optimization with prototype-guided shape rewards.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive benchmarks across Seen, Unseen, and Sequential splits with in-depth component and hyperparameter ablations.
- Writing Quality: โญโญโญโญโญ Crisp problem framing, elegant mathematical definitions, and tight alignment between architectural figures and technical prose.
- Value: โญโญโญโญโญ Provides a robust, practical alignment framework for bridging foundation model planning with physical robotic manipulation.