Skip to content

Geo-DPO: Aligning Semantic Intent with Geometry for 3D Affordance Segmentation

Conference: ECCV 2026
Paper: ECCV Official
Area: Segmentation
Keywords: 3D affordance segmentation, point cloud, multimodal large language model, direct preference optimization, hierarchical geometry adapter

TL;DR

Addressing the severe geometric ambiguity and boundary overflow caused by the over-compression of the <SEG> token in 3D MLLM sequential affordance segmentation, Geo-DPO introduces a Hierarchical Geometry Adapter to inject multi-scale physical features into the semantic query and Contrastive Affordance Preference Optimization to penalize mask overflow via hybrid position and shape prototype rewards.

Background & Motivation

3D language-guided affordance segmentation aims to identify and localize actionable, functional regions of objects within 3D scenes based on natural language instructions. This capability is pivotal for embodied AI and robotic manipulation, as it directly translates high-level human intent into precise physical contact surfaces necessary for real-world interactions. Recently, 3D Multimodal Large Language Models (3D MLLMs) have demonstrated impressive common-sense reasoning on sequential affordance tasks, where an agent must deduce a logical series of functional parts to carry out complex, multi-step instructions (such as locating the handle of a watering can to lift it before localizing the spout to pour). To perform dense grounding, prevailing 3D MLLM frameworks append a special <SEG> token to the end of the generated text response, compressing the high-level semantic reasoning into a single vector that serves as a query for downstream mask decoders.

However, this paradigm suffers from a critical structural bottleneck: condensing rich, high-dimensional 3D spatial reasoning into a single 1D hidden state inevitably strips away fine-grained geometric structures and precise coordinate cues. This semantic over-compression induces severe "geometric ambiguity" in the output masks. Rather than isolating clean part boundaries, the predicted masks routinely bleed into physically distinct adjacent componentsโ€”such as overflowing from a mug's handle onto its body. Existing frameworks trained under isolated point-wise objectives (such as Binary Cross Entropy and Dice loss) lack holistic structural awareness to penalize such boundary leakage, causing compounding geometric errors across multi-step execution.

To eliminate geometric ambiguity, the core insight is that semantic intent must be explicitly aligned with physical 3D geometries from both structural feature restoration and preference-based boundary regularization perspectives. Core idea: propose the Geo-DPO framework, which structurally restores fine-grained geometric awareness by injecting multi-scale 3D visual encoder features into the <SEG> token via a Hierarchical Geometry Adapter (HGA), and introduces Contrastive Affordance Preference Optimization (CAPO) with hybrid Soft-IoU position and prototype shape rewards to penalize mask overflow via contrastive ranking.

Method

Overall Architecture

The input to Geo-DPO comprises an object 3D point cloud \(X \in \mathbb{R}^{N \times 3}\) and a natural language instruction \(T\). The point cloud is first processed by a frozen pre-trained 3D visual encoder (Uni3D) to extract multi-scale dense geometric representations, which are projected into the embedding space of a 3D MLLM backbone (ShapeLLM-7B, adapted via LoRA). The MLLM performs sequential reasoning over the multimodal input and auto-regressively produces an output sequence ending with an initial semantic query \(h_{\text{raw}} \in \mathbb{R}^{1 \times D}\). To overcome semantic over-compression, the Hierarchical Geometry Adapter (HGA) queries low-level and intermediate feature maps from the 3D encoder in a cascaded fashion, refining \(h_{\text{raw}}\) into a geometry-grounded token \(h_{\text{geo}}\). This refined query interacts with high-resolution point features via dot-product decoding to yield the point-wise prediction mask \(P\). During training, the Contrastive Affordance Preference Optimization (CAPO) objective supervises the decoder alongside standard text generation and segmentation losses, regularizing boundary precision through in-batch preference ranking.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: 3D Point Cloud and Text Instruction"] --> B["3D Vision Encoder & MLLM Reasoning<br/>Auto-regressively generate h_raw"]
    B --> C["Hierarchical Geometry Adapter HGA<br/>Cascaded cross-attention with multi-scale 3D features"]
    C --> D["Geometry-Aware Query Decoding<br/>Dot-product generates point-wise mask P"]
    D --> E["Contrastive Affordance Preference Optimization CAPO<br/>Hybrid position and shape rewards penalize overflow"]
    E --> F["Output: Physically Aligned 3D Affordance Mask"]

Key Designs

1. Hierarchical Geometry Adapter: cascaded cross-attention for structural grounding To resolve the loss of spatial coordinates within the compressed <SEG> token, the Hierarchical Geometry Adapter (HGA) extracts multi-scale representations directly from the intermediate stages of the frozen 3D visual encoder. Because early encoder layers excel at capturing sharp boundaries and local curvatures while intermediate layers retain part-level geometric context, HGA employs a cascaded cross-attention mechanism to progressively inject structural priors into the semantic query:

\[\mathbf{h}_1 = \mathbf{h}_{\text{raw}} + \text{Attention}(\mathbf{Q}=\mathbf{h}_{\text{raw}}, \mathbf{K}=\mathbf{F}_{\text{low}}, \mathbf{V}=\mathbf{F}_{\text{low}})$$ $$\mathbf{h}_{\text{geo}} = \mathbf{h}_1 + \text{Attention}(\mathbf{Q}=\mathbf{h}_1, \mathbf{K}=\mathbf{F}_{\text{mid}}, \mathbf{V}=\mathbf{F}_{\text{mid}})\]

where \(F_{\text{low}}\) and \(F_{\text{mid}}\) denote shallow and intermediate 3D features linearly projected into the language hidden dimension \(D\). By attending first to fine-grained edges and subsequently to broader local part geometry, \(h_{\text{raw}}\) is transformed into an anchored physical token \(h_{\text{geo}}\). Finally, \(h_{\text{geo}}\) serves as the conditional query taking a dot product with point-wise dense features \(F_{\text{point}} \in \mathbb{R}^{N \times D}\) followed by a Sigmoid function \(\sigma(\cdot)\) to generate mask \(P = \sigma(h_{\text{geo}} F_{\text{point}}^\top)\).

2. Hybrid Preference Reward: disentangling spatial overlap and physical shape Because standard point-wise cross-entropy losses fail to penalize structural bleeding into adjacent functional parts, CAPO introduces a hybrid reward function \(R(P, G)\) comparing predicted mask \(P\) with ground truth \(G \in \{0, 1\}^N\). To assess macroscopic spatial overlap, the Position Reward \(R_{\text{pos}}\) is computed via Soft-IoU:

\[R_{\text{pos}}(P, G) = \frac{\sum_{i=1}^N P_i G_i}{\sum_{i=1}^N P_i + \sum_{i=1}^N G_i - \sum_{i=1}^N P_i G_i + \epsilon}\]

However, Soft-IoU alone is insensitive to geometric distortion when masks expand into geometrically distinct components. To counter this, a feature-level Shape Reward \(R_{\text{shape}}\) is formulated. Utilizing intermediate features \(F_{\text{mid}}\), mask-weighted pooling extracts the physical prototype vectors \(v_P, v_G \in \mathbb{R}^{C_2}\) for the predicted and ground-truth regions:

\[v_P = \frac{\sum_{i=1}^N P_i F_{\text{mid}, i}}{\sum_{i=1}^N P_i}, \quad v_G = \frac{\sum_{i=1}^N G_i F_{\text{mid}, i}}{\sum_{i=1}^N G_i}\]

The shape reward is defined as their cosine similarity \(R_{\text{shape}}(P, G) = \frac{v_P^\top v_G}{\|v_P\|_2 \|v_G\|_2}\). When the predicted mask bleeds into adjacent surfaces with differing curvatures or structures, \(v_P\) diverges markedly from \(v_G\), drastically reducing the shape reward. The combined hybrid reward is \(R = R_{\text{pos}} + \alpha R_{\text{shape}}\) (where \(\alpha\) balances the two rewards).

3. Contrastive Preference Objective: in-batch negative sampling for boundary alignment Rather than directly applying complex reinforcement learning, CAPO frames mask refinement as a direct preference ranking problem. Using an In-Batch Negative Sampling strategy across batch size \(B\), for the \(i\)-th instruction-scene pair, the matched ground truth \(G_i\) serves as the positive preference target, while non-matching ground truths from other batch instances \(G_j (j \neq i)\) act as negative targets. The preference objective maximizes the reward margin:

\[\mathcal{L}_{\text{CPO}} = - \frac{1}{B(B-1)} \sum_{i=1}^B \sum_{j \neq i} \log \sigma \left( \beta \left( R(P_i, G_i) - R(P_i, G_j) \right) \right)\]

where \(\beta\) is a temperature hyperparameter. By amplifying the reward margin between true physical affordances and negative candidate masks, this contrastive push-pull mechanism directly penalizes spatial leakage and enforces sharp boundary adherence.

Loss & Training

The entire Geo-DPO architecture is optimized end-to-end via a multi-task objective: $\(\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{txt}} + \mathcal{L}_{\text{seg}} + \lambda \mathcal{L}_{\text{CPO}}\)$ where \(\mathcal{L}_{\text{txt}}\) is the autoregressive cross-entropy loss for language reasoning, \(\mathcal{L}_{\text{seg}}\) is the composite point-wise segmentation loss combining Binary Cross Entropy (BCE) and Dice loss, and \(\mathcal{L}_{\text{CPO}}\) is the preference ranking objective scaled by \(\lambda = 0.2\). The model builds upon ShapeLLM-7B adapted via LoRA (rank 8), with the Uni3D vision backbone frozen. Training utilizes AdamW with an initial learning rate of \(2 \times 10^{-4}\), zero weight decay, a 3% warm-up ratio under a cosine annealing schedule, and FlashAttention acceleration.

Key Experimental Results

Main Results

On the comprehensive SeqAfford benchmark spanning single-step and sequential affordance tasks, Geo-DPO is evaluated against prior state-of-the-art methods across Seen, Unseen, and Sequential splits using mIoU (%), AUC (%), Similarity (SIM), and Mean Absolute Error (MAE).

Task Setting Method mIoU (โ†‘) AUC (โ†‘) SIM (โ†‘) MAE (โ†“)
Seen ReferTrans (NeurIPS 2021) 11.4 77.2 0.449 0.135
Seen ReLA (CVPR 2023) 12.1 76.3 0.480 0.130
Seen 3D-SPS (CVPR 2022) 10.1 75.2 0.413 0.141
Seen IAGNet (ICCV 2023) 14.2 81.7 0.510 0.117
Seen LASO (CVPR 2024) 16.3 84.3 0.568 0.108
Seen SeqAfford (CVPR 2025) 19.5 86.9 0.594 0.098
Seen Geo-DPO (Ours) 20.9 88.1 0.610 0.091
Unseen ReferTrans 9.1 67.4 0.427 0.151
Unseen ReLA 9.3 68.2 0.423 0.147
Unseen 3D-SPS 7.1 66.9 0.397 0.162
Unseen IAGNet 11.7 73.6 0.438 0.143
Unseen LASO 12.4 76.1 0.502 0.132
Unseen SeqAfford 13.8 82.4 0.518 0.128
Unseen Geo-DPO (Ours) 14.8 85.6 0.528 0.126
Sequential ReferTrans* 10.8 74.5 0.425 0.142
Sequential ReLA* 11.4 74.8 0.463 0.136
Sequential 3D-SPS* 9.9 73.1 0.407 0.148
Sequential IAGNet* 13.5 78.2 0.496 0.131
Sequential LASO* 14.3 80.7 0.521 0.124
Sequential SeqAfford 14.6 84.2 0.573 0.118
Sequential Geo-DPO (Ours) 15.1 86.5 0.601 0.115

(Note: Baselines marked with * receive ground-truth sequential order to compensate for their inability to natively perform sequential affordance reasoning.)

Ablation Study

Ablations on the Sequential split systematically evaluate the core architectural and training contributions of Geo-DPO.

1. Core Component Contribution

HGA Adapter CAPO Objective mIoU (โ†‘) AUC (โ†‘) SIM (โ†‘) MAE (โ†“) Note
โœ— โœ— 14.5 84.0 0.572 0.117 Baseline without geometric enhancement
โœ“ โœ— 14.7 84.9 0.585 0.116 Structural feature restoration alone
โœ— โœ“ 14.8 85.4 0.591 0.117 Preference boundary regularization alone
โœ“ โœ“ 15.1 86.5 0.601 0.115 Full model combining both designs

2. Multi-Scale Feature Fusion in HGA

Fusion Strategy mIoU (โ†‘) AUC (โ†‘) SIM (โ†‘) MAE (โ†“) Note
Addition 14.4 84.1 0.572 0.118 Fails to exploit spatial correlation
Concatenation 14.5 84.5 0.579 0.118 Static channel aggregation
Cascaded Cross-Attention (Ours) 15.1 86.5 0.601 0.115 Adaptive context-aware structural querying

3. Robustness Across Instruction Reasoning Steps

Reasoning Steps Baseline SeqAfford mIoU Geo-DPO mIoU Improvement
1 Step (Single) 19.5 20.9 +1.4
2 Steps 14.8 15.2 +0.4
\(\ge 3\) Steps 14.5 15.1 +0.6

Key Findings

  • Complementary Dual Formulation: Introducing HGA alone (+0.2% mIoU) or CAPO alone (+0.3% mIoU) provides solid gains, but their combination achieves 15.1% mIoU, demonstrating that structural spatial feature recovery and reward-based boundary penalty reinforce each other.
  • Error Accumulation Suppression: In multi-step sequential tasks (\(\ge 3\) steps), baseline SeqAfford drops by 0.3% mIoU due to compounding spatial confusion, whereas Geo-DPO exhibits only a 0.1% decline, highlighting that strict physical boundaries prevent error accumulation across long-horizon planning.
  • Preference Scaling Balance: Sweeping \(\lambda\) reveals an optimum at 0.2; excessively high values (\(\lambda = 0.4\)) degrade mIoU to 14.8% by destabilizing base semantic token grounding.

Highlights & Insights

  • Adapting DPO to 3D Physical Geometry: Rather than aligning with subjective human textual feedback, Geo-DPO cleverly treats the underlying 3D point cloud surfaces and geometric prototypes as objective preference ground truth, establishing an elegant unsupervised alignment paradigm.
  • Feature Prototype Shape Reward: Using masked-pooling over intermediate 3D encoder features allows the model to detect subtle curvature discrepancies that point-wise IoU metrics miss, creating an effective critic against mask bleeding.
  • Non-Invasive Architecture: The HGA module and CAPO loss integrate seamlessly without altering the core autoregressive language model pipeline, offering high transferability to general 3D vision-language grounding models.

Limitations & Future Work

  • Autoregressive Inference Latency: Relying on a 7B MLLM incurs substantial inference latency, posing challenges for real-time high-frequency closed-loop robotic control. Compressing the model via knowledge distillation into a lightweight feedforward architecture is an important future direction.
  • Restriction to Rigid Objects: CAPO currently evaluates static, rigid shapes; extending physical prototype matching to highly deformable items (e.g., cloth manipulation) or articulated mechanisms remains challenging.
  • Generalization to Generic 3D Segmentation: The preference-based boundary regularization is directly applicable to broader 3D instance and semantic part segmentation benchmarks, which remains to be validated.
  • vs SeqAfford (CVPR 2025): SeqAfford introduced 3D MLLMs to sequential affordance grounding but suffered from severe mask overflow due to isolated point-wise supervision and <SEG> token over-compression. Geo-DPO eliminates boundary confusion via HGA feature injection and CAPO preference ranking.
  • vs LASO (CVPR 2024) / 3D-SPS: Early grounding methods relied purely on shallow text-point cross-modal attention, lacking the high-level semantic reasoning necessary for multi-step affordance planning. Geo-DPO bridges high-level reasoning with low-level geometric fidelity.

Rating

  • Novelty: โญโญโญโญโญ Formulates 3D affordance boundary alignment as direct preference optimization with prototype-guided shape rewards.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive benchmarks across Seen, Unseen, and Sequential splits with in-depth component and hyperparameter ablations.
  • Writing Quality: โญโญโญโญโญ Crisp problem framing, elegant mathematical definitions, and tight alignment between architectural figures and technical prose.
  • Value: โญโญโญโญโญ Provides a robust, practical alignment framework for bridging foundation model planning with physical robotic manipulation.