IP-SAM: Rethinking Prompt-Conditioned Segmentation for Prompt-Absent Deployment¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Medical Imaging
Keywords: Prompt-Space Conditioning, Prompt-Absent Deployment, SAM2 Adaptation, Camouflaged Object Detection, Polyp Segmentation
TL;DR¶
Addressing the structural interface mismatch where foundation segmenters lack external prompts during automatic deployment, IP-SAM proposes a prompt-space conditioning paradigm that synthesizes complementary intrinsic foreground/background prompts through a frozen prompt encoder and applies asymmetric prompt-space gating to suppress background false positives, achieving state-of-the-art performance with only 21.26M trainable parameters.
Background & Motivation¶
Prompt-conditioned vision foundation models, most notably SAM and SAM2, have established a dominant paradigm for interactive segmentation by steering mask decoding with explicit spatial cues such as points, bounding boxes, or rough masks. However, in many real-world deployment scenarios, segmentation pipelines must run fully autonomously without human guidance or upstream detector-generated bounding boxes, placing the model in a strictly "prompt-absent deployment" regime. This induces an architectural paradox: the mask decoder is pre-trained to depend heavily on prompt-conditioned signals for spatial grounding, yet during automatic inference, no external prompts exist to steer the decoding process.
Existing efforts to adapt promptable segmenters to automated tasks (such as camouflaged object detection and medical polyp segmentation) predominantly follow a "feature-space adaptation" route—inserting parameter-efficient adapters, lateral fusion layers, or specialized decoder heads directly into intermediate image feature streams. While feature-level tuning provides task alignment, it systematically bypasses the foundation model's native prompt encoder and prompt-conditioned decoding interface, leaving the pre-trained prompt manifold dormant and failing to harness its rich geometric and boundary priors. Conversely, attempts to automatically synthesize discrete geometric cues (e.g., predicting pseudo-points or bounding boxes) tend to fail catastrophically under severe visual camouflage and low tissue contrast, as ambiguous boundaries cause coordinate misregistration and inject heavy false positives into the decoder; furthermore, symmetric feature mixing often amplifies deceptive background textures.
To resolve this mismatch, this paper revisits downstream adaptation from a prompt-space perspective. Rather than bypassing the prompt interface, task-specific spatial requirements should be translated directly into native prompt conditions that the pre-trained decoder inherently understands. Core idea: develop a Self-Prompt Generator to synthesize complementary dense and sparse intrinsic foreground/background prompts, project continuous logits through SAM2's frozen prompt encoder to reactivate the native prompt manifold, and enforce asymmetric Prompt-Space Gating before decoding to suppress background false positives without external prompts.
Method¶
Overall Architecture¶
The core architecture of IP-SAM reframes automatic prompt-absent segmentation into an end-to-end pipeline: intrinsic prompt generation \(\to\) frozen manifold projection \(\to\) prompt-space asymmetric gating \(\to\) prompt-conditioned decoding. Given an input image \(I\), high-level visual semantic embeddings \(E \in \mathbb{R}^{C \times H \times W}\) (at \(1/16\) input resolution) are extracted using the SAM2 image encoder adapted with lightweight LoRA. A dual-branch Self-Prompt Generator (SPG) driven by learnable task queries distills \(E\) into complementary intrinsic foreground and background prompt representations. The continuous real-valued dense logits and sparse tokens are directly fed into SAM2's strictly frozen prompt encoder to produce native prompt embeddings. Prior to mask decoding, Prompt-Space Gating (PSG) leverages the negative background prompt embedding to asymmetrically suppress deceptive activations in the positive prompt embedding. Finally, a customized mask decoder generates fine-grained segmentation masks guided by the purified prompt conditions.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image I"] --> B["SAM2 Image Encoder (LoRA)<br/>Extract High-Level Semantic Embedding E"]
B --> C["Intrinsic Prompt Generation & Frozen Projection<br/>SPG synthesizes complementary prompts into frozen encoder"]
C --> D["Prompt-Space Asymmetric Gating (PSG)<br/>Construct suppressive gate from background embedding"]
D --> E["Prompt-Conditioned Decoding & Semantic Anchoring<br/>Drive mask decoder with purified prompts and propagated tokens"]
E --> F["Coarse Mask Mc & Residual Refined Mask Output"]
Key Designs¶
1. Intrinsic prompt generation and frozen manifold projection: translating task cues into native prompt manifold In camouflaged scenes and low-contrast lesions, heuristic pseudo-prompts (e.g., predicted points or boxes) are fragile and prone to spatial drift. To provide continuous, uncertainty-aware spatial guidance, IP-SAM introduces a dual-branch Self-Prompt Generator (SPG). SPG utilizes a 2-layer TwoWayTransformer decoder driven by learnable task queries \(Q \in \mathbb{R}^{N \times d}\) to cross-attend over the flattened image embedding \(E\). Instead of outputting discrete coordinates, SPG synthesizes complementary dense prompt logits \(P^+, P^- \in \mathbb{R}^{1 \times 128 \times 128}\), continuous sparse prompt tokens \(S^+, S^- \in \mathbb{R}^{K \times C}\), and specialized propagated output tokens \(T_{\text{prop}} \in \mathbb{R}^{K_m \times C}\) matching the dimensionality of SAM2 mask tokens.
Crucially, rather than thresholding dense predictions into binary masks, continuous real-valued logits \(P^\pm\) are routed directly into SAM2's strictly frozen prompt encoder \(\mathcal{E}\), yielding dense prompt embeddings \(Z^\pm = \mathcal{E}(P^\pm) \in \mathbb{R}^{C \times H_e \times W_e}\). Continuous sparse tokens \(S^\pm\) are combined with pre-trained point-type embeddings to form sparse conditions \(U^\pm\). Empirical analysis reveals that explicit binarization or temperature scaling degrades performance, whereas continuous logits preserve subtle spatial uncertainty while remaining fully compatible with the frozen prompt manifold. Keeping the prompt encoder frozen prevents task-specific noise from corrupting pre-trained geometric priors.
2. Prompt-space asymmetric gating: preemptively suppressing deceptive background false positives Under severe foreground-background ambiguity, standard symmetric feature fusion tends to amplify deceptive background textures. Even when the background prior \(P^-\) is embedded into prompt space as \(Z^-\), feeding positive and negative dense embeddings symmetrically into the decoder causes false-positive leakage. To address this, Prompt-Space Gating (PSG) designates the negative branch strictly as an asymmetric suppressive constraint prior to decoding.
PSG first transforms the negative dense embedding \(Z^-\) into an unactivated background feature \(B = \phi_{\text{feat}}(Z^-)\) and a localized gate decision map \(G\): $\(G = \sigma(\phi_{\text{gate}}(B)), \quad \widetilde{Z}^+ = Z^+ \odot (1 - G)\)$ Here, \((1 - G)\) selectively attenuates background-correlated activations within the positive prompt embedding \(Z^+\), producing a purified embedding \(\widetilde{Z}^+\). Next, a convolutional fusion block \(\psi\) maps the concatenated representation \([\widetilde{Z}^+, B]\) back to \(C\) channels, stabilized by an anchored residual connection: $\(Z = \psi([\widetilde{Z}^+, B]) + Z^+\)$ Crucially, the residual is anchored to the original unsuppressed \(Z^+\) rather than the gated feature \(\widetilde{Z}^+\). Because spatial gating can inadvertently attenuate fine foreground boundaries, anchoring to \(Z^+\) preserves positive structural topology, allowing \(\psi\) to focus on cancelling localized prompt-space noise.
3. Prompt-conditioned decoding and semantic anchoring: multi-condition synergy with multi-scale refinement Following prompt-space purification, the conditioned prompt embedding drives downstream mask decoding. To exploit SAM2's prompt-conditioned decoder interface, IP-SAM constructs a unified prompt condition: sparse tokens from both branches are concatenated into \([U^+; U^-]\), and SAM2's default mask tokens are replaced with the propagated semantic tokens \(T_{\text{prop}}\) from SPG, while retaining the native IoU token \(T_{\text{iou}}\). The customized mask decoder \(\mathcal{D}\) processes image features \(E\), purified dense prompt embedding \(Z\), concatenated sparse tokens, and semantic output tokens to predict a robust coarse mask \(M_c\): $\(M_c = \mathcal{D}(E, Z, [U^+; U^-], [T_{\text{iou}}; T_{\text{prop}}])\)$ Two optional lightweight modules provide further boundary enhancement: an auxiliary lateral inhibition module (Lat.) that uses the detached coarse mask \(\sigma(M_c)\) to soft-gate multi-scale FPN features, suppressing out-of-target distractors; and a pixel refinement head (Ref.) that fuses coarse masks, purified prompt embeddings, and image features through a bottleneck to predict a residual offset \(M = M_c + \mathcal{R}(\cdot)\) via zero-initialized convolution.
Loss & Training¶
The entire framework is trained end-to-end. SPG is supervised with complementary binary cross-entropy targets on foreground and background prompts: $\(\mathcal{L}_{\text{SPG}} = \text{BCE}(\sigma(P^+), Y) + \text{BCE}(\sigma(P^-), 1 - Y)\)$ where \(Y \in \{0, 1\}^{1 \times H_0 \times W_0}\) is the ground-truth mask. The mask decoding loss \(\mathcal{L}_{\text{mask}}\) is a linear combination of standard BCE, IoU loss, and L1 loss. The overall optimization objective is: $\(\mathcal{L} = \lambda_{\text{spg}}\mathcal{L}_{\text{SPG}} + \lambda_c \mathcal{L}_{\text{mask}}(M_c, Y) + \lambda_r \mathcal{L}_{\text{mask}}(M, Y)\)$ with weights set to \(\lambda_{\text{spg}} = \lambda_c = \lambda_r = 1.0\). The prompt encoder remains strictly frozen, while LoRA (rank \(r = 32\)) is applied to the QKV and output projection layers of the SAM2 image encoder. Trainable parameters across SPG, PSG, LoRA, and the mask decoder total only 21.26M. Training is conducted on a single RTX 3090 GPU with AdamW (learning rate \(1 \times 10^{-4}\), batch size 1) for 100 epochs.
Key Experimental Results¶
Main Results¶
Quantitative evaluations across four standard COD benchmarks (CAMO, CHAMELEON, COD10K, NC4K) under a deterministic prompt-absent protocol demonstrate consistent state-of-the-art accuracy:
| Dataset | Metric | IP-SAM (Ours) | ZoomNeXt (Specialist SOTA) | SAM2-Adapter (SAM2 Adaptation) | SAM2 (Zero-Prompt Baseline) |
|---|---|---|---|---|---|
| COD10K | MAE \(\downarrow\) / \(S_m \uparrow\) | 0.017 / 0.906 | 0.018 / 0.898 | 0.020 / 0.891 | 0.132 / 0.669 |
| COD10K | \(F_\beta^\omega \uparrow\) / \(E_\phi \uparrow\) | 0.838 / 0.947 | 0.834 / 0.939 | 0.814 / 0.936 | 0.287 / 0.531 |
| CAMO | MAE \(\downarrow\) / \(S_m \uparrow\) | 0.032 / 0.912 | 0.041 / 0.892 | 0.048 / 0.884 | 0.198 / 0.658 |
| CAMO | \(F_\beta^\omega \uparrow\) / \(E_\phi \uparrow\) | 0.878 / 0.946 | 0.857 / 0.935 | 0.817 / 0.915 | 0.348 / 0.506 |
| CHAMELEON | MAE \(\downarrow\) / \(S_m \uparrow\) | 0.018 / 0.920 | 0.020 / 0.903 | 0.025 / 0.909 | 0.182 / 0.660 |
| NC4K | MAE \(\downarrow\) / \(S_m \uparrow\) | 0.026 / 0.911 | 0.030 / 0.902 | 0.030 / 0.906 | 0.159 / 0.722 |
Cross-dataset medical polyp segmentation results (trained on Kvasir-SEG and evaluated on CVC-ClinicDB and ETIS without task-specific modifications):
| Dataset | Metric | IP-SAM (Ours) | LACFormer (Specialist SOTA) | SAM2-Adapter | MDSAM |
|---|---|---|---|---|---|
| Kvasir-SEG (In-domain) | mDice \(\uparrow\) / mIoU \(\uparrow\) | 0.913 / 0.861 | 0.906 / 0.851 | 0.899 / 0.846 | 0.906 / 0.851 |
| Kvasir-SEG (In-domain) | \(S_m \uparrow\) / MAE \(\downarrow\) | 0.929 / 0.025 | 0.922 / 0.029 | 0.918 / 0.031 | 0.922 / 0.028 |
| CVC-ClinicDB (Cross-domain) | mDice \(\uparrow\) / mIoU \(\uparrow\) | 0.840 / 0.778 | 0.826 / 0.758 | 0.801 / 0.738 | 0.796 / 0.724 |
| ETIS (Cross-domain) | mDice \(\uparrow\) / \(S_m \uparrow\) | 0.764 / 0.868 | 0.798 / 0.884 | 0.733 / 0.847 | 0.753 / 0.858 |
Ablation Study¶
Incremental component ablation on COD10K and CAMO benchmarks:
| Config | T-Param (M) | COD10K MAE \(\downarrow\) | COD10K \(S_m \uparrow\) | CAMO MAE \(\downarrow\) | CAMO \(S_m \uparrow\) | Note |
|---|---|---|---|---|---|---|
| Baseline (LoRA + Decoder + Null Prompt) | 9.15 | 0.024 | 0.857 | 0.043 | 0.859 | Severe degradation without prompts |
| + Self-Prompt Generator (SPG) | 17.00 | 0.022 | 0.872 | 0.041 | 0.875 | Manifold reactivated, errors shrink |
| + Prompt-Space Gating (PSG) | 18.98 | 0.019 | 0.893 | 0.034 | 0.902 | Major leap in CAMO via noise filtering |
| + Lateral Inhibition (Lat.) | 19.13 | 0.018 | 0.899 | 0.032 | 0.907 | Suppresses multi-scale edge clutter |
| + Pixel Refinement (Ref., full model) | 21.26 | 0.017 | 0.906 | 0.032 | 0.912 | Fine boundary alignment, best overall |
Comparison of prompt-space interaction operators and residual anchoring:
| Operator / Anchoring Design | T-Param (M) | COD10K MAE \(\downarrow\) | CAMO MAE \(\downarrow\) | Note |
|---|---|---|---|---|
| None (discard background \(Z^-\)) | 19.27 | 0.020 | 0.040 | Lacks background constraint, false alarms rise |
| Subtraction (\(Z^+ - Z^-\)) | 19.93 | 0.019 | 0.038 | Symmetric subtraction erodes weak targets |
| Concatenation (\([Z^+; Z^-]\)) | 20.37 | 0.019 | 0.038 | Symmetric mixing amplifies deceptive textures |
| Anchoring to gated \(\widetilde{Z}^+\) | 21.26 | 0.019 | 0.039 | Gating attenuates foreground, hurting residual |
| Asymmetric Gate (Anchor to \(Z^+\), Ours) | 21.26 | 0.017 | 0.032 | Filters background false alarms while preserving topology |
Key Findings¶
- Crucial value of prompt manifold reactivation: Directly evaluating SPG output as a segmentation head (Full-SPG Direct) yields an inferior COD10K MAE of 0.025; however, routing the exact same logits through SAM2's frozen prompt encoder and decoding in prompt space cuts MAE to 0.017. This confirms that gains stem from reactivating pre-trained prompt-conditioned geometric priors rather than raw capacity.
- Asymmetric negative constraint is indispensable: PSG delivers the largest stepwise error reduction, cutting CAMO MAE from 0.041 to 0.034. Feature energy visualizations confirm that PSG forms a localized "isolation ring" around concealed objects, cutting off false-positive leakage into surrounding clutter.
- Frozen prompt encoder preserves geometric stability: Unfreezing the prompt encoder degrades COD10K MAE from 0.017 to 0.019. Freezing the prompt manifold acts as a regularizing anchor, preventing extreme camouflage ambiguity from warping the pre-trained embedding space.
Highlights & Insights¶
- Paradigm shift (Feature-space bypass \(\to\) Prompt-space conditioning): Overturns the standard practice of bypassing foundation prompt interfaces with ad-hoc feature adapters, demonstrating that routing synthesized continuous logits back into the frozen prompt encoder is substantially more effective.
- Asymmetric suppressive negative prompting: Rather than symmetrically fusing foreground and background features, PSG reformulates negative prompt embeddings as an asymmetric gating mechanism, cleanly excising background distractions before decoding.
- Parameter-efficient cross-domain transferability: With only 21.26M trainable parameters, IP-SAM outperforms heavy specialist models on four COD benchmarks and transfers directly to medical endoscopic polyp segmentation without domain-specific hyperparameter retuning.
Limitations & Future Work¶
- Foreground-background binary scope: The framework relies on dual complementary foreground-background queries, making it directly suitable for binary tasks (COD, SOD, polyp segmentation); extending to COCO-style multi-class or instance-level segmentation requires redesigning prompt allocation.
- Sensitivity to foundation model capacity: In smaller SAM2-T and SAM2-B variants, intrinsic prompt separability declines, occasionally leading to over-suppression of delicate target regions by the background branch. Uncertainty-aware adaptive gating remains an open direction.
Related Work & Insights¶
- vs SAM-Adapter / SAM2-Adapter: These models insert lightweight adapter modules into transformer layers but leave the prompt encoder dormant; IP-SAM reactivates prompt-conditioned decoding via intrinsic prompts, yielding superior structural consistency.
- vs AutoSAM / RSPrompter: Explicit prompt generators predict discrete point/box coordinates that easily drift under camouflage; IP-SAM routes continuous dense logits through the frozen encoder, preserving spatial uncertainty and fine boundary transitions.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Elegant shift from feature-space hacking to prompt-space conditioning with asymmetric suppressive gating.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 4 COD and 3 medical polyp benchmarks, with rigorous controlled ablations and internal feature visualizations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear exposition of the prompt-absent interface mismatch, followed by well-structured methodology and self-consistent empirical evidence.
- Value: ⭐⭐⭐⭐⭐ Provides a highly practical, parameter-efficient blueprint for deploying interactive foundation segmenters in fully automated environments.