Skip to content

Iterative Refinement of Semantic and Spatial Representations for Open-Vocabulary Camouflaged Object Segmentation

Conference: ECCV 2026
Paper: ECCV 2026
Area: Segmentation
Keywords: Open-Vocabulary Camouflaged Object Segmentation / Segment Anything Model 2 / Iterative Refinement / Category Re-ranker / Segmentation Modulator

TL;DR

Addressing error accumulation from unidirectional one-shot semantic injection in open-vocabulary camouflaged object segmentation, ISSR establishes a closed-loop bidirectional interaction between a spatial-structure-aware category re-ranker and a semantic-context modulator, iteratively refining masks and category representations to boost cSm by 14.4% on OVCamo.

Background & Motivation

Camouflaged Object Segmentation (COS) aims to detect and accurately segment objects that blend seamlessly into their surroundings. While conventional COS methods achieve remarkable precision, they strictly rely on the closed-set assumption with predefined categories and fixed distributions, causing severe performance collapse in open-world environments where novel classes continuously appear. To overcome this limitation, Open-Vocabulary Camouflaged Object Segmentation (OVCOS) has emerged, seeking to harness the open-world generalizability of large-scale vision-language models such as CLIP to localize and segment unseen camouflaged entities without task-specific retraining.

However, existing OVCOS frameworks typically adhere to a unidirectional two-stage paradigm with one-shot semantic injection. Specifically, CLIP is queried to predict category semantics, which are subsequently treated as fixed static priors to guide downstream mask generation. This pipeline implicitly presumes that category classification and spatial delineation can be cleanly decoupled in a single forward pass, and that semantic assignments require no retrospective correction. In camouflaged scenarios, foreground targets share high visual and textural similarity with background clutter, causing category semantics and spatial boundaries to be intimately coupled. A minor semantic misclassification at the initial stage is magnified during mask decoding, resulting in catastrophic segmentation collapse. More critically, prior architectures lack any feedback channel to exploit the emerging spatial structures to revise earlier semantic hypotheses, trapping the system in an irreversible error cascade.

Empirical observations reveal that while CLIP's Top-1 prediction frequently fluctuates due to camouflage interference, the ground-truth category consistently resides with high probability within the Top-N candidate pool. Hence, semantic failure stems not from an absence of learned concepts, but from brittle one-shot ranking without adequate spatial constraints. Concurrently, the progressive spatial masks produced by modern foundation models like SAM 2 offer salient structural cues for foreground-background separation that can effectively filter background distractors. Core idea: formulate category semantics and spatial masks as dynamically evolving state variables, constructing a spatial-structure-aware category re-ranker and a semantic-allocation segmentation modulator that interact in a closed-loop alternating optimization process, enabling semantics-guided spatial segmentation and mask-driven semantic calibration to iteratively self-correct and mutually reinforce.

Method

Overall Architecture

ISSR (Iterative Semantic and Spatial Refinement) freezes the pre-trained CLIP image/text encoders and the SAM 2 image encoder, focusing parameter updates on lightweight interaction modules between them. The framework comprises two core modules: the Spatial-Structure-Aware Category Re-ranker and the Semantic-Allocation Segmentation Modulator, supported by a shared Continuous Learnable Prompt mechanism.

During inference, ISSR executes a multi-round alternating closed-loop schedule. At iteration \(t\), the current semantic state modulates the multi-scale visual features within SAM 2 to decode an updated spatial mask. Next, this spatial mask acts as an explicit geometric prior that focuses the category re-ranker onto foreground regions, re-evaluating and re-ranking the Top-N candidate categories using intermediate multi-layer CLIP visual features. The updated category semantic representation is then passed into iteration \(t+1\).

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image & Candidate Texts"] --> B["Continuous Learnable Prompt<br/>Shared context tokens with class embeddings"]
    B --> C["Spatial-Structure-Aware Category Re-ranker<br/>Rerank Top-N classes via mask prior & multi-layer features"]
    C --> D["Semantic-Allocation Segmentation Modulator<br/>Adaptive balance of fine-grained interaction & global priors"]
    D --> E["SAM 2 Mask Decoder<br/>Generate refined foreground spatial mask"]
    E -->|Spatial mask feedback refines classification| C
    E -->|Reach iteration budget t=3| F["Output final class prediction & high-quality mask"]

Key Designs

1. Spatial-Structure-Aware Category Re-ranker: calibrating Top-N category confidence via mask-guided multi-layer features To mitigate the vulnerability of CLIP's global image embedding to pervasive background camouflage, the re-ranker avoids relying on raw zero-shot similarities and instead re-estimates confidence over a localized Top-N candidate set under explicit spatial constraints. First, normalized image features \(f\) extracted by the frozen CLIP visual backbone are matched against the text matrix \(W\) to retrieve the Top-N candidates \(\{(c_i, s_i)\}_{i=1}^N\). For each candidate \(c_i\), a compact 3D ranking descriptor is constructed: $\(r_i = \left[ s_i, \frac{i}{N}, s_i - s_1 \right]^\top\)$ which simultaneously encodes the candidate's absolute score, relative rank position, and margin against the current Top-1 choice. To recover mid-level structural cues lost during global pooling, intermediate visual tokens from layers 12, 16, and 18 of the CLIP encoder are pooled into compact representations \(\{v^{(m)}\}\). Crucially, the spatial mask \(M\) generated by SAM 2 in the current iteration serves as an explicit spatial filter, masking out irrelevant background texture and confining feature aggregation to the foreground. The re-ranker is optimized using a margin-based pairwise ranking loss over the candidate set: $\(\mathcal{L}_{cls} = \frac{1}{|\mathcal{P}|} \sum_{(p, n) \in \mathcal{P}} \max \left(0, \gamma - (\hat{s}_p - \hat{s}_n) \right)\)$ Restricting the loss to Top-N pairs eliminates disruptive gradients from millions of irrelevant open-vocabulary negatives and stabilizes ranking trajectory.

2. Semantic-Allocation Segmentation Modulator: dynamically balancing cross-modal interaction and category-level priors In modulating segmentation backbones with linguistic concepts, fine-grained cross-modal attention is vital for pinpointing locally consistent patterns, whereas category-level priors are needed to suppress false alarms across ambiguous backgrounds. Because their optimal trade-off shifts across varying camouflage intensities, the modulator incorporates an adaptive allocation mechanism between the SAM 2 encoder and mask decoder. Modality-specific gating networks first project SAM 2 features \(F_{sam}\) and CLIP text features \(F_{clip}\) into an aligned subspace as \(V_i\) and \(S_t\). An element-wise additive feature \(F_{fus}\) serves as Query, while \(V_i\) serves as Key and Value in a multi-head cross-attention layer to produce a spatially detailed representation \(F_{att}\). Concurrently, a cross-modal correlation map \(R = V_i \odot S_t\) is computed to filter attention tokens: \(F_{rel} = F_{att} \odot R\). The re-ranked Top-1 text vector is flattened into a global conditioning vector \(c_{top}\) and applied as a categorical prior: \(F_{cat} = F_{rel} \odot \text{Flatten}(\text{CLIP}(\text{Top-1}))\). A learnable scalar \(s\) is mapped through a sigmoid function \(\lambda = \sigma(s)\) to dynamically determine the convex blending weights \(\alpha = \lambda\) and \(\beta = 1 - \lambda\): $\(F_{mod} = \alpha \cdot F_{att} + \beta \cdot F_{cat}\)$ This data-driven weighting dynamically balances localized semantic-spatial correspondence with global categorical supervision.

3. Continuous Learnable Prompt: replacing discrete handcrafted templates to eliminate manual bias Conventional handcrafted prompt engineering (e.g., "a photo of a [CLS]") introduces subjective stylistic biases and struggles to characterize the obscure appearance variations typical of camouflaged objects. ISSR replaces discrete textual templates with a continuous learnable prompt structure. For each base class \(c \in \mathcal{C}_B\), the textual input sequence is parameterized as: $\(P_c = [p_1, p_2, \dots, p_l, cls_c]\)$ where \(\{p_i\}_{i=1}^l\) denotes a sequence of \(l\) learnable context vectors initialized randomly with identical dimensionality to CLIP text embeddings. These context vectors are shared across all categories and trained via back-propagation. When transitioning to open-vocabulary inference on novel classes, the learned context tokens are preserved while the class-name embedding is swapped with \(cls_{novel}\), enabling lightweight and robust adaptation without altering foundational language representations.

4. Semantic-Spatial Bidirectional Iteration: closed-loop alternating optimization for recursive self-correction To fundamentally circumvent unidirectional error accumulation, inference is formulated as an alternating state-space dynamical system. Let \(s^{(t)}\) denote the semantic representation and \(M^{(t)}\) denote the spatial mask at step \(t\). The semantic state is initialized as a confidence-weighted mixture of Top-N text embeddings: $\(s^{(0)} = \sum_{i=1}^N p_i \, e(c_i), \quad p_i = \text{softmax}(\text{CLIP}(I, c_i))\)$ The states are then updated recursively through coupled closed-loop equations: $\(\begin{cases} s^{(t+1)} = \mathcal{R}\left(\{e(c_i)\}_{i=1}^N, f_I, M^{(t)}\right) \\ M^{(t+1)} = \mathcal{S}\left(I, s^{(t+1)}\right) \end{cases}\)$ At step \(t\), the coarse mask \(M^{(t)}\) acts as a spatial attention filter for the re-ranker \(\mathcal{R}\), isolating the object territory and rectifying category confusion. The newly validated semantic state \(s^{(t+1)}\) is subsequently fed to the spatial operator \(\mathcal{S}\) inside SAM 2, synthesizing an updated mask \(M^{(t+1)}\) with sharpened boundaries and reduced background artifacts. This reciprocal guidance progressively steers both representations toward optimal alignment.

Loss & Training

The framework is optimized end-to-end under a unified multi-task objective: $\(\mathcal{L}_{total} = \mathcal{L}_{BCE}^\omega + \mathcal{L}_{IoU}^\omega + \mathcal{L}_{dice} + \mathcal{L}_{cls}\)$ The mask generation loss combines weighted binary cross-entropy (\(\mathcal{L}_{BCE}^\omega\)), weighted intersection-over-union (\(\mathcal{L}_{IoU}^\omega\)), and Dice loss (\(\mathcal{L}_{dice}\)), enforcing multi-scale pixel, regional, and structural alignment on camouflaged boundaries. The category re-ranker is supervised via margin-based pairwise ranking loss (\(\mathcal{L}_{cls}\)). During training, ground-truth masks are fed into the re-ranker to guarantee stable gradients, whereas at test time, the system runs purely on masks autonomously predicted by SAM 2. The model is trained with the AdamW optimizer (initial learning rate \(1 \times 10^{-6}\), weight decay \(5 \times 10^{-4}\)) on an NVIDIA RTX 3090 GPU with batch size 4 and image resolution \(384 \times 384\).

Key Experimental Results

Main Results

On the benchmark OVCamo dataset, ISSR is benchmarked against leading open-vocabulary semantic segmentation (OVSS) models and dedicated OVCOS architectures (OVCoser, SuCLIP). Evaluation metrics include structure measure (\(cSm\)), weighted F-measure (\(cF_\beta^\omega\)), mean absolute error (\(cMAE\)), mean F-measure (\(cF_\beta\)), enhanced-alignment measure (\(cEm\)), and intersection-over-union (\(cIoU\)).

Model Visual Backbone / Prompt Setup cSm ↑ cFβw ↑ cMAE ↓ cFβ ↑ cEm ↑ cIoU ↑
SimSeg CLIP-ViT-B/16 / CoOp 0.053 0.049 0.921 0.056 0.098 0.047
OVSeg CLIP-ViT-L/14 / Handcrafted 0.024 0.046 0.954 0.056 0.130 0.046
ODISE CLIP-ViT-L/14 / Diffusion-based 0.187 0.119 0.700 0.211 0.298 0.167
SAN CLIP-ViT-L/14 / Side Adapter 0.275 0.202 0.612 0.220 0.318 0.189
CAT-Seg CLIP-ViT-L/14 / Cost Aggregation 0.181 0.106 0.719 0.123 0.196 0.094
FC-CLIP CLIP-ConvNeXt-L / Conv-frozen 0.080 0.076 0.872 0.090 0.191 0.072
OVCoser CLIP-ConvNeXt-L / CamoPrompts 0.579 0.490 0.336 0.520 0.616 0.443
SuCLIP CLIP-ConvNeXt-L / Context-Aware 0.667 0.594 0.242 0.633 0.722 0.540
ISSR (Ours) CLIP-ViT-L/14 / Learnable Prompt 0.723 0.663 0.211 0.693 0.762 0.611

Beyond open-vocabulary evaluation, ISSR is tested on standard closed-set COS benchmarks (CAMO, COD10K, NC4K) against specialized state-of-the-art models:

Method CAMO (Sm ↑ / Fβw ↑ / MAE ↓) COD10K (Sm ↑ / Fβw ↑ / MAE ↓) NC4K (Sm ↑ / Fβw ↑ / MAE ↓)
SINet 0.745 / 0.712 / 0.092 0.776 / 0.667 / 0.043 0.808 / 0.768 / 0.058
ACUMEN 0.886 / 0.850 / 0.039 0.852 / 0.761 / 0.026 0.874 / 0.826 / 0.036
VSCode 0.873 / 0.820 / 0.046 0.869 / 0.780 / 0.025 0.882 / 0.841 / 0.032
ZoomNeXt 0.874 / 0.839 / 0.047 0.887 / 0.818 / 0.019 0.892 / 0.852 / 0.030
SAM2-UNet 0.884 / 0.861 / 0.042 0.880 / 0.789 / 0.021 0.901 / 0.863 / 0.029
CamoDiffusion 0.878 / 0.853 / 0.042 0.881 / 0.814 / 0.020 0.893 / 0.859 / 0.029
CFF-KDNet-P2 0.870 / 0.833 / 0.048 0.889 / 0.821 / 0.021 0.897 / 0.854 / 0.030
ISSR (Ours) 0.896 / 0.870 / 0.039 0.898 / 0.832 / 0.019 0.905 / 0.883 / 0.028

Ablation Study

The contribution of each individual module is systematically dissected by progressively integrating them into the frozen baseline:

No. Module Configuration cSm ↑ cFβw ↑ cMAE ↓ cFβ ↑ cEm ↑ cIoU ↑ Note
1 Baseline (CLIP + SAM 2) 0.658 0.615 0.268 0.620 0.711 0.534 Unidirectional non-iterative baseline
2 + Re-ranker (R) 0.678 0.634 0.256 0.645 0.730 0.557 Corrects Top-1 semantic errors
3 + Modulator (M) 0.697 0.642 0.242 0.660 0.743 0.576 Enhances cross-modal spatial alignment
4 + Learnable Prompt (L) 0.711 0.651 0.230 0.679 0.749 0.599 Eliminates handcrafted template bias
5 + Iterative Refinement (I) (Full Model) 0.723 0.663 0.211 0.693 0.762 0.611 Multi-round reciprocal refinement

Key ablation insights regarding hyperparameters and inputs include: 1. Iteration Budget \(t\): Varying \(t \in \{1, 2, 3, 4, 5\}\) shows steady gains up to \(t=3\) (\(cSm\) improves from 0.714 to 0.723, \(cIoU\) rises from 0.597 to 0.611). Beyond \(t=3\), slight performance degradation occurs (\(t=5\) yields \(cSm=0.711\), \(cIoU=0.586\)), indicating that while multi-round guidance cleanses errors, excessive cycling invites drift and erodes semantic diversity. 2. Candidate Pool Size Top-N: Scaling \(N \in \{2, 4, 6, 8, 10, 12, 14\}\) yields peak performance at \(N=10\) (\(cSm=0.723\)). Setting \(N < 10\) risks omitting positive classes, while \(N > 10\) injects distractor negatives that dilute ranking sharpness. 3. Prompt Context Length \(l\): Context lengths \(l \in \{4, 8, 16, 32, 64\}\) demonstrate that \(l=16\) strikes the best balance (\(cSm=0.723\)). Shorter lengths provide insufficient capacity, whereas longer contexts complicate optimization. 4. Re-ranker Input Signals: Inputting only ranking scores achieves \(cSm=0.711\); adding visual features \(I\) brings it to 0.715; adding mask features \(M\) achieves 0.717; fusing scores, visual features, and spatial masks \((S, I, M)\) achieves the top score of 0.723.

Key Findings

  • Bidirectional interaction breaks unidirectional performance ceilings: Compared with the unidirectional baseline (No. 1 vs. No. 5), the closed-loop iterative architecture achieves a net gain of +6.5% in \(cSm\) and +7.7% in \(cIoU\), eliminating background artifacts on camouflaged chameleons, cheetahs, and insects.
  • Spatial masks are essential for semantic re-ranking: Without mask-guided spatial cropping, multi-scale features are overwhelmed by background camouflage patterns. Supplying the mask provides an effective geometric focus for linguistic comparison.
  • Superior dual-task capability: Leveraging the strong pre-trained geometry of SAM 2 alongside targeted modulation, ISSR not only sets a new state-of-the-art in OVCOS, but also outperforms dedicated supervised models on closed-set benchmarks (e.g., \(F_\beta^\omega\) reaching 0.883 on NC4K).

Highlights & Insights

  • Shifting from static priors to dynamic state evolution: Rather than freezing semantic guidance after a single pass, ISSR models class labels and spatial masks as coupled, recursively refined state variables, conferring intrinsic self-correction capability.
  • Parsimonious 3D ranking descriptor: Rather than deploying cumbersome transformer cross-layers for re-ranking, the re-ranker utilizes a lightweight 3-dimensional ranking vector alongside pooled features to calibrate confidence at negligible computational cost.
  • Broad cross-task transferability: The principle of leveraging coarse segmentation masks to filter visual noise for categorical verification, and recursively updating masks with calibrated semantics, is directly applicable to weakly-supervised object localization, few-shot segmentation, and infrared small-target detection.

Limitations & Future Work

  • Constraint to single-category scenarios: ISSR is predominantly tailored for scenes with a single camouflaged foreground category. It lacks explicit instance separation and multi-label reasoning when handling multi-category camouflage compositions within a single frame.
  • Multi-round computational latency: An iteration budget of \(t=3\) requires 3 serial passes through the SAM 2 decoder and re-ranker. While foundational backbones remain frozen, inference wall-clock time is inherently longer than single-stage feedforward models.
  • Future directions: Developing an adaptive early-exit mechanism based on inter-iteration mask IoU convergence, and extending the framework to open-vocabulary camouflaged panoptic segmentation and video camouflage tracking.
  • vs OVCoser: OVCoser pioneered OVCOS using static one-shot category injection into a convolutional branch without feedback; ISSR integrates Top-N re-ranking and three-round closed-loop iteration, lifting \(cSm\) from 0.579 to 0.723 (+14.4%).
  • vs SuCLIP: SuCLIP introduced context-aware prompt tuning to reduce textual ambiguity but retains a strictly feedforward structure; ISSR adds mask-to-semantics feedback, achieving a +7.1% gain in \(cIoU\).
  • vs SAM2-UNet: SAM2-UNet excels at closed-set segmentation via SAM 2 but lacks open-vocabulary and vision-language interaction; ISSR equips SAM 2 with precise linguistic modulation while preserving its sharp boundary delineation.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Elegantly formulates semantic-spatial interaction as a closed-loop iterative refinement process, addressing the core vulnerability of open-vocabulary camouflage segmentation.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous validation spanning complete OVCOS metrics, three traditional closed-set benchmarks, and comprehensive hyperparameter ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Well-structured narrative with precise mathematical formulations and intuitive visual schematics.
  • Value: ⭐⭐⭐⭐☆ Offers valuable methodology for bidirectional cross-modal alignment under severe visual camouflage and occlusion.