Skip to content

Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement

Conference: NeurIPS2026
arXiv: 2609.34528
Code: https://github.com/blue-531/pref-ovss
Area: Segmentation
Keywords: open-vocabulary semantic segmentation, prompt disagreement, preference learning, region-localized optimization, domain adaptation

TL;DR

The paper converts localized segmentation disagreement between prompt templates into binary supervision and adapts open-vocabulary segmentation with region-localized preference optimization and outside-region consistency; with a ground-truth-based preference oracle, CAT-Seg-L improves its mean MESS mIoU from 35.26 to 45.88.

Background & Motivation

Open-vocabulary semantic segmentation (OVSS) lets users specify class names at inference time and grounds them at the pixel level through the alignment of a vision-language model (VLM). Models such as SAN and CAT-Seg already have useful zero-shot capabilities on everyday scenes, but chest radiographs, agriculture, remote sensing, and industrial inspection can differ substantially in appearance, class granularity, and terminology. Drawing dense masks for every new domain is expensive and requires expertise, while source-domain imageโ€“text alignment alone may not reliably distinguish the intended target concepts.

Preference learning offers another supervision interface: rather than drawing correct boundaries, annotators choose which of two predictions better matches the target. However, a whole-image comparison can involve one candidate being better on the left and the other on the right, leaving the feedback spatially ambiguous. Candidate generation is another concern: nearly identical predictions provide little information, while variations that change the input content can make the comparison semantically inconsistent.

The authors exploit the native prompting interface of OVSS. Keeping the image and vocabulary fixed while varying sentence structure and scale modifiers produces systematically different segmentation hypotheses. Selecting a locally uncertain region and then the most disagreeing pair within it makes the comparison concrete. Core Idea: use prompt disagreement to generate candidates and locate queries, turn one regional preference into class-balanced local optimization, and constrain unqueried regions with outside-region pseudo-label consistency.

Method

Overall Architecture

The inputs are sequential target-domain training images, a target vocabulary, and 14 fixed prompt templates; the output is an adapted OVSS model. Each image produces one regional query, one binary preference, and one gradient update, without replaying images or collecting preferences for multi-epoch training. The three core components are preference query mining, Region-Localized Preference Optimization (RLPO), and outside-region consistency.

The original OVSS parameters remain frozen, and only a vision LoRA and a residual text adapter are updated. A frozen reference model preserves the pre-adaptation state. Candidates come from the current model, while the reference anchors scores in the preference objective. Main experiments obtain preferences from a ground-truth oracle; human regional comparisons are the proposed practical interface, not an already completed large-scale human-driven adaptation experiment.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Target image and vocabulary<br/>14 template candidates"] --> B["Preference query mining"]
    B -->|Regional binary supervision| C["Region-localized<br/>preference optimization"]
    O["Training preference<br/>GT oracle or human"] -.-> C
    F["Frozen reference model"] -.-> C
    C --> D["Outside-region consistency"]
    D --> E["Single adapter update"]
    E -->|Inference with default templates| G["Adapted segmentation"]

The diagram's preferences and reference model provide training supervision and score anchoring, not additional inference inputs. After adaptation, prediction uses the baseline's default configuration: SAN uses its default ViLD prompt pool, and CAT-Seg uses โ€œa photo of a [CLASS] in the scene.โ€ The method does not request preferences for every test image or select templates using test ground truth.

Key Designs

1. Preference query mining: determine where to ask before choosing which hypotheses to compare

Each class name is inserted into 14 templates, producing 14 pixelwise class-probability maps. Templates vary both sentence structure and modifiers such as small, medium, and large, while preserving the target vocabulary. Unlike candidates generated through image perturbation, these are predictions for the same image and semantic target. This makes preferences comparable without training an additional candidate generator.

Query localization uses prompt-ensemble entropy: average the 14 class distributions at each pixel, then compute the Shannon entropy of that average. Binarize the entropy at its image-level 0.95 quantile and use the bounding box of the largest connected component as the query region \(R\). Thus, \(R\) is a rectangle rather than the original high-entropy mask and can contain low-entropy pixels; the protocol does not necessarily query exactly 5% of the image area.

After localization, convert each probability map into argmax labels, count label disagreements for every template pair inside \(R\), and select the pair with the largest count. Entropy identifies an uncertain location, whereas pairwise disagreement identifies the hypotheses worth comparing; these are different quantities. Ensemble entropy can also increase when all templates are individually uncertain, so it is not strictly a measure of only between-template disagreement.

The oracle or annotator judges which prediction better matches the intended segmentation inside the region, defining a winner template \(w\) and loser template \(l\). Local queries avoid competing regional advantages in whole-image preferences, but supervision remains limited to the available hypotheses. If every template misidentifies the target, a preference can only select the less incorrect result rather than supply a missing class or correct boundary.

2. Region-localized preference optimization: compare regional confidence changes relative to a reference

A language model's full-response probability cannot be transferred directly to a segmentation map without image size and class area dominating the likelihood. RLPO partitions \(R\) into class buckets using the winner's hard prediction. The same buckets are used for the winner, loser, current model, and reference model, so the comparison does not silently change its spatial partition. Averaging within each present class and then across classes prevents large backgrounds from overwhelming small structures.

Let \(R_c\) contain pixels predicted as class \(c\) by the winner within the region, let \(U_R\) contain the nonempty class buckets, and let \(\hat Y^k\) be the current model's hard prediction under template \(k\) for this step. Equation (7) defines the class-balanced score as:

\[ S_\theta^k(R)=\frac{1}{|U_R|}\sum_{c\in U_R}\frac{1}{|R_c|}\sum_{u\in R_c}\log P_\theta^k\bigl(\hat Y^k(u)\mid x,u\bigr). \]

Although the winner defines the buckets, each template scores its own hard labels rather than always scoring the winner's labels. The reference score replaces current probabilities with frozen-model probabilities, retaining the same step-fixed candidate labels and class buckets. Hard candidates remain fixed during one update; the updated model generates fresh candidates for the next image, making this an online preference-learning procedure.

The reference is the pre-adaptation model, not the previous-step model or an external teacher. It anchors score changes to each template's initial behavior, rather than rewarding a template simply for having higher initial confidence. The preference loss encourages a larger score increase relative to the reference for the winner than for the loser. It is neither pixelwise cross-entropy against the oracle's ground truth nor dense ground-truth supervision using the winner mask throughout the query rectangle.

The design borrows DPO's Bradleyโ€“Terry form, but its practical objective should be distinguished from a strict derivation. Appendix A explicitly treats class-balanced averaging as a deliberate modification for class imbalance, not a necessary consequence of full-map likelihood factorization. Furthermore, winner and loser arise from different template contexts, so the normalizer cancellation used in standard shared-context DPO does not automatically apply.

Appendix A.2 argues that the cross-template normalization offset can be omitted inside the log-sigmoid without affecting gradients because it is parameter-independent. This argument is insufficient: a constant inside a nonlinear function has zero derivative itself but still changes that function's gradient weighting of the score difference. The empirical gains support the practical objective, but RLPO should not be presented as an established strict equivalent of the original KL-regularized DPO objective.

3. Outside-region consistency: local preference correction should not destabilize unqueried predictions

Regional preferences constrain scores inside \(R\), but shared adapter updates can still change predictions elsewhere. The authors use the winner's hard prediction as a pseudo-label to constrain the loser probability map outside the region. They do not unconditionally impose the winner over the entire image or request additional outside-region annotations.

The supervision mask intersects two conditions: a pixel must lie outside \(R\), and the winner's confidence must exceed \(\tau_{\mathrm{conf}}\). Lovรกszโ€“Softmax loss is computed only on these pixels, with a default threshold of 0.8. This separates preference supervision from stabilization: inside the rectangle the model learns which hypothesis is preferable, while outside it high-confidence pseudo-labels limit loser-branch drift.

This remains model-based self-supervision, not a guarantee of correct pseudo-labels. Confident errors may persist. Moreover, shared parameters transmit changes across space, so spatially masked loss computation does not mean parameter updates affect only masked pixels. Ablations show a complementary benefit, while RLPO drives the larger gain.

A Worked Example

Consider a chest radiograph and its target vocabulary. The model produces structure predictions under 14 templates. Ensemble entropy identifies a high-entropy connected region near a boundary, and the interface shows its bounding box alongside the two most disagreeing candidates within it. An annotator chooses the closer prediction without drawing a boundary. This is an illustrative walkthrough, not a claim that the paper reports a particular annotated image outcome.

After selecting the winner, the method scores each candidate's own hard labels within winner-defined class buckets and compares those scores against the frozen reference. Outside the rectangle, only pixels where winner confidence exceeds 0.8 constrain the loser through pseudo-labels. The combined objective updates the adapters once before the next training image arrives. Final inference returns to the baseline's default template configuration and predicts a complete segmentation, rather than returning the training-time preferred candidate.

Loss & Training

With \(\Delta_k=S_\theta^k(R)-S_{\mathrm{ref}}^k(R)\), the core objectives in Equations (8) and (10) can be written compactly as:

\[ \mathcal L_{\mathrm{RLPO}}=-\log\sigma\!\left(\beta[\Delta_w-\Delta_l]\right),\qquad \mathcal L_{\mathrm{total}}=\mathcal L_{\mathrm{RLPO}}+\lambda_{\mathrm{cons}}\mathcal L_{\mathrm{cons}}. \]

Defaults are \(\beta=0.1\) and \(\lambda_{\mathrm{cons}}=0.1\), with batch size 1, AdamW, and weight decay \(10^{-4}\). Rank-4 vision LoRA attaches to the Q and V projections of the last four CLIP Transformer blocks. The text adapter is a rank-4 low-rank residual on the mean-template classifier embedding, with parameters shared across classes and templates.

CAT-Seg uses a default learning rate of \(3\times10^{-3}\), reduced to \(10^{-3}\) for CUB-200, ATLANTIS, iSAID, and PST900. SAN uses \(10^{-3}\), reduced to \(3\times10^{-4}\) for CUB-200, ATLANTIS, and PST900. These are the actual configurations in Appendix C and Table 6, so learning rates are not identical across all domains.

Key Experimental Results

Main Results

MESS spans general scenes, earth monitoring, medical sciences, engineering, and agriculture & biology. The paper excludes Dark Zurich, DRAM, ISPRS Potsdam, and CryoNuSeg because they lack training splits or are no longer publicly accessible, leaving 18 datasets. Overall Mean averages across datasets, not simply across the five group scores. Adaptation samples come from training splits, and evaluation uses official held-out splits.

The default budget is 64 adaptation images per dataset, with one update per image; CHASE_DB1 uses only 8 images and CWFID uses 32. Adaptation results average three runs with independently sampled images, and the main table includes standard deviations. Crucially, the main oracle compares candidates using regional ground-truth IoU. The optimizer receives a binary outcome rather than dense ground truth, but producing experimental feedback still depends on ground-truth masks.

The following table selects overall mIoU (%) from Table 1. Gain is a percentage-point difference from the same backbone's zero-shot baseline, not an improvement over a universal previous SOTA.

Backbone Zero-shot Ours Dense-mask single-step reference Gain
SAN-B 25.29 32.39 ยฑ 0.09 32.41 ยฑ 0.46 +7.10
CAT-Seg-B 33.12 39.67 ยฑ 0.08 44.26 ยฑ 0.49 +6.55
SAN-L 27.40 33.99 ยฑ 0.90 37.64 ยฑ 0.15 +6.59
CAT-Seg-L 35.26 45.88 ยฑ 0.38 46.67 ยฑ 0.74 +10.62

Dense-mask preserves the same single-step streaming protocol and changes only supervision to dense ground truth; it is not a fully trained upper bound. Appendix E, Table 9 reports fully supervised CAT-Seg-L prompt tuning on the same 64 images for 200 epochs at 53.17 ยฑ 0.50. The proposed method does not match this stronger optimization-budget reference.

Ablation Study

All entries below use CAT-Seg-L and are selected from Tables 2 and 3 and Appendix B, Table 5. The first four rows compare losses, while the remaining rows examine candidate sources or template diversity; these are different diagnostic types rather than one uniform module-removal experiment.

Config Overall mIoU (%) Note
Full model 45.88 14 templates, RLPO + consistency
Without RLPO 40.51 5.37 points below full model
Without consistency 43.45 2.43 points below full model
Zero-shot 35.26 No adaptation
MC Dropout candidates 39.09 14 candidates, dropout probability 0.1
Test-time augmentation candidates 44.66 14 candidates, flips and 0.75โ€“1.25 scaling
Sentence-only templates 43.79 5 templates without scale modifiers
Scale-only templates 43.65 Fixed sentence, 4 scale-modifier settings

Appendix D, Table 7 separately tests humanโ€“oracle agreement: 20 participants judge 30 pairs each, with 10 pairs per difficulty tier, totaling 600 judgments at an average of 9.4 seconds each. This validates preference agreement, not a complete adaptation run driven by those human judgments.

Tier Regional IoU margin Judgments Humanโ€“oracle agreement
Easy โ‰ฅ0.18 200 0.967
Medium [0.04, 0.18) 200 0.953
Hard <0.04 200 0.840
Overall โ€” 600 0.920

Key Findings

  • RLPO provides the main gain, and consistency supplies additional stabilization. Prompt candidates outperform MC Dropout and augmentation overall, but medical-group augmentation reaches 53.15 versus 51.83 for prompt candidates; prompts are not strongest in every group.
  • Table 4 reports mean mIoU of 38.26, 42.31, 45.88, and 45.79 for 4, 16, 64, and 128 images. Gains are not strictly monotonic with sample count, and average performance saturates beyond 64 images.
  • The main text says medical sciences has the largest improvement for every backbone, but Table 1 gives SAN-L a medical gain of only 6.62 points, versus 11.98 in engineering and 11.91 in agriculture & biology. The table facts are retained rather than repeating that overgeneralization; CAT-Seg-L's medical gain of 22.31 does agree with the table.
  • Appendix D, Table 8 targets near-tie queries: inverting every query with an IoU margin below 0.15 still yields 42.43, above 35.26. Figure 4's roughly 8-point gain under 20% uniformly flipped preferences is an approximate plotted finding, not an exact measurement supplied here.

Highlights & Insights

  • Candidate diversity can serve as an interaction interface rather than merely noise to remove at inference time. Direct prompt ensembling in Appendix E reaches only 34.74, below the 35.26 baseline, so gains cannot be attributed simply to combining more templates.
  • A binary preference can influence many pixels through a query region, class buckets, and candidate confidence. โ€œDense supervisionโ€ describes the spatial effect of optimization, not the information content of a complete correct mask in one binary judgment.
  • The division into local querying, local preference correction, and outside-region stabilization can transfer to other interactive dense-prediction tasks. Candidates must remain semantically comparable, and regions without human supervision must be identified explicitly.

Limitations & Future Work

  • The authors acknowledge that preferences become weak when all templates lack useful hypotheses. Expert templates, domain-specific descriptions, or learned templates could extend the candidate space beyond natural-image-style sentences.
  • Main results rely on a ground-truth oracle. The 600-judgment study does not establish expert usability across all specialized domains, long-term adaptation gains, or total annotation costs. The 9.4-second figure covers individual judgments, not training compute, system latency, or expert preparation.
  • Confident outside-region pseudo-label errors may become entrenched. Abstention, skipping near-ties, or multiple queries are possible extensions, but additional feedback must be included in the supervision budget rather than treated as free.
  • The nonlinear constant-offset issue in Appendix A.2 warrants a stricter objective interpretation or an offset ablation. Empirical performance and theoretical equivalence should be evaluated separately.
  • Appendix H, Table 11 reports 1,616 ms per CAT-Seg-L step with 7,335 MiB peak memory, and 418 ms per SAN-L step with 8,392 MiB, on one H200 with BDD100K. Both train only 71,681 parameters, but 14 candidate forwards still incur compute costs; obtaining diversity from an existing interface does not mean zero computational overhead.
  • vs SAN / CAT-Seg: These provide the OVSS backbone. The paper adds lightweight adaptation and preference objectives without replacing the main architecture. Gains across two architectures support some generality, not coverage of all OVSS systems.
  • vs DPO: The method retains winnerโ€“loser comparison relative to a frozen reference without training an explicit reward model. Candidates arise from different prompt contexts, however, and scores are localized and balanced by winner-defined classes; strict equivalence requires further argument.
  • vs fixed-target segmentation preference optimization: The paper distinguishes prior medical or small closed-label tasks from open-vocabulary adaptation. The additional contribution is constructing local queries through template disagreement, not introducing preferences to visual segmentation for the first time.
  • vs dense-mask fine-tuning: The goal is a lower-burden proposed interaction interface, not proof that binary feedback beats every fully trained mask-supervised method. Equal image budgets and equal optimization-step budgets must also be discussed separately.

Rating

  • Novelty: 4/5 โ€” Combines prompt disagreement, local queries, and class-balanced preference learning with a clear task-specific purpose.
  • Experimental Thoroughness: 4/5 โ€” Four backbone configurations, ablations, and human agreement tests are substantial, but actual human-driven adaptation remains insufficiently validated.
  • Writing Quality: 3/5 โ€” The pipeline is clear, but a group-level gain claim conflicts with the table and the cross-template derivation has an argumentative gap.
  • Value: 4/5 โ€” Offers a reusable interface for low-annotation-budget open-vocabulary adaptation; deployment value depends on candidate quality and feedback reliability.