Skip to content

Calibrate Before Adapt: Training-Free Pseudo-Label Calibration for Semi-Supervised Cross-Domain Few-Shot Detection

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/xSheep123/Semi-CDFSOD.git
Area: Object Detection
Keywords: Cross-Domain Few-Shot Object Detection, Semi-Supervised Learning, Pseudo-Label Calibration, Visual Self-Prompting, Training-Free

TL;DR

To tackle severe semantic misclassification and low recall in pseudo-labels generated by open-set detectors for semi-supervised cross-domain few-shot object detection (Semi-CDFSOD), this paper introduces a training-free pseudo-label calibration framework that couples center-weighted prototype semantic calibration with SAM-based visual self-prompting to purify and complement unlabeled target-domain supervision.

Background & Motivation

Cross-domain few-shot object detection (CDFSOD) aims to adapt detectors to novel, disparate target domains under severe label scarcity. However, mainstream paradigms operate under a restrictive assumption disconnected from real-world deployments: they assume the target domain offers nothing beyond a tiny set of labeled support examples, while completely ignoring the abundant, annotation-free unlabeled data readily accessible in practical environments. Combining semi-supervised learning with cross-domain few-shot detection (Semi-CDFSOD) to bridge substantial domain gaps and label distribution shifts offers immense real-world value, yet remains largely unexplored.

A straightforward avenue to exploit unlabeled target data is pseudo-labeling via open-set detectors (e.g., MM-Grounding-DINO) fine-tuned on few-shot target supports. Nevertheless, under severe domain shifts and data scarcity, this conventional strategy suffers from two compounding intrinsic defects. The first is semantic misclassification: while open-set detectors localize novel domain objects reasonably well, their feature representations inadvertently activate source-domain decision boundaries, causing confident yet erroneous category assignments or false background suppression. The second is low recall: the severe distribution shift causes target representations to drift outside the detector's expected manifold, leading to the omission of the majority of foreground objects. Conventional confidence thresholding creates a dilemma: low thresholds flood training with noisy labels, whereas high thresholds retain only trivial instances and fail to rectify category errors, causing corrupted supervision to cascade in self-training.

Rather than passively relying on complex, noise-robust adaptation losses after pseudo-labels are already tainted, pseudo-label calibration must take precedence. The core insight is that scarce labeled support samples, while insufficient to train a detector from scratch, provide pristine semantic anchors to calibrate noisy pseudo-labels; concurrently, foundation models such as the Segment Anything Model (SAM) offer class-agnostic cross-domain promptable segmentation, serving as an independent second opinion to retrieve undetected objects. The core idea is to propose a training-free pseudo-label calibration framework that rectifies semantic misclassifications via center-weighted class prototypes built from supports and confident pseudo-labels, while leveraging calibrated detections as visual self-prompts for SAM to recover missed instances, delivering high-quality pseudo-labels for stable detector adaptation.

Method

Overall Architecture

The framework comprises two complementary, training-free modules: the Semantic Calibrator and Missed Object Recovery via Visual Self-Prompting.

In the pipeline, initial candidate pseudo-labels are first generated on unlabeled target images using an open-set detector (MM-Grounding-DINO) fine-tuned on target-domain few-shot supports. Next, center-weighted regional feature representations are extracted, and robust target-domain class prototypes are constructed by combining ground-truth supports with top-confidence pseudo-labels to recalibrate category assignments via cosine similarity. Subsequently, the highest-confidence calibrated detection per class serves as a primary visual promptโ€”augmented with class-aware attention guidanceโ€”to prompt SAM to discover previously overlooked instances of the same class. Recovered candidates are verified by the semantic calibrator, and non-maximum suppression (NMS) merges all detections into a refined, high-recall pseudo-label pool combined with the support set for detector fine-tuning.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Unlabeled Target Images + Few-Shot Supports"] --> B["Initial Candidate Proposal Generation<br/>MM-Grounding-DINO Coarse Screening"]
    B --> C["Center-Weighted Regional Feature Extraction & Prototype Construction<br/>Merging Supports & High-Confidence Pseudo-Labels"]
    C --> D["Prototype-Guided Category Calibration<br/>Cosine Similarity Category Assignment"]
    D --> E["Missed Object Recovery via Visual Self-Prompting<br/>High-Quality Boxes as Prompts to SAM"]
    E --> F["Calibration Verification & NMS Duplicate Removal<br/>Purified High-Recall Pseudo-Label Set"]
    F --> G["Final Detector Adaptation & Fine-Tuning"]

Key Designs

1. Center-Weighted Regional Feature Extraction and Class Prototype Construction: Suppressing Peripheral Noise and Anchoring Target Representations

Conventional average RoI pooling introduces peripheral background noise into region features, degrading semantic discriminability under domain shift. Since the core object typically occupies the center of a localized bounding box, dense feature maps are extracted using a frozen DINOv3 (ViT-L/16) backbone, followed by a distance-decaying center-weighted aggregation strategy. Let \(\mathcal{O}\) denote the set of patch coordinates in the RoI and \((c_x, c_y)\) its center. For each coordinate \((x, y) \in \mathcal{O}\), a center-aware weight is assigned:

\[w_{\mathrm{center}}(x, y) = \exp\left(-\frac{(x - c_x)^2 + (y - c_y)^2}{2\sigma^2}\right)\]

After normalization \(\hat{w}(x, y) = \frac{w_{\mathrm{center}}(x, y)}{\sum_{(i, j) \in \mathcal{O}} w_{\mathrm{center}}(i, j) + \epsilon}\), the regional feature is computed as \(\mathbf{f}_{\mathrm{box}} = \sum_{(x, y) \in \mathcal{O}} \hat{w}(x, y) \mathbf{F}(x, y)\). To construct representative target-domain class prototypes, each class \(c\) incorporates its labeled few-shot supports alongside the top-3 highest-confidence initial pseudo-labels, computing their mean feature \(\mathbf{p}_c = \frac{1}{N_c} \sum_{i=1}^{N_c} \mathbf{f}_i^c\). This fuses the pristine ground-truth semantics of the supports with the empirical distribution diversity of confident target-domain pseudo-labels.

2. Prototype-Guided Category Calibration: Rectifying Semantic Misclassifications via Cosine Similarity

Due to source-domain classifier bias, initial detections frequently suffer from category confusion despite accurate localization. Because the constructed prototypes are grounded in DINOv3's rich visual representation space, category reassignment bypasses the detector's biased classification head entirely. For any region feature \(\mathbf{f}\) extracted from an initial proposal, its cosine similarity against all class prototypes \(\{\mathbf{p}_c\}_{c=1}^C\) is evaluated, assigning the class with maximum alignment:

\[\hat{y} = \arg\max_{c \in \{1, \dots, C\}} \frac{\mathbf{f}^\top \mathbf{p}_c}{\|\mathbf{f}\|_2 \|\mathbf{p}_c\|_2}\]

This parameter-free mechanism preserves high-quality geometric bounding boxes while correcting misassigned labels (e.g., confusing industrial defect types or confusing vehicles with background clutters) at near-zero computational cost.

3. Missed Object Recovery via Visual Self-Prompting: Guiding SAM to Overcome Detector Recall Bottlenecks

To resolve the low recall inherent in open-set detectors under distribution drift, textual prompts often fail due to cross-domain linguistic and visual ambiguity. Instead, calibrated high-confidence pseudo-labels are repurposed as visual prompts for SAM. For each class \(c\), the pseudo-label with highest confidence \(z_*^c = \arg\max_{z_i \in \hat{\mathcal{Z}}^c} s_i\) acts as the primary visual prompt box. To encourage SAM to explore beyond a single localized region, a class-aware spatial enhancement mask \(m^c(x, y)\) is defined over remaining candidate boxes belonging to class \(c\). Scaling this map by \(\lambda = 5\) produces \(\mathbf{E}^c = \lambda m^c(x, y)\), which is injected into the relative position bias:

\[\mathbf{A}^c = \mathbf{A}_{\mathrm{rpb}}(z_*^c) + \mathbf{E}^c\]

This guided cross-attention bias directs SAM to discover visually similar instances of the same category across the scene. All newly recovered proposals are routed through the semantic calibrator to prune false positives, before NMS merges them with calibrated detections, achieving a closed-loop pseudo-label set that is both accurate and comprehensive.

Loss & Training

The pseudo-label calibration and missed-object recovery phases are completely training-free, requiring zero back-propagation or weight updates. In the subsequent target-domain adaptation stage, the base detector MM-Grounding-DINO with Swin-B (initialized from LLMDet) is adapted using the combined supervision of the calibrated pseudo-labels and labeled few-shot supports. Fine-tuning runs for 500 iterations using the AdamW optimizer with a base learning rate of \(2 \times 10^{-5}\) (\(1 \times 10^{-4}\) for ArTaxOr and DIOR), and the backbone learning rate scaled down by 0.1 to maintain stable cross-domain adaptation.

Key Experimental Results

Main Results

The authors establish the first standardized Semi-CDFSOD benchmark spanning 6 diverse domains: ArTaxOr (biological art/macro), DIOR (optical remote sensing), Clipart1k (artistic style transfer), UODD and DeepFish (underwater marine environments), and NEU-DET (industrial steel defect inspection). The table below summarizes detection performance (mAP) under 1-shot, 5-shot, and 10-shot settings.

Method Type 1-shot Avg. mAP 5-shot Avg. mAP 10-shot Avg. mAP
Detic-FT Traditional CDFSOD Fine-tuning 6.6 13.3 16.5
DE-ViT-FT Traditional CDFSOD Visual Tuning 10.1 22.3 25.2
CD-ViTO Enhanced Open-Set CDFSOD SOTA 13.9 26.1 29.6
MM-Grounding-DINO Open-Set Detector Baseline (Few-shot FT) 31.2 40.6 45.1
Threshold Filtering (thr=0) Traditional Semi-supervised PL (All retained) 19.7 21.0 23.6
Threshold Filtering (thr=0.5) Traditional Semi-supervised PL (Standard threshold) 31.8 41.3 45.3
Threshold Filtering (thr=0.9) Traditional Semi-supervised PL (Strict threshold) 32.8 41.1 44.5
SAM3 (Text Prompts) Foundation Model Prompting Baseline 28.8 33.8 37.9
SAM3 (Visual Prompts) Foundation Model Prompting Baseline 32.1 40.5 44.1
SAM3 (Text & Visual Prompts) Foundation Model Prompting Baseline 31.3 37.9 40.1
Ours Training-Free Calibration Framework 35.9 44.1 46.7

Per-dataset performance comparisons under 1-shot and 5-shot settings:

Dataset Domain Shift 1-shot Baseline 1-shot Ours 1-shot Gain 5-shot Baseline 5-shot Ours 5-shot Gain
ArTaxOr Biological Taxonomy 44.0 56.2 +12.2 67.6 78.6 +11.0
Clipart1k Artistic Style 49.6 50.1 +0.5 55.6 55.8 +0.2
DeepFish Underwater Species 39.0 46.4 +7.4 41.0 43.3 +2.3
DIOR Aerial Remote Sensing 18.6 23.7 +5.1 29.6 32.5 +2.9
NEU-DET Industrial Steel Surface 13.9 16.8 +2.9 22.2 23.8 +1.6
UODD Underwater Multi-Species 22.2 22.0 -0.2 27.8 30.3 +2.5
Average (Avg.) All Domains 31.2 35.9 +4.7 40.6 44.1 +3.5

Ablation Study

A progressive ablation study on DIOR (20 classes, remote sensing) and NEU-DET (6 classes, steel surface defects) demonstrates the complementary individual and joint benefits of the calibrator and visual self-prompting modules.

Dataset Config / Module Combination 1-shot mAP 5-shot mAP 10-shot mAP Overall Gain vs. Baseline
DIOR Baseline (MM-Grounding-DINO) 18.6 29.6 35.7 -
DIOR + Calibrator 20.6 32.1 36.0 +2.0 / +2.5 / +0.3
DIOR + Calibrator & Visual Self-Prompting (Ours) 23.7 32.5 37.3 +5.1 / +2.9 / +1.6
NEU-DET Baseline (MM-Grounding-DINO) 13.9 22.2 24.3 -
NEU-DET + Calibrator 16.5 23.5 25.6 +2.6 / +1.3 / +1.3
NEU-DET + Calibrator & Visual Self-Prompting (Ours) 16.8 23.8 25.9 +2.9 / +1.6 / +1.6

Key Findings

  • Semantic calibrator brings immediate gains: In extreme 1-shot scenarios, introducing the prototype calibrator alone yields a +2.0 mAP increase on DIOR and +2.6 mAP on NEU-DET. Confusion matrix comparisons confirm that off-diagonal misclassifications are noticeably suppressed while diagonal correct predictions are boosted.
  • Visual self-prompting resolves recall bottlenecks: Building on calibrated detections, SAM guided by class-aware spatial attention unearths previously missed low-contrast and challenging objects, driving an additional +3.1 mAP surge on DIOR 1-shot.
  • Visual prompts distinctly outperform text prompts: Foundation model prompting with visual patches achieves 32.1/40.5/44.1 mAP compared to 28.8/33.8/37.9 mAP for text prompts. This highlights that natural language descriptions encounter severe semantic ambiguity under extreme domain shifts, whereas in-domain visual prompt exemplars deliver far more reliable grounding.

Highlights & Insights

  • "Calibrate Before Adapt" paradigm shift: Traditional semi-supervised detection methods typically compensate for label noise during fine-tuning via complex loss reweighting or heuristic threshold adjustments. This paper establishes that pseudo-label quality itself is the root bottleneck in Semi-CDFSOD, proving that offline, training-free geometric and visual calibration prior to adaptation yields superior performance with simpler mechanics.
  • Decoupled representation via center-weighted DINOv3: Leveraging frozen, pre-trained dense visual features with Gaussian distance weighting effectively isolates object semantics from background noise, enabling robust class prototype construction without requiring target-domain fine-tuning.
  • Closed-loop self-prompting: Recycling confident calibrated boxes as visual prompts for SAM effectively bridges the gap between foundation segmentation models (which lack semantic classification) and open-set detectors (which lack recall under domain drift).

Limitations & Future Work

  • Reliance on initial bounding box localization: The semantic calibration module assumes that candidate proposals roughly intersect the target instance to benefit from center-weighted pooling. Grossly mislocalized or severely drifted proposals may contaminate prototype matching.
  • Sequential multi-stage offline overhead: Extracting dense DINOv3 features and querying SAM across multiple classes on massive unlabeled target datasets involves non-negligible offline computation before detector adaptation can begin.
  • Future directions: Investigating online dynamic prototype updating and integrating lightweight visual prompt-guided segmentation heads directly into single-stage architectures could enable real-time, streaming pseudo-label calibration.
  • vs. CD-ViTO & Traditional CDFSOD: Prior CDFSOD methods strictly depend on target labeled supports, ignoring unlabeled data. Expanding to the Semi-CDFSOD paradigm delivers substantial margins (e.g., boosting 1-shot average mAP from 13.9 to 35.9).
  • vs. Threshold-based Pseudo-Labeling: Conventional thresholding fails to address domain-shift-induced confident misclassifications and missed detections. This paper breaks through the fixed-threshold upper bound via prototype alignment and prompt-driven retrieval.
  • vs. Naรฏve Foundation Model Prompting: Direct prompt-based detection using uncalibrated pseudo-labels suffers from prompt noise amplification. Performing prototype calibration before visual prompting prevents cascading errors.

Rating

  • Novelty: โญโญโญโญโ˜† Pinpoints the core failure modes of pseudo-labels in Semi-CDFSOD and designs an elegant, training-free calibration and retrieval framework.
  • Experimental Thoroughness: โญโญโญโญโญ Establishes a standardized 6-domain benchmark with comprehensive baselines and ablations across 1/5/10-shot settings.
  • Writing Quality: โญโญโญโญโญ Cohesive narrative, rigorous formulation, and self-consistent empirical analysis.
  • Value: โญโญโญโญโญ Delivers an effective, plug-and-play recipe for real-world few-shot domain adaptation using low-cost unlabeled data.