When Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-Labeling¶
Conference: NeurIPS2026
arXiv: 2609.33668
Code: https://github.com/pingggg516/FTC-Seg
Area: Segmentation
Keywords: semi-supervised semantic segmentation, long-tailed learning, pseudo-labeling, prototype reconstruction, adaptive thresholds
TL;DR¶
FTC-Seg combines prototype-based residual feature reconstruction with pseudo-label thresholds driven by labeled accuracy and class-distribution bias in teacher–student segmentation, improving DINOv2-S over UniMatch V2 by 2.09–8.54 mIoU percentage points across four benchmarks with 2% labels, without winning every configuration.
Background & Motivation¶
Semi-supervised semantic segmentation typically lets a teacher predict on a weakly augmented unlabeled image and passes confident pixel labels to a student processing a strongly augmented view. FixMatch and UniMatch use thresholds to exclude unreliable predictions, but this filtering also determines which classes can receive additional supervision. In sonar, underwater optical imagery, and adverse weather, object features are harder to separate from the background, and even correct predictions may have insufficient confidence. Being rejected by a threshold therefore does not necessarily mean being incorrect.
Long-tailed distributions concentrate this problem on rare classes: fewer labeled pixels weaken representation learning, while imaging noise further suppresses confidence. A fixed high threshold removes these classes first; the resulting gap in unlabeled supervision further weakens their representations, creating a self-reinforcing cycle. Resampling or distribution adjustment cannot recover supervision already lost before admission, while consistency training cannot train pixels that never pass the filter. The paper studies this interaction through a \(2\times2\) design crossing original/synthetically corrupted images with balanced/long-tailed subsets. However, balancing changes the sample composition through subsampling, and the original sonar and underwater images are not noise-free.
Core Idea: intervene at both ends of pseudo-label generation—first improve separability with foreground-gated prototype residual reconstruction, then determine class-specific admission thresholds from learning difficulty and distribution bias so that low-confidence rare classes are not permanently excluded.
Method¶
Overall Architecture¶
Inputs consist of pixel-labeled images and weakly and strongly augmented views of the same unlabeled images. The student processes labeled images and strong views, while the teacher processes weak views. Both use the same encoder–decoder architecture, but teacher parameters are updated through an exponential moving average (EMA) of student parameters rather than independent back-propagation.
Orthogonal Prototype Reconstruction (OPR) sits between the encoder and decoder in both branches. Its default placement is the two deepest encoder stages: learnable prototypes shared across samples provide residual corrections to local features, rather than producing a denoised image. After decoding, Adaptive Threshold Calibration (ATC) uses learning statistics on labeled data and prediction distributions on unlabeled data to select class-dependent thresholds for teacher predictions.
Accepted teacher labels supervise student predictions on the strong views, while labeled images provide ground-truth supervision. Student updates are followed by teacher EMA updates. Inference only requires the encoder, OPR, and decoder to output a class map; ATC selects training supervision and is not a test-time probability calibrator.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
X["Labeled images and<br/>two unlabeled views"] --> E["Student encoder and<br/>weak-view EMA teacher"]
E --> O["Orthogonal Prototype<br/>Reconstruction: OPR"]
O --> D["Student and teacher decoding<br/>pixel class predictions"]
D --> A["Adaptive Threshold<br/>Calibration: ATC"]
G["Labeled accuracy and<br/>labeled/unlabeled distributions"] -.-> A
A -->|Filter teacher pseudo-labels| L["Training supervision<br/>strong-view consistency and ground truth"]
L -.->|Parameter EMA after student update| E
D -->|Inference bypasses ATC| Y["Output segmentation map"]
Key Designs¶
1. Orthogonal Prototype Reconstruction: preserve original detail while adding foreground-related global feature corrections
Local pixel features are vulnerable to imaging noise, so OPR provides a more stable reference using a set of learnable vectors shared across the dataset. The default number of prototypes is \(K=16\), and each has the channel dimension of the encoder features. Each pixel first computes cosine similarities with all prototypes, which then determine its correction directions. These prototypes are not one center per class: they have no predefined fixed class identities, and their dominant classes can change during training.
Not every global pattern should be added back. OPR applies the segmentation head's linear classification weights and biases to each prototype and identifies its strongest class response. Only when that class belongs to the predefined foreground set \(\mathcal C_{fg}\) does the gate \(\gamma_k\) equal 1; otherwise it is 0. This gate comes from the segmentation head's class decision, not from thresholding prototype similarity or directly reading a ground-truth foreground mask for the pixel. Admitting background prototypes could reintroduce the interference that reconstruction is intended to suppress.
The central transformation retains the original pixel feature \(F_{(h,w)}\):
Cosine similarities can be positive or negative. The original equation directly sums signed contributions, without softmax or a normalized convex combination. The residual can strengthen or cancel directions, so OPR should not be interpreted as replacing a pixel with a selected class-dictionary center. Retaining the original features means the global correction need not carry all spatial detail. In the ablation, removing the residual hurts overall segmentation despite high tail-class pseudo-label accuracy.
Prototypes that converge to similar directions respond redundantly and lose diversity. The paper applies a soft orthogonality regularizer based on the sum of squared pairwise inner products of normalized prototypes to discourage directional collapse. It encourages approximate orthogonality rather than imposing an exactly orthogonal basis at every step. Teacher EMA smooths student parameters, whereas OPR improves the representation used by both branches; these mechanisms address different issues.
The foreground set still requires a task definition. Adding background to the gating set or excluding Diver harms results in the appendix, so OPR is not an unconditional denoiser for arbitrary label semantics. Shallow features may also be unsuitable for this gate: segmentation-head weights primarily encode deep semantics, and applying them to shallow prototypes can incorrectly classify background prototypes as foreground. The default therefore does not place OPR at every stage.
2. Adaptive Threshold Calibration: correct admission bias using labeled learning state and class priors
Confidence on unlabeled data alone can make repeated over-prediction look like successful learning. ATC maintains three EMA statistics per class: its proportion in labeled data \(P_c^l\), classification accuracy on labeled data \(A_c^l\), and predicted proportion on unlabeled data \(P_c^u\). The statistics use EMA momentum 0.999 to reduce mini-batch fluctuations, especially in rare-class counts. This statistical smoothing is a separate state update from teacher-parameter EMA.
The difficulty signal measures the gap in labeled accuracy: a poorly learned class needs a less restrictive supervision entrance. The bias signal compares the unlabeled prediction proportion with the labeled prior, tightening admission for relative over-prediction and relaxing it for under-prediction. Together they modify the base threshold rather than uniformly lowering thresholds for every rare class:
The defaults are \(\lambda_D=0.1\) and \(\lambda_B=0.05\), with \(\epsilon\) stabilizing the denominator. For a difficult but severely over-predicted class, the two corrections can offset one another, so its final threshold need not decrease. For a well-learned but under-predicted class, the bias term can still relax admission. Rescuing tail classes is therefore an intended effect supported by observations, not a hard rule triggered by tail-class identity.
The teacher first selects the maximum-probability class and compares its probability with the threshold for that predicted class. Only pixels whose maximum probability is strictly greater than the threshold form valid one-hot pseudo-labels. ATC does not change class probabilities or apply temperature scaling; it changes the pixels participating in consistency loss. OPR improves feature quality, while ATC changes access to supervision, leaving room for complementary gains.
The prior reference also has a cost: if labeled and unlabeled data have different true distributions, the ratio may reflect genuine distribution shift rather than prediction error. The paper does not specify threshold clipping, the numerical value of \(\epsilon\), or complete handling of batches missing classes. A \([0,1]\) clamp or other implementation safeguards should not be invented. The main text also does not clearly establish a universal base-threshold default; the fixed threshold of 0.95 in a particular appendix OPR-gating experiment should not automatically be assigned to all ATC main experiments.
A Worked Example¶
Consider a sonar image containing a rare Diver target. Its local pixel features are mixed with the background but retain some similarity to several foreground prototypes. OPR first uses the segmentation head to determine which prototypes can participate, then adds signed prototype corrections to the original features. The decoder subsequently produces teacher class predictions on the weak view; no clean sonar image is generated before segmentation.
To illustrate ATC, suppose the current statistics for this class are labeled accuracy 0.60, labeled proportion 0.01, and unlabeled predicted proportion 0.005. Set \(\tau_0=0.95\) only for this example. With the default two coefficients and a negligible \(\epsilon\), the difficulty term is approximately 0.40, the bias term approximately −0.50, and the threshold approximately 0.885. A teacher prediction of Diver with probability 0.90 would fail a fixed 0.95 threshold but could supervise the student's strong view under the illustrative threshold.
Conversely, a class with labeled accuracy 0.90 and a predicted proportion approximately twice its prior would receive a threshold of approximately 0.99 under the same illustrative setup, rejecting even a probability of 0.97. ATC can thus both recover low-confidence predictions and restrict over-prediction rather than relaxing all thresholds. These numbers demonstrate the mechanism; they are not paper samples or reproduced results.
Loss & Training¶
The student uses cross-entropy on labeled pixels and unsupervised cross-entropy on filtered pseudo-labels, with a soft orthogonality loss constraining learnable prototypes. The unsupervised loss is normalized by the full image area \(HW\). Rejected pixels have zero pseudo-label vectors and contribute no loss; this should not be silently rewritten as normalization by the number of accepted pixels.
The soft orthogonality term uses normalized prototypes \(\hat\mu_k\):
The teacher receives no back-propagation. The paper describes supervised, unsupervised, and prototype-regularization terms but does not provide enough information to establish every total-loss weight. No complete objective with assumed coefficients is constructed here.
Experiments train for 240 epochs with AdamW, learning rate \(5\times10^{-6}\), weight decay 0.01, and polynomial learning-rate decay with power 0.9 on one RTX 3090. Labeled proportions are 2%, 5%, and 10%, sampled to include at least one labeled instance per class; the remaining images are unlabeled.
Key Experimental Results¶
Main Results¶
The four benchmarks cover acoustic speckle in FLSMD and FSSG, underwater optical scattering in SUIM, and adverse weather in ACDC. Main tables report mean ± standard deviation over three random seeds, with mIoU in percent. Differences below are percentage points and compare methods within the same backbone rather than merging different backbones into a single SOTA claim.
| Dataset / Labels | Backbone | FTC-Seg | Comparison | Comparison Result | Difference |
|---|---|---|---|---|---|
| FLSMD / 2% | DINOv2-S | 62.55 ± 1.22 | UniMatch V2 | 57.24 ± 0.49 | +5.31 |
| FSSG / 2% | DINOv2-S | 63.20 ± 0.98 | UniMatch V2 | 58.78 ± 1.32 | +4.42 |
| SUIM / 2% | DINOv2-S | 66.91 ± 1.35 | UniMatch V2 | 58.37 ± 0.92 | +8.54 |
| ACDC / 2% | DINOv2-S | 61.68 ± 1.42 | UniMatch V2 | 59.59 ± 2.48 | +2.09 |
| SUIM / 2% | RN-101 | 54.71 ± 0.53 | SemiVL | 55.99 ± 0.79 | −1.28 |
| FSSG / 10% | RN-101 | 70.13 ± 1.14 | SemiVL | 65.24 ± 0.97 | +4.89 |
| FSSG / 10% | DINOv2-S | 67.53 ± 0.68 | UniMatch V2 | 64.41 ± 0.75 | +3.12 |
FTC-Seg trails SemiVL with RN-101 and 2% labels on SUIM. On FSSG with 10% labels, its RN-101 result also exceeds its own DINOv2-S result. The conclusion's broad superiority wording is stronger than the main tables support: the method is competitive in most configurations, with clear benefits in low-label DINOv2-S settings.
Ablation Study¶
The following component chain comes from main-text Table 5(b). Its endpoints of 35.61 and 63.20 correspond to the FSSG, DINOv2-S, 2% main results, and the appendix efficiency table explicitly identifies these values as FSSG references. Main-text Tables 3 and 5 do not restate the dataset in each caption. The value 62.55 alone must not be used to assign every ablation to FLSMD, where it also appears as a full-model main result.
| Config | mIoU (%) | Difference from EMA Baseline | Note |
|---|---|---|---|
| Labeled only M1 | 35.61 | Not applicable | No pseudo-label framework |
| EMA teacher–student M2 | 58.75 | 0.00 | Semi-supervised baseline |
| EMA + OPR M3 | 62.55 | +3.80 | Feature calibration only |
| EMA + ATC M4 | 61.90 | +3.15 | Threshold calibration only |
| EMA + OPR + ATC M5 | 63.20 | +4.45 | Joint calibration, 0.65 above the stronger single module |
Table 4's pseudo-label diagnostics show distinct effects from the two interventions. Only tail-related columns are retained below. Accuracy and filtering rates follow the authors' diagnostic conventions and are not validation tail mIoU; the paper does not fully formalize every diagnostic denominator here.
| Config | Tail Pseudo-Label Accuracy (%) | Tail Filtered Rate (%) | Correctly Kept GT Rate (%, Tail) |
|---|---|---|---|
| Baseline | 56.56 | 17.83 | 46.47 |
| + OPR | 87.80 | 11.24 | 77.93 |
| + ATC | 88.65 | 10.39 | 79.44 |
| FTC-Seg | 89.03 | 8.43 | 80.61 |
Key Findings¶
- In the internal OPR ablation, removing the residual reduces mIoU from 62.55 to 56.93 even as tail pseudo-label accuracy rises from 87.80 to 89.61. Better local pseudo-label diagnostics do not guarantee better overall segmentation; original features still carry spatial detail.
- The OPR placement table peaks at two stages with 62.55 mIoU, falling to 61.47 and 60.98 at three and four stages. More prototype modules are not necessarily better. The main text attributes this to semantic mismatch between shallow prototypes and the deep segmentation head.
- Appendix Table 7 uses three seeds that differ from the main tables but are matched across modules within each setting. Joint calibration improves over the stronger individual module by 1.96–3.85 points. After paired tests and Holm correction across 12 settings, ACDC 2% has \(p=0.0769\), above 0.05; significance does not hold for every setting.
- The controlled interaction is \(I=CL-OL-CB+OB\), where the four quantities are tail filtering rates under the respective conditions. Mild fog on ACDC gives −0.11 in the appendix, and SUIM is not monotonic with corruption severity. Super-additive exclusion has experimental support but is not universal at every noise severity.
- Appendix inference efficiency uses an RTX 4090, DINOv2-S + DPT, and \(518\times518\) inputs: UniMatch V2 reaches 122.62 FPS and FTC-Seg 120.93 FPS. OPR adds a small computational cost, while ATC adds no inference computation. The main-text millisecond table and appendix FPS table should not be treated as one measurement protocol.
Highlights & Insights¶
- Feature quality and access to supervision are separate bottlenecks. Raising confidence cannot replace a suitable admission rule, while lowering thresholds does not guarantee separable representations for recovered predictions; joint calibration targets this feedback chain.
- Learnable prototypes need not map one-to-one to classes. In the appendix's same-seed comparison, free-prototype OPR reaches 62.55 mIoU versus 54.87 for class-anchored OPR, supporting shared feature patterns rather than hard class centers.
- Labeled data can provide an external reference for pseudo-label filtering in addition to cross-entropy supervision. Monitoring learning state and prediction bias relative to priors is transferable, but prior reliability should first be checked under domain shift.
Limitations & Future Work¶
- The foreground set requires manual definition and is sensitive to mistakes. In the appendix's same-seed FSSG experiment, the correct set gives 62.55 mIoU, while adding background and omitting Diver gives 53.09. More robust soft gating is a possible extension, not a module in this paper.
- Labeled priors can be distorted by sparse sampling or genuine domain shift. Small-class denominators are particularly sensitive; EMA reduces batch variation but cannot eliminate systematic prior bias.
- Controlled balancing changes sample counts and composition, while original sonar and underwater images retain intrinsic noise. Synthetic corruption controls added degradation only. These experiments provide mechanism evidence without fully separating natural noise from all sample-selection effects.
- Main results use only three seeds, and some appendix gating, prototype, and Pascal VOC experiments use one seed. Joint gains are not statistically significant in every setting. More seeds, natural distribution shifts, and independent-domain tests would strengthen the conclusions.
- Details such as total-loss weights and threshold-boundary safeguards are incomplete. Performance and efficiency also involve differing table protocols; reproduction requires checking code rather than inferring unreported settings from the mechanism equations.
Related Work & Insights¶
- vs UniMatch / UniMatch V2: Weak-to-strong consistency and strong backbones establish the semi-supervised baseline. FTC-Seg changes both feature reconstruction and filtering within that paradigm, targeting noisy long-tailed data rather than introducing a new teacher-update scheme.
- vs FlexMatch / FreeMatch: Both address class-dependent thresholds. ATC distinguishes itself by combining labeled accuracy with labeled/unlabeled distribution ratios. The additional anchor can identify over-prediction but also introduces prior-mismatch risk.
- vs CReST / DASO: Distribution rebalancing targets class proportions, whereas this paper emphasizes improving representations and admission before filtering. Combining the approaches warrants testing for redundancy or complementarity rather than assuming additive benefits.
Rating¶
- Novelty: 4/5 — Makes the noise–long-tail filtering cycle explicit and provides interpretable interventions at two points.
- Experimental Thoroughness: 4/5 — Four domains, three label ratios, and varied ablations are substantial, but three-seed and several single-seed experiments limit evidence strength.
- Writing Quality: 3/5 — The mechanism is intuitive, but some dataset boundaries, implementation settings, and broad superiority claims need more careful qualification.
- Value: 4/5 — Useful for low-label segmentation in noisy environments, subject to task-specific validation of gating and prior assumptions.