SCDL: Synergistic Confidence-Dispersion Learning for Semi-Supervised Video Polyp Segmentation¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/Yuanqin-He/SCDL
Area: Medical Imaging / Video Understanding
Keywords: Semi-Supervised Learning, Video Polyp Segmentation, Confidence-Dispersion Synergy, Spectral Convex Optimization Pseudo-Label Selection, Contextual Consistency Perturbation
TL;DR¶
To tackle low polyp-mucosa contrast and error propagation from overconfident pseudo-labels in colonoscopy videos, SCDL introduces prototype-guided adaptive refinement, spectral convex optimization over confidence-dispersion space, and contextual consistency perturbation, creating a closed-loop semi-supervised framework that establishes new SOTA results across all label ratios on SUN-SEG.
Background & Motivation¶
Early detection and resection of colorectal adenomatous polyps is the cornerstone clinical intervention for preventing colorectal cancer (CRC). Although recent fully-supervised deep learning methods have achieved substantial success on static colonoscopy images, they inherently neglect inter-frame temporal dependencies and frequently suffer from temporal instability when deployed on continuous video streams. Video polyp segmentation (VPS) models have thus emerged to exploit temporal coherence. Nevertheless, dense pixel-level video annotation is notoriously labor-intensive and expensive, resulting in a severe shortage of large-scale labeled video datasets in clinical practice.
Semi-supervised video polyp segmentation (SSVPS) aims to alleviate this annotation bottleneck by leveraging abundant unlabeled video frames alongside a modest fraction of annotated sequences. However, conventional semi-supervised semantic segmentation methods operate on the fundamental assumption that high prediction confidence reliably indicates high pseudo-label quality. This premise severely breaks down in endoscopic video domains. Endoscopic videos exhibit unique physical challenges, including inherently low visual contrast between lesions and surrounding mucosal tissues, drastic inter-frame deformations driven by rapid camera motion, and recurring sequences of degraded frames plagued by specular reflections or motion blur. Under such conditions, deep networks become blindly overconfident, generating high-probability yet completely erroneous pseudo-labels along ambiguous boundaries. Worse still, temporal propagation mechanisms iteratively compound these noisy errors over time, while standard fixed confidence thresholding chronically neglects uncertain pixels, inducing severe regional supervision bias.
Overcoming these intertwined challenges requires more than hand-crafted heuristic thresholds or isolated temporal modules; it demands a closed-loop formulation connecting reliability quantification with hard-region representation learning. Core idea: formulate pseudo-label selection as a spectral convex optimization separation problem over a joint confidence-residual dispersion space, coupled with Prototype-guided Adaptive Refinement (PAR) and Contextual Consistency Perturbation (CCP) that masks high-confidence areas to force semantic reconstruction from ambiguous boundaries, establishing a synergistic, self-reinforcing semi-supervised learning loop.
Method¶
Overall Architecture¶
SCDL unifies feature refinement, pseudo-label selection, and perturbation-driven self-supervision into an interdependent closed loop. Given a small set of labeled frames alongside large unlabeled video sequences, the framework implements a weak-to-strong consistency baseline integrated with three synergistic components: Prototype-guided Adaptive Refinement (PAR), Spectral Convex Optimization Separation (SCOS), and Contextual Consistency Perturbation (CCP). First, PAR extracts high-resolution spatial features and fuses them with historical prototype memory to produce discriminative spatio-temporal representations. Next, SCOS maps pixel-wise maximum probabilities and residual dispersions into a two-dimensional feature space, solving a convex optimization problem via spectral relaxation to adaptively produce reliability indicators and smooth soft-weight maps without manual thresholds. Finally, CCP leverages these reliable indicators to randomly mask high-confidence spatial patches, compelling the network to reconstruct global semantics from the remaining ambiguous regions and temporal priors.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Video Sequences<br/>Labeled + Unlabeled Frames"] --> PAR["Prototype-guided Adaptive Refinement (PAR)<br/>IOF Shallow-Deep Fusion + EOF Cross-Attention + OPG Clustering"]
PAR --> SCOS["Spectral Convex Optimization Separation (SCOS)<br/>Joint Confidence-Dispersion Spectral Relaxation for Adaptive Masking"]
SCOS -->|Soft Weighting Map| LossA["Standard Weak-to-Strong Consistency Loss<br/>Weighted Suppression of Ambiguous Noise"]
SCOS -->|Binary Reliable Region Indicator| CCP["Contextual Consistency Perturbation (CCP)<br/>Masking High-Confidence Patches to Drive Ambiguous Reconstruction"]
CCP --> LossM["Perturbed Reconstruction Consistency Loss<br/>Semantic Inference from Surrounding Context"]
LossA & LossM --> Opt["Closed-Loop Joint Optimization<br/>Backward Gradients Continuously Enhance PAR Discrimination"]
Key Designs¶
1. Prototype-guided Adaptive Refinement (PAR): Decoupling Spatio-Temporal Priors to Dispel Low-Contrast Ambiguity Low tissue contrast and abrupt camera jitter severely compromise single-frame representations. PAR resolves these ambiguities via a three-tier progressive fusion mechanism. First, Implicit Object-aware Fusion (IOF) addresses the semantic discrepancy between shallow detail textures and deep abstractions. A FIFO memory bank of capacity \(T=8\) retains historical high-level encoder representations; cross-attention is executed using current deep features as Query and memory features as Keys and Values to produce temporal context \(F_t^{mem}\), which is then compressed via convolutional layers and point-wise summed with multi-scale shallow features. Second, Explicit Object-aware Fusion (EOF) concatenates this fused representation with the prior frame context \(R_{t-1}^c\), followed by cross-attention against explicit historical polyp prototypes \(\mathcal{P}_t\), leveraging compact object-level priors to regularize spatial deformations. Third, Object-aware Prototype Generation (OPG) samples \(k\) seed points within the predicted polyp mask using Farthest Point Sampling (FPS), runs a single \(k\)-means clustering step, and calculates cluster spatial feature means as updated prototypes to refresh the temporal memory bank.
2. Residual Dispersion Theory and Spectral Convex Optimization Separation (SCOS): Heuristic-Free Pseudo-Label Selection Conventional pseudo-label filtering relies exclusively on maximum class probability \(p_n(k^\star)\), which frequently misclassifies overconfident predictions in low-contrast boundaries. Rooted in the entropy minimization principle, the paper provides a formal theoretical decomposition: for an ideal unimodal target distribution, second-order Taylor expansion of the cross-entropy yields: $$ \mathrm{CE}(p_n, \tilde{q}) \approx -\log p_n(k^\star) + \varepsilon \frac{(C-1)^3}{2(1-p_n(k^\star))^2} v_n $$ where residual dispersion is rigorously defined as \(v_n \triangleq -\frac{1}{C-1} \sum_{k \neq k^\star} \Delta_k^2\) (\(\Delta_k\) measuring probability deviation from non-maximum mean). As maximum confidence approaches unity (\(p_n(k^\star) \to 1\)), the multiplier preceding \(v_n\) grows rapidly, establishing residual dispersion as a sensitive indicator for separating truly accurate predictions from overconfident errors. SCOS constructs a two-dimensional feature attribute \(h_n = [p_n(k^\star), v_n]^T\) for each pixel and casts pseudo-label partitioning as maximizing spectral separability: $$ \max_{S} \mathrm{Tr}(S^T \Phi^T \Phi S), \quad \text{s.t. } S \in {0, 1}^{HW \times 2} $$ Solving this via spectral relaxation through the Ky Fan theorem circumvents heuristic threshold search. A Gaussian kernel mapping then generates sample-adaptive soft weights \(M_t\), assigning full weight to high-confidence/high-dispersion pixels while smoothly attenuating ambiguous candidates.
3. Contextual Consistency Perturbation (CCP): Forcing Semantic Reconstruction via Targeted Masking Standard semi-supervised pipelines that supervise only confident pixels inadvertently teach models to ignore challenging boundary regions, entrenching severe regional supervision bias. CCP turns this dynamic around: utilizing the binary reliable indicator \(M_t = 1\), it applies patch-wise masking with patch size \(s=32\) and masking ratio \(\theta=0.5\) exclusively over confident regions to generate perturbed input \(\tilde{x}_t^u = x_t^u \odot R_t\). Because the clear polyp core is masked out, the model cannot rely on trivial cues and is forced to infer the missing semantics by aggregating surrounding low-contrast contextual margins and PAR's temporal memory priors. This mechanism actively channels mutual information from ambiguous boundaries into the supervision stream, bridging the semantic gap between easy and hard regions.
Loss & Training¶
The overall training objective couples supervised cross-entropy with dual-stream unsupervised consistency objectives: $$ \mathcal{L} = \mathcal{L}S + \mathcal{L}_U = \mathcal{L}_S + \lambda_1 \mathcal{L}_u^a + \lambda_2 \mathcal{L}_u^m $$ where \(\mathcal{L}_S\) denotes standard supervised cross-entropy on labeled frames, \(\mathcal{L}_u^a\) is the standard weak-to-strong consistency loss, and \(\mathcal{L}_u^m\) evaluates reconstruction consistency under CCP masking. Both unsupervised losses are pixel-wise modulated by the SCOS adaptive weight map \(M_t\): $$ \mathcal{L}_u^a = \frac{1}{B_U} \sum}^{B_U} M_t \odot \mathrm{CE}\big(\hat{yt^u, \mathcal{F}(A_s(A\omega(x_t^u)))\big) $$ $$ \mathcal{L}u^m = \frac{1}{B_U} \sum_t^u)\big) $$ The loss weights are balanced at }^{B_U} M_t \odot \mathrm{CE}\big(\hat{y}_t^u, \mathcal{F}(\tilde{x\(\lambda_1 = 0.5, \lambda_2 = 0.5\). Training is conducted with SGD optimization using a batch size of 8, weight decay of \(1\times 10^{-4}\), and initial backbone learning rate of \(1\times 10^{-3}\) (with the decoder rate set \(10\times\) higher). Inputs are randomly cropped to \(352 \times 352\) over 60 epochs. Crucially, SCOS and CCP serve strictly as training-phase supervision engines and are completely discarded at test time, incurring zero inference latency.
Key Experimental Results¶
Main Results¶
On the comprehensive SUN-SEG benchmark (stratified into Easy and Hard evaluation subsets encompassing seen and unseen video sequences), SCDL outperforms competing semi-supervised methods spanning Natural Image Segmentation (NIS), Image Polyp Segmentation (IPS), Natural Video Segmentation (NVS), and Video Polyp Segmentation (VPS) across all labeled data partitions:
| Model | Backbone | Paradigm | 1/16 Labels (Easy) | 1/16 Labels (Hard) | 1/8 Labels (Hard) | 1/4 Labels (Hard) | 1/2 Labels (Hard) |
|---|---|---|---|---|---|---|---|
| Corrmatch | Res2Net-50 | NIS | \(S_\alpha\): 80.92 / Dice: 71.25 | \(S_\alpha\): 79.85 / Dice: 69.98 | \(S_\alpha\): 80.82 / Dice: 72.45 | \(S_\alpha\): 84.52 / Dice: 76.02 | \(S_\alpha\): 84.95 / Dice: 77.42 |
| ACL-Net | Res2Net-50 | IPS | \(S_\alpha\): 81.69 / Dice: 73.43 | \(S_\alpha\): 81.38 / Dice: 71.76 | \(S_\alpha\): 81.53 / Dice: 73.21 | \(S_\alpha\): 86.04 / Dice: 77.32 | \(S_\alpha\): 86.08 / Dice: 78.93 |
| PedSemiSeg | Res2Net-50 | IPS | \(S_\alpha\): 81.42 / Dice: 71.79 | \(S_\alpha\): 80.73 / Dice: 70.84 | \(S_\alpha\): 81.93 / Dice: 73.74 | \(S_\alpha\): 85.24 / Dice: 76.74 | \(S_\alpha\): 85.63 / Dice: 78.35 |
| IFR | Res2Net-50 | NVS | \(S_\alpha\): 81.05 / Dice: 71.38 | \(S_\alpha\): 80.12 / Dice: 70.25 | \(S_\alpha\): 81.05 / Dice: 72.68 | \(S_\alpha\): 84.73 / Dice: 76.23 | \(S_\alpha\): 85.18 / Dice: 77.65 |
| TDC | Res2Net-50 | NVS | \(S_\alpha\): 81.39 / Dice: 71.72 | \(S_\alpha\): 80.51 / Dice: 70.64 | \(S_\alpha\): 81.29 / Dice: 72.92 | \(S_\alpha\): 85.09 / Dice: 76.59 | \(S_\alpha\): 85.43 / Dice: 77.90 |
| TCCNet | Res2Net-50 | VPS | \(S_\alpha\): 80.63 / Dice: 74.18 | \(S_\alpha\): 79.87 / Dice: 73.42 | \(S_\alpha\): 81.43 / Dice: 76.62 | \(S_\alpha\): 84.78 / Dice: 79.02 | \(S_\alpha\): 85.64 / Dice: 80.83 |
| SSTAN | Res2Net-50 | VPS | \(S_\alpha\): 81.93 / Dice: 75.27 | \(S_\alpha\): 81.37 / Dice: 74.51 | \(S_\alpha\): 82.89 / Dice: 77.70 | \(S_\alpha\): 86.18 / Dice: 80.14 | \(S_\alpha\): 87.16 / Dice: 81.96 |
| PSDNet | PVTv2-b2 | VPS | \(S_\alpha\): 87.09 / Dice: 83.07 | \(S_\alpha\): 86.76 / Dice: 81.29 | \(S_\alpha\): 87.52 / Dice: 82.54 | \(S_\alpha\): 87.64 / Dice: 83.36 | \(S_\alpha\): 88.10 / Dice: 83.77 |
| SCDL (Ours) | Res2Net-50 | VPS | 87.52 / 83.48 | 87.17 / 81.66 | 88.05 / 82.98 | 88.13 / 83.94 | 88.62 / 84.28 |
| SCDL (Ours) | PVTv2-b2 | VPS | 89.37 / 85.91 | 88.59 / 83.76 | 89.04 / 84.62 | 89.31 / 85.44 | 89.51 / 86.09 |
Ablation Study¶
On the SUN-SEG-Hard split under the 1/2 labeled data partition, systematic ablation isolates module contributions alongside computational overhead:
| Config | Dice (%) | GFLOPs | Params (M) | Notes |
|---|---|---|---|---|
| Baseline | 80.21 | 92.82 | 41.38 | Standard fixed-threshold self-training |
| w/ PAR | 82.36 | 107.29 (+14.47) | 42.43 (+1.05) | Adds prototype-guided spatio-temporal refinement |
| w/ SCOS + CCP (w/o PAR) | 83.24 | 94.90 (+2.08) | 41.38 (+0) | Pseudo-label filtering and perturbation only |
| w/ PAR + SCOS (w/o CCP) | 83.72 | 107.29 (+14.47) | 42.43 (+1.05) | Reliable filtering without inverse perturbation |
| SCDL (Full Model) | 84.28 | 109.37 (+16.55) | 42.43 (+1.05) | Complete synergistic closed-loop framework |
Key Findings¶
- Training-Only Supervision Synergy: SCOS and CCP operate entirely during training and impose zero runtime cost during inference. PAR adds only 1.05M parameters and 14.47 GFLOPs, making SCDL highly practical for clinical real-time endoscopy pipelines.
- Robustness in Annotation-Scarce Settings: Under the extreme 1/16 labeled setting, SCDL achieves 81.66% Dice on SUN-SEG-Hard, outpacing the standard baseline by +11.68% and demonstrating that filtering overconfident errors is paramount when supervision is scarce.
- Superior Zero-Shot Cross-Dataset Generalization: Evaluated on the completely unseen CVC-ClinicDB dataset without any fine-tuning, SCDL attains 94.12% \(S_\alpha\) and 91.53% Dice, comfortably outperforming the previous state of the art PSDNet (91.84% / 89.67%) and indicating minimal domain overfitting.
Highlights & Insights¶
- Residual Dispersion as a Countermeasure to Overconfidence: The formal mathematical derivation proves that non-maximum probability dispersion provides an indispensable orthogonal indicator to raw confidence, resolving pseudo-label overconfidence in ambiguous domains.
- Counter-Intuitive Confident Masking: Instead of conventional random cutouts, CCP selectively masks highly reliable predictions, transforming what would be ignored ambiguous boundaries into active supervisory drivers.
- Compact Prototype Geometry: Restricting prototype extraction to \(k=5\) cosine-clustered centers efficiently captures complex morphology while keeping cross-attention overhead minimal.
Limitations & Future Work¶
- Vulnerability to Specular Flares and Diminutive Polyps: Intense endoscopic light reflections can create local saturated whiteouts devoid of texture, while ultra-small flat polyps risk feature entanglement during spatial pooling downsampling.
- Future Directions: Exploring reflection-invariant decomposition techniques and multi-scale sub-pixel attention mechanisms to preserve tiny flat lesion signatures across video sequences.
Related Work & Insights¶
- vs TCCNet / SSTAN: While prior semi-supervised VPS works rely on spatio-temporal feature attention to propagate representations across frames, they lack explicit mechanisms to prevent noisy pseudo-labels from snowballing over time; SCDL stops error propagation at the source through convex optimization.
- vs UniMatch / SoftMatch: Traditional semi-supervised semantic segmentation methods rely on static heuristic cutoffs or restrictive Gaussian distribution assumptions; SCDL derives sample-adaptive filtering boundaries mathematically via spectral graph cuts.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Rigorous residual dispersion formulation coupled with spectral convex optimization for threshold-free selection]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive label fractions, comprehensive module ablations, computational auditing, and zero-shot cross-dataset evaluation]
- Writing Quality: ⭐⭐⭐⭐⭐ [Mathematically self-contained with well-articulated closed-loop dynamics]
- Value: ⭐⭐⭐⭐⭐ [Directly cuts clinical annotation costs for endoscopic computer-assisted intervention systems]