Deep Noise Label Learning via Effective Rank Reduction¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: AI Safety
Keywords: Noisy Label Learning, Effective Rank Reduction, Transition Matrix, Nuclear Norm Regularization, Newton-Schulz Iteration
TL;DR¶
Addressing the failure of existing transition-matrix methods that overlook semantic class clustering and assume an unconstrained full rank, this paper introduces the Low-Effective-Rank Noisy-Label Learning (LENL) framework, which enforces nuclear-norm regularization on the noise transition matrix and develops a decomposition-free Newton–Schulz optimizer to eliminate SVD gradient instability while providing provable generalization guarantees.
Background & Motivation¶
In large-scale visual classification, imperfect annotations caused by human mistakes, ambiguous visual concepts, and automated web-crawling pipelines are ubiquitous. Deep neural networks, equipped with vast expressive capacity, readily memorize corrupted labels, leading to severe degradation in test-time generalization. Among noisy-label learning paradigms, transition-matrix-based forward correction explicitly models the class-dependent corruption process through a transition matrix \(T \in \mathbb{R}^{C \times C}\) (where \(T_{ij} = P(\tilde{Y}=j \mid Y=i)\)). Multiplying the network's predicted clean class posterior by \(T^\top\) aligns the predictions with the observed noisy distribution, providing a principled foundation for unbiased clean posterior recovery.
However, conventional transition-based approaches model \(T\) as a completely unconstrained full-rank matrix, optimizing all \(C^2\) entries independently. This assumption substantially mismatches real-world label noise: human confusion is rarely random or uniformly spread across the label space. Annotators systematically confuse visually or semantically similar classes (e.g., oak tree, pine tree, and willow tree). In datasets like CIFAR-100N, pairs of classes within the same superclass exhibit row cosine similarities exceeding 0.75 in \(T\), demonstrating strong row collinearity. Such semantic clustering causes the singular value spectrum of \(T\) to decay rapidly, concentrating spectral energy into a small subset of dominant modes and exhibiting a pronounced "reduced effective rank" structure. Unconstrained estimation wastes model capacity on noise-dominated tail modes, leading to statistical inefficiency and severe overfitting.
Directly constraining the rank during deep neural network training presents both theoretical and numerical hurdles: the discrete effective rank is non-differentiable, while its tightest convex surrogate—the nuclear norm—causes fatal numerical instability when optimized via standard differentiable SVD. Because backpropagating through SVD introduces terms proportional to \((\sigma_i^2 - \sigma_j^2)^{-1}\), the abundance of near-zero singular values induced by low-rank regularizers triggers zero-division and NaN gradients. Core idea: incorporate the low-effective-rank inductive bias into forward correction via nuclear-norm convex regularization to tighten generalization risk bounds, and devise a decomposition-free Newton–Schulz optimizer that evaluates the regularizer solely via matrix multiplications to completely eradicate SVD numerical instability.
Method¶
Overall Architecture¶
LENL establishes an end-to-end joint optimization framework for the classifier network parameters \(\theta\) and the noise transition matrix \(T\). Given an input image \(x\), the deep backbone produces a clean class probability distribution \(p_\theta(x) \in \Delta^{C-1}\). This distribution is linearly transformed by the transpose of the row-stochastic transition matrix into a forward-corrected noisy prediction \(T^\top p_\theta(x)\), which is trained against observed noisy labels via cross-entropy loss. Concurrently, nuclear-norm regularization \(\|T\|_*\) is imposed on the transition matrix to suppress spurious spectral modes. For gradient backpropagation through the regularizer, LENL incorporates a decomposition-free Newton–Schulz iterative approximation of the matrix square root, followed by row-wise softmax projection to preserve row stochasticity during SGD updates.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Noisy Sample (x, y~)"] --> B["Classifier Prediction p_theta(x)"]
B --> C["Forward Correction T^T p_theta(x)"]
C --> D["Forward Cross-Entropy Loss"]
D --> E["Nuclear Norm Spectral Regularization"]
E --> F["Decomposition-Free Newton-Schulz Optimizer"]
F --> G["Joint SGD & Row-Stochastic Projection"]
Key Designs¶
1. Transition Matrix Nuclear Norm Spectral Regularization: Inducing Low Effective Rank via Convex Relaxation
To overcome the overfitting of unconstrained full-rank estimation on tail noise modes, LENL introduces an explicit spectral capacity constraint. The effective rank is defined as the minimum number of singular components required to capture 90% of the total spectral energy: \(r_{\text{eff}} = \min \{r : E_r(T) \ge 0.90\}\). Because discrete rank metrics are intractable for gradient-based training, LENL employs the nuclear norm \(\|T\|_* = \sum_{i=1}^C \sigma_i(T)\)—the tightest convex relaxation of matrix rank—as a continuous surrogate. The overall objective is formulated as: $\(\min_{\theta, T} \mathcal{L}(\theta, T) := \mathbb{E}_{(x, \tilde{y}) \sim \mathcal{D}_{\text{noisy}}} \left[ -\tilde{y}^\top \log \left( T^\top p_\theta(x) \right) \right] + \lambda \|T\|_*\)$ where \(\lambda > 0\) balances noisy-label fitting and spectral complexity. To prevent trivial low-rank collapse (such as collapsing to rank-1 identical rows), \(T\) is constrained to be row-stochastic (\(\sum_j T_{ij} = 1, T_{ij} \ge 0\)). If \(T\) collapses into identical rows, \(T^\top p_\theta(x)\) becomes restricted to an uninformative low-dimensional convex hull insensitive to inputs, incurring severe penalties under the forward cross-entropy loss.
2. Theoretical Generalization Guarantees: Bounding True Clean Risk via Controlled Spectral Complexity
To theoretically substantiate the benefits of reduced effective rank, the paper establishes generalization error bounds demonstrating that both noisy empirical risk and true clean risk are directly bounded by the nuclear norm of \(T\). Let the hypothesis space be \(\mathcal{F} \subseteq \{f : \mathcal{X} \to \Delta^{C-1}\}\), and assume forward-corrected predictions satisfy \((T^\top f(x))_j \ge c > 0\) for all \(j \in \{1, \dots, C\}\). Given \(n\) i.i.d. noisy samples, the empirical forward-corrected risk is \(\widehat{R}_n^{\text{forw}}(f) = -\frac{1}{n} \sum_{i=1}^n \tilde{y}_i^\top \log (T^\top f(x_i))\). With probability at least \(1-\delta\), the population noisy risk satisfies: $\(R_{\text{noisy}}(f) \le \widehat{R}_n^{\text{forw}}(f) + \frac{2}{c} \|T\|_* \widehat{\mathfrak{R}}_n(\mathcal{F}) + M \sqrt{\frac{\log(2/\delta)}{2n}}\)$ where \(\widehat{\mathfrak{R}}_n(\mathcal{F})\) is the empirical Rademacher complexity of \(\mathcal{F}\), and \(M = \log(C/c)\). Furthermore, when \(\sigma_{\min}(T) > 0\), the expected risk under the underlying clean distribution \(\min_f R_{\text{clean}}(f)\) is similarly upper-bounded by terms scaling with \(\|T\|_*\). Minimizing the nuclear norm of \(T\) thus provides rigorous theoretical control over the hypothesis capacity of the composite model.
3. Decomposition-Free Newton–Schulz Optimizer: Eradicating SVD Gradient Singularities
Conventional differentiable SVD implementations exhibit fatal numerical instability when applied to matrices with reduced effective rank. Backpropagation through SVD requires computing \((\sigma_i^2 - \sigma_j^2)^{-1}\); as effective rank reduction clusters many singular values near zero, these denominators vanish, causing gradient explosion or NaN values. To resolve this, LENL leverages the identity \(\|T\|_* = \text{tr}((T T^\top)^{1/2})\) and approximates the matrix square root via the Newton–Schulz iteration using purely matrix multiplications. With scaling factor \(\alpha = 1 / \|T T^\top + \epsilon I\|_F\) and \(\epsilon = 10^{-6}\), initialization is defined as \(Y_0 = \alpha (T T^\top + \epsilon I)\) and \(Z_0 = \alpha I\), updated via coupled equations: $\(Y_{k+1} = \frac{1}{2} Y_k (3I - Z_k Y_k), \quad Z_{k+1} = \frac{1}{2} (3I - Z_k Y_k) Z_k\)$ Under the condition \(\|I - Y_0 Z_0\| < 1\), the sequence converges quadratically to \((T T^\top + \epsilon I)^{1/2}\). In practice, setting \(k=5\) iterations yields excellent numerical precision. Because the computational graph comprises only matrix multiplications and additions, gradients propagate smoothly without requiring eigenvalue decomposition, ensuring complete stability even on 1,000-class benchmarks.
Loss & Training¶
The framework is optimized end-to-end via joint mini-batch stochastic gradient descent. Backbone parameters \(\theta\) and transition matrix parameters \(T\) are updated simultaneously. The network is optimized using SGD with momentum 0.9. Row-stochasticity of \(T\) is strictly maintained via row-wise softmax projection after each update. Across all synthetic and real-world benchmarks, the regularization coefficient is fixed at \(\lambda = 5 \times 10^{-4}\), demonstrating remarkable hyperparameter stability without requiring task-specific tuning.
Key Experimental Results¶
Main Results¶
Evaluation was conducted on synthetic noise benchmarks (CIFAR-100, ImageNet-1K) and real-world noisy benchmarks (CIFAR-100N, Noisy Ostracods, Food-101N). Noise settings include uniform symmetric flipping (20%, 50%) and semantically correlated pair flipping (20%, 45%).
| Dataset | Noise Setting | Ours (LENL) | Prev. SOTA (PLM / ILDE) | Gain / Note |
|---|---|---|---|---|
| CIFAR-100 | Sym-20% | 70.42 ± 0.18% | 70.08 ± 0.31% (ILDE) | +0.34% |
| CIFAR-100 | Sym-50% | 61.35 ± 0.21% | 60.95 ± 0.42% (ILDE) | +0.40% |
| CIFAR-100 | Pair-20% | 73.21 ± 0.15% | 72.68 ± 0.24% (ILDE) | +0.53% |
| CIFAR-100 | Pair-45% | 62.88 ± 0.35% | 61.90 ± 1.82% (PLM) | +0.98% |
| ImageNet-1K | Sym-20% | 73.35 ± 0.08% | 66.69 ± 0.34% (Co-Teaching) | +6.66% (VolMinNet crashed) |
| ImageNet-1K | Sym-50% | 69.37 ± 0.56% | 63.40 ± 0.50% (Co-Teaching) | +5.97% |
| ImageNet-1K | Pair-20% | 72.59 ± 0.21% | 67.01 ± 1.28% (Co-Teaching) | +5.58% |
| ImageNet-1K | Pair-45% | 65.59 ± 0.89% | 64.08 ± 0.57% (Co-Teaching) | +1.51% |
| CIFAR-100N | Real-world | 61.51 ± 0.14% | 60.52 ± 0.32% (PLM) | +0.99% (+6.01% vs CE 55.50%) |
| Food-101N | Real-world | 83.30% | 80.89% (PLM) | +2.41% |
On the fine-grained real-world Noisy Ostracods benchmark (52 classes of marine fossil ostracods), multi-metric evaluation demonstrates the comprehensive superiority of LENL-NS:
| Model / Method | Accuracy (%) | Precision (%) | Recall (%) | F1-Score (%) | Execution Status |
|---|---|---|---|---|---|
| Standard CE | 95.98 | 88.50 | 77.80 | 79.51 | Completed |
| VolMinNet | - | - | - | - | Crashed (SVD numerical NaN) |
| Co-Teaching | 95.79 | 79.01 | 71.79 | 73.23 | Completed |
| PLM | 96.77 | 77.77 | 72.74 | 73.46 | Completed |
| ILDE | 96.23 | 77.49 | 72.08 | 73.05 | Completed |
| LENL-NS (Ours) | 97.20 | 87.21 | 77.93 | 79.75 | Best across metrics, stable |
Ablation Study¶
| Configuration / Dataset | Matrix Dimension \(C\) | Estimated eRank \(r_{\text{eff}}\) | Ratio \(r_{\text{eff}} / C\) | Empirical Observation / Note |
|---|---|---|---|---|
| CIFAR-100 (Pair-20%) | 100 | 86 | 0.86 | Moderate spectral concentration |
| CIFAR-100 (Pair-45%) | 100 | 69 | 0.69 | Heavy noise induces tighter effective rank |
| ImageNet-1K (Pair-45%) | 1000 | 686 | 0.686 | Pronounced spectral clustering at large scale |
| CIFAR-100N (Real-world) | 100 | 68 | 0.68 | Removing rank penalty drops accuracy by -4.6% |
| Noisy Ostracods (Real-world) | 78 | 39 | 0.50 | Extreme fine-grained confusion; SVD fails, NS succeeds |
| LENL (w/o Rank Regularization) | 100 | 100 (unconstrained) | 1.00 | Overfits tail noise modes; degraded generalization |
Key Findings¶
- Ubiquity of Low Effective Rank in Noise Matrices: Human annotations in natural datasets yield effective rank ratios between 0.50 and 0.68, confirming that semantic clustering systematically compresses the spectral dimensionality of real-world label noise.
- Newton–Schulz as an Enabler of Scalable Optimization: Traditional SVD-based methods (VolMinNet) fail completely on ImageNet-1K and real-world noisy datasets due to zero-division in backward gradients. In contrast, the Newton–Schulz iteration delivers robust quadratic convergence and superior training speeds.
- Persistent Regularization Effect Throughout Training: Tracking the effective rank of estimated transition matrices across epochs reveals that the low-effective-rank property is actively sustained from initial epochs to convergence, continuously safeguarding the backbone against label memorization.
Highlights & Insights¶
- From Full-Rank Overparameterization to Structured Inductive Bias: Unveils the intrinsic low-effective-rank geometry of label confusion matrices, breaking away from the conventional paradigm that treats all \(C^2\) entries as unconstrained independent parameters.
- Numerically Robust and Fully Differentiable Matrix Function Engine: Replaces fragile differentiable SVD routines with a 5-step matrix multiplication Newton–Schulz solver, providing a plug-and-play solution for nuclear-norm constraints in deep architectures.
- Rigorous Theoretical and Empirical Alignment: Connects the empirical observation of row collinearity directly to Rademacher generalization bounds, demonstrating how spectral complexity regularization directly contracts the clean risk upper bound.
Limitations & Future Work¶
- Reliance on In-Distribution Noise Assumption: The current framework assumes all noisy samples stem from in-distribution classes. Real-world open-set environments often feature out-of-distribution (OOD) outliers that lack semantic proximity to any class, potentially breaking the low-effective-rank structure. Future work should combine out-of-distribution filtering with low-rank constraints.
- Class-Conditional Rather Than Instance-Dependent Formulation: The transition matrix \(T\) is modeled at the population class level. Integrating sample-specific visual hardness into the low-effective-rank formulation represents a promising avenue for further investigation.
Related Work & Insights¶
- vs VolMinNet: While VolMinNet introduces volume minimization constraints, its dependence on exact SVD causes catastrophic numerical failures on large-scale datasets; LENL leverages nuclear-norm convex relaxation and the Newton–Schulz solver to achieve superior numerical stability and generalization.
- vs Sample Selection (Co-teaching, DivideMix): Sample selection heuristics require complex multi-network pipelines and lack formal clean-risk theoretical bounds; LENL achieves competitive or superior robustness through a single, elegant spectral regularization term.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First study revealing and theoretically analyzing the reduced effective-rank property of noise transition matrices, combined with an innovative Newton–Schulz optimizer.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation spanning 5 major synthetic and real-world benchmarks, including 1,000-class ImageNet.
- Writing Quality: ⭐⭐⭐⭐⭐ Highly coherent narrative bridging physical motivation, rigorous mathematical analysis, and robust algorithmic execution.
- Value: ⭐⭐⭐⭐⭐ Delivers a general, elegant, and computationally robust spectral regularization tool for weak supervision and noisy label learning.