Why Linear Probing Works: Non-Vacuous Generalization Bounds via Effective Dimension¶
Conference: ECCV 2026
Paper: ECCV Official
Full Text Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5959.txt
Area: Interpretability
Keywords: linear probing, generalization bounds, PAC-Bayes, effective dimension, vision foundation models
TL;DR¶
Resolving the theoretical discrepancy where linear probing requires exponentially fewer samples than classical bounds predict, this paper leverages steep covariance spectral decay to construct covariance-aligned data-dependent Gaussian priors, deriving the first non-vacuous PAC-Bayes generalization certificates governed by effective dimension \(d_{\text{eff}} \ll d\) and classification margin \(\gamma\).
Background & Motivation¶
In modern computer vision benchmarking, linear probing on frozen representations has emerged as the universal evaluation protocol. Across foundation models such as DINOv2 and CLIP, training a simple linear classifier over 768-dimensional features routinely achieves competitive accuracy on ImageNet with as few as 10 labeled examples per class. Yet from a statistical learning perspective, this empirical efficiency remains fundamentally paradoxical: standard multiclass sample complexity bounds for linear classifiers in \(\mathbb{R}^d\) prescribe \(\Omega(dk \log k / \epsilon^2)\) samples, which for \(d=768\) and \(k=1000\) exceeds practical sample regimes by orders of magnitude; standard dimension-dependent PAC-Bayes bounds of order \(\mathcal{O}(\sqrt{d/n})\) vacuously evaluate to values strictly greater than 1.
The root cause of this theoretical disconnect lies in the egalitarian treatment of dimensions within classical learning theory. In conventional VC-dimension and Rademacher complexity formulations, every dimension in ambient space is penalized equally, regardless of whether it encodes dominant class separability or isotropic noise. While empirical literature has long acknowledged that foundation models project inputs onto low-dimensional manifolds, existing formal analyses either impose restrictive conditional independence assumptions violated by visual pipelines, rely on augmentation graphs unavailable during downstream deployment, or offer heuristic transferability scores rather than rigorous probabilistic generalization certificates.
This paper tackles the challenge by interrogating the algebraic eigenspectrum of empirical feature covariances in frozen vision backbones: natural representations do not uniformly fill the ambient \(\mathbb{R}^d\) space, but rather exhibit steep power-law spectral decay where variance concentrates in an ultra-low-dimensional subspace. The core idea is to introduce the participation ratio effective dimension into the PAC-Bayes framework and align a data-dependent Gaussian prior with the unlabeled feature covariance, ensuring that hypothesis complexity is bounded by effective dimension \(d_{\text{eff}} \ll d\) and classification margin \(\gamma\) instead of ambient dimension \(d\), thereby establishing the first non-vacuous generalization certificates for linear probing.
Method¶
Overall Architecture¶
The framework establishes an analytical and computational pipeline for certifying the generalization of linear classifiers trained on frozen features. Given \(\ell_2\)-normalized feature representations extracted by a frozen encoder, the method first estimates the empirical feature covariance matrix over unlabeled data and computes its spectral participation ratio (effective dimension \(d_{\text{eff}}\)). Within the PAC-Bayes setting, it constructs a data-dependent Gaussian prior aligned with the principal axes of the covariance and pairs it with an eigenbasis water-filled posterior distribution to prevent multiclass and trace divergence. Combining the empirical classification margin loss with the resulting spectral complexity, the framework yields certifiable, non-vacuous risk bounds, which further support distribution-shift sensitivity theorems and a zero-label representation quality score.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Frozen Visual Feature Input<br/>ℓ2-normalized feature z"] --> B["Covariance Eigendecomposition & Effective Dimension"]
B --> C["Covariance-Aligned Spectral Data-Dependent Prior"]
C --> D["Eigenbasis Water-Filling Posterior & Complexity Decomposition"]
D --> E["Margin Conversion & Spectral PAC-Bayes Bound Derivation"]
E --> F["Distribution-Shift Analysis & Quality Score Derivation"]
Key Designs¶
1. Spectral Decay and Effective Dimension (Participation Ratio): Quantifying Subspace Concentration
To eliminate ambient dimensional penalties on redundant feature directions, the representation's true active degrees of freedom must be formalized. For \(\ell_2\)-normalized feature embeddings \(z_i \in \mathbb{R}^d\) (\(\|z_i\|=1\)), the empirical covariance is \(\hat{\Sigma} = \frac{1}{n} \sum_{i=1}^n z_i z_i^\top\), with eigendecomposition \(\hat{\Sigma} = \sum_{j=1}^d \lambda_j u_j u_j^\top\) where \(\lambda_1 \ge \cdots \ge \lambda_d \ge 0\).
Borrowing the participation ratio from condensed-matter physics for quantifying eigenstate localization, the effective dimension \(d_{\text{eff}}\) is defined as: $\(d_{\text{eff}}(\hat{\Sigma}) = \frac{(\operatorname{tr}(\hat{\Sigma}))^2}{\|\hat{\Sigma}\|_F^2} = \frac{(\sum_{j=1}^d \lambda_j)^2}{\sum_{j=1}^d \lambda_j^2}\)$ This quantity satisfies \(1 \le d_{\text{eff}} \le d\), collapsing to 1 when variance lies strictly on a 1D line and reaching \(d\) only for isotropic spherical noise. For modern discriminative foundation models exhibiting power-law decay \(\lambda_j \propto j^{-\alpha}\) with \(\alpha > 1\), \(d_{\text{eff}} = \Theta(1)\) becomes asymptotically independent of ambient dimension \(d\) (e.g., \(d_{\text{eff}} \approx 35\) for DINOv2-B despite \(d=768\)), creating a mathematically grounded substrate for compression.
2. Covariance-Aligned Spectral Data-Dependent Prior: Evading Ambient Dimensional KL Penalties
Conventional PAC-Bayes certificates utilize isotropic Gaussian priors \(\pi = \mathcal{N}(0, \sigma^2 I)\), which uniformly penalize divergence across all \(d\) coordinate axes, rendering bounds vacuous. The paper proposes a data-dependent Gaussian prior scaled by the unlabeled feature covariance: $\(\pi = \mathcal{N}\left(0, \frac{\sigma^2}{d_{\text{eff}}} \hat{\Sigma}\right)\)$ The core mechanism is spectral variance matching: directions with large eigenvalues \(\lambda_j\) receive broad prior variance, allowing posterior weights to align with major signals without incurring high KL divergence penalties. Conversely, near-zero noise directions receive narrow prior variance, strictly penalizing unnecessary posterior dispersion. Thus, the effective KL cost scales only with the \(d_{\text{eff}}\) active directions rather than the ambient dimension \(d\). Because the prior relies exclusively on unlabeled feature statistics, validity is preserved via an independent data-split argument.
3. Eigenbasis Water-Filling Posterior and Spectral Complexity Decomposition: Taming Multiclass Trace Explosion
When formulating the posterior distribution for the linear classifier \(\hat{W} = [\hat{w}_1, \dots, \hat{w}_k] \in \mathbb{R}^{d \times k}\), a naive isotropic perturbation forces the trace and log-determinant correction term \(T_{\text{lower}}\) into \(\Theta(d)\). To achieve numerical non-vacuity, the authors introduce a per-class water-filled posterior \(\rho_c = \mathcal{N}(\hat{w}_c, \operatorname{diag}(p_j^2))\).
By optimizing posterior variances \(p_j^2\) under a total margin constraint \(\sum_j \lambda_j p_j^2 \le \bar{\gamma}^2 / (8 \log(2kn))\), the spectral complexity decomposes cleanly into: $\(\Phi(\hat{W}, \hat{\Sigma}) = \frac{1}{d_{\text{eff}}} \sum_{j=1}^d \frac{\frac{1}{k} \|\hat{W}^\top u_j\|^2}{\lambda_j + \varepsilon_0}\)$ Because an \(\ell_2\)-regularized probe \(\hat{W}\) naturally aligns its projection energy with leading eigenvectors, the dominant terms possess large denominators \(\lambda_j\), suppressing \(\Phi(\hat{W}, \hat{\Sigma})\) and preventing high-dimensional multiclass inflation.
4. Spectral PAC-Bayes Bound Derivation and Margin Conversion: Producing Certifiable Guarantees
By integrating per-class PAC-Bayes aggregation with the empirical classification margin \(\bar{\gamma} = \frac{1}{n} \sum_{i=1}^n (w_{y_i}^\top z_i - \max_{c \ne y_i} w_c^\top z_i)\), whenever posterior variance satisfies \(\varepsilon \le \bar{\gamma}^2 / (8 \log(2kn))\), the stochastic classification error is strictly dominated by the empirical margin loss \(\hat{R}_{\bar{\gamma}}(\hat{W}) = \frac{1}{n} \sum_{i=1}^n \mathbf{1}[\gamma(z_i, y_i) < \bar{\gamma}]\). With probability at least \(1-\delta\), the population 0-1 risk satisfies: $\(R(\hat{W}) \le \hat{R}_{\bar{\gamma}}(\hat{W}) + \sqrt{\frac{d_{\text{eff}} \cdot \Phi(\hat{W}, \hat{\Sigma}) + T_{\text{lower}}(\varepsilon_{\text{post}}, \sigma^2) + \log\frac{2k\sqrt{n}}{\delta}}{2(n-1)}}\)$ Under spectral subspace alignment, this simplifies to the rate \(\mathcal{O}\left(\sqrt{\frac{d_{\text{eff}} \cdot k}{n \cdot \bar{\gamma}^2}}\right)\), confirming that effective dimension and margin supersede ambient dimension as the arbiters of sample efficiency.
5. Representation Quality Score and Distribution-Shift Formalism: Bridging Theory and Deployment
Building upon the bound's leading terms, the authors formulate a label-efficient representation score: $\(\mathcal{S}(f_\theta) = \frac{\bar{\gamma}^2}{d_{\text{eff}}}\)$ Here, \(d_{\text{eff}}\) requires zero labels and computes in seconds from unlabeled features, while \(\bar{\gamma}\) stabilizes with as few as 50 labeled instances. For target distribution shift, under the benign-shift condition where target covariance \(\Sigma_T\) has bounded variance in the source null space \(U_r^\perp\) (\(r = \lceil d_{\text{eff}}^S \rceil\)), the target effective dimension satisfies \(d_{\text{eff}}^T \ge d_{\text{eff}}^S + \Delta_{\text{shift}} \cdot c_T\), explaining precisely why and when linear probing certificates loosen under covariate shift.
Loss & Training¶
Linear probing optimizes the empirical risk under standard \(\ell_2\)-regularized cross-entropy loss: $\(\hat{W} = \arg\min_W \frac{1}{n} \sum_{i=1}^n \ell_{\text{CE}}(W^\top z_i, y_i) + \frac{\lambda_{\text{reg}}}{2} \|W\|_F^2\)$ For ImageNet-1K experiments, \(C = 1/\lambda_{\text{reg}} = 0.316\) is selected via held-out cross-validation. Hyperparameters \((\varepsilon, \sigma^2)\) are numerically determined via a joint \(50 \times 50\) grid search directly minimizing the PAC-Bayes upper bound expression.
Key Experimental Results¶
Main Results¶
On ImageNet-1K with the full training split (\(n=1,281,167\), \(k=1000\)) and confidence parameter \(\delta=0.05\), the spectral quantities and generalization bounds were evaluated across 12 vision foundation models spanning discriminative, predictive, language-supervised, supervised, and generative pretraining.
| Model | Pretraining Paradigm | Dimension \(d\) | Effective Dim \(d_{\text{eff}}\) | Margin \(\bar{\gamma}\) | Bound | kl-inv Bound | True Error | Bound / Error Ratio |
|---|---|---|---|---|---|---|---|---|
| DINOv2-g | Discriminative SSL | 1536 | 22 | 2.81 | 0.112 | 0.098 | 0.065 | 1.72× |
| DINOv2-L | Discriminative SSL | 1024 | 28 | 2.58 | 0.126 | 0.111 | 0.072 | 1.75× |
| DINOv2-B | Discriminative SSL | 768 | 35 | 2.30 | 0.114 | — | 0.081 | 1.40× |
| DINOv2-S | Discriminative SSL | 384 | 28 | 1.92 | 0.175 | 0.155 | 0.105 | 1.67× |
| I-JEPA-g | Predictive SSL | 1536 | 38 | 2.18 | 0.155 | 0.138 | 0.090 | 1.72× |
| I-JEPA-H | Predictive SSL | 1280 | 40 | 2.05 | 0.162 | 0.145 | 0.095 | 1.71× |
| CLIP-L | Language-Supervised | 768 | 50 | 1.97 | 0.198 | 0.178 | 0.120 | 1.65× |
| CLIP-B | Language-Supervised | 512 | 55 | 1.72 | 0.245 | 0.220 | 0.150 | 1.63× |
| Sup-ViT-L | Supervised | 1024 | 65 | 1.81 | 0.268 | 0.242 | 0.153 | 1.75× |
| Sup-ViT-B | Supervised | 768 | 72 | 1.48 | 0.315 | 0.285 | 0.188 | 1.68× |
| MAE-L | Generative SSL | 1024 | 150 | 0.83 | 0.782 | 0.710 | 0.280 | 2.79× |
| MAE-B | Generative SSL | 768 | 180 | 0.61 | > 1.00 | > 1.00 | 0.320 | Vacuous |
Ablation Study¶
On DINOv2-B / ImageNet-1K, the ablation isolates the impact of the spectral prior, dimensionality reduction heuristics, and \(\ell_2\) regularization on certificate tightness.
| Configuration | Effective Dim \(d_{\text{eff}}\) | Bound | Ratio | Status | Description |
|---|---|---|---|---|---|
| Full method (Theorem 3) | 35 | 0.114 | 1.40× | Non-vacuous | Covariance-aligned prior and water-filled posterior |
| Isotropic prior \(\mathcal{N}(0, \sigma^2 I)\) | 768 | 2.340 | — | Vacuous | Lacks spectral alignment; KL diverges with ambient \(d\) |
| Hard PCA to \(d_{\text{eff}}\) dims + standard bound | 35 | 0.172 | 2.12× | Non-vacuous | Omits soft spectral weighting and within-subspace structure |
| No \(\ell_2\) regularization (\(\lambda_{\text{reg}}=0\)) | 35 | 0.205 | 2.53× | Non-vacuous | Weights leak into noise directions, inflating \(\Phi\) |
| Top-50 eigenvectors only | 25 | 0.138 | 1.70× | Non-vacuous | Hard truncation impairs residual feature expressivity |
Key Findings¶
- Universal Non-Vacuity across Discriminative Models: All 10 discriminative, predictive, and language-supervised encoders achieve non-vacuous certificates (\(< 1\)). The DINOv2-B certificate attains a 1.40× bound-to-error ratio (bound 0.114 vs. held-out error 0.081), turning non-vacuous at sample size \(n \approx 8000\).
- Spectral Divergence of Generative Pretraining: Masked autoencoding (MAE) disperses representation variance uniformly across space (\(\alpha < 1\)), yielding high effective dimensions (\(d_{\text{eff}}=180\) for MAE-B) and smaller classification margins, which collapses the non-vacuity condition \(\Phi \ll k/\bar{\gamma}^2\).
- Label-Free Model Ranking: Effective dimension \(d_{\text{eff}}\) alone correlates with empirical accuracy at a Spearman rank coefficient of \(|\rho|=0.85\) (negative correlation: smaller \(d_{\text{eff}}\) indicates superior linear separability). The composite score \(\mathcal{S} = \bar{\gamma}^2/d_{\text{eff}}\) achieves \(\rho = 0.91\), outperforming PACTran (0.82), NLEEP (0.80), LogME (0.78), and H-score (0.76).
- Subspace Misalignment under Covariate Shift: Evaluating DINOv2-B across 7 diverse transfer targets shows that subspace misalignment \(\Delta_{\text{shift}}\), target effective dimension \(d_{\text{eff}}^T\), and the target generalization bound are strictly co-monotonic (\(\rho = +1.00\)).
Highlights & Insights¶
- Formulates the first mathematically rigorous bridge explaining the sample efficiency of linear probing in visual foundation models through spectral concentration rather than heuristic manifold arguments.
- Replaces naive isotropic priors with a covariance-aligned Gaussian prior and water-filled posterior, successfully neutralizing the multiclass \(k\)-fold and ambient-dimensional \(d\)-fold trace explosion.
- Derives a practical representation quality score \(\mathcal{S} = \bar{\gamma}^2 / d_{\text{eff}}\) where the dominant term \(d_{\text{eff}}\) is completely label-free, computing in under 30 seconds to provide sound, theoretically backed model selection.
Limitations & Future Work¶
- Reliance on Covariance Estimation: Precision requires reliable estimation of empirical covariance \(\hat{\Sigma}\); while small unlabeled samples yield non-vacuous bounds, small sample regimes overestimate \(d_{\text{eff}}\).
- Coarse-Grained Shift Dynamics: Theorem 9 establishes monotonic certificate loosening under distribution shift but does not capture fine-grained per-class margin structures (e.g., Flowers-102 retains 98.2% test accuracy despite elevated \(d_{\text{eff}}^T=42\) due to low class count \(k=102\)).
- Extension Beyond Linear Heads: The theoretical framework is tailored to frozen representations with linear heads; extending spectral generalization bounds to parameter-efficient fine-tuning (PEFT/LoRA) and non-linear probing architectures remains an open direction.
Related Work & Insights¶
- vs Saunshi et al. (ICML 2019) & HaoChen et al. (NeurIPS 2021): Prior generalization bounds for self-supervised learning depended on multi-view conditional independence or augmentation graphs unavailable at test time; this work evaluates arbitrary frozen backbones post-hoc via measurable empirical spectra.
- vs PACTran (ECCV 2022): PACTran formulated heuristic transferability metrics without obtaining formal PAC-Bayes generalization certificates or incorporating effective dimension and classification margin.
- vs Lotfi et al. (ICML 2024): While Lotfi et al. achieved non-vacuous bounds for large language models via weight compression, this paper delivers the first non-vacuous PAC-Bayes certificates for computer vision foundation models under linear probing, achieving a substantially tighter bound/error ratio of 1.40×.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First theoretical framework establishing non-vacuous PAC-Bayes bounds via spectral effective dimension for linear probing.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation spanning 12 models, 4 pretraining paradigms, and 7 benchmark datasets.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical derivation, well-contextualized theoretical boundaries, and clear intuition.
- Value: ⭐⭐⭐⭐⭐ Reconciles deep empirical practice with statistical learning theory while delivering a fast, label-free representation metric.