Skip to content

๐Ÿ“ Learning Theory

๐Ÿง  NeurIPS2026 ยท 9 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (1) ยท ๐Ÿ”ฌ ICLR2026 (293) ยท ๐Ÿงช ICML2026 (45) ยท ๐Ÿค– AAAI2026 (3) ยท ๐Ÿง  NeurIPS2025 (25) ยท ๐Ÿงช ICML2025 (16)

Adaptive inference for functionals of M-estimands

This paper extends inference for smooth functionals of nonparametric M-estimands to adaptive data with known sampling policies, combining one-step bias correction with per-round conditional variance stabilization, or self-normalization when variance converges to a random nonzero limit, to construct asymptotically valid confidence intervals; dynamic pricing simulations show undercoverage for ordinary GLMs and GLMs using inverse propensity weighting alone.

Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS

The paper turns significance testing for partial least squares (PLS) supervised subspaces into tests of held-out predictive correlation, supplies an NB-corrected fast approximation with an explicit empirical scope and a finite-sample valid permutation test under the omnibus independence null, and separates ordered inference on original components from interpretation of rotated coordinates.

Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models

COCOCO separately calibrates the concepts and task labels of neuro-symbolic concept-based models, then removes unsupported candidates through one deductionโ€“abduction intersection step, producing smaller, mutually consistent sets with explicit coverage-loss bounds and supporting adaptive size budgets through COCOCO*.

Cost-Aware Best-LLM Identification using Dueling Feedback

The paper formulates fixed-confidence identification of the best model under pairwise preferences as a heterogeneous-cost dueling bandit, allocates comparisons using rejection evidence per unit cost, tracks sampling targets, and stops through best-arm hypothesis testing, with an almost-sure asymptotic cost guarantee; however, the confidence-interval variant DCTAC costs less than the main algorithm DCTAS in the finite experiments.

Estimation of the Label-Noise Transition Matrix with Performance Guarantees via Selective Classification

The paper learns high-purity subsets through one-sided selective classification and estimates the transition matrix from the empirical frequencies of all noisy labels within those subsets, decomposing error into subset contamination, learning suboptimality, and accepted-sample fluctuations without requiring accurate pointwise class-posterior recovery.

Even Sharper Bounds for Transductive Learning and Its Applications

The paper proves Bernstein-type supremum concentration through a two-parameter entropy closure for the swap walk, then uses local complexity to analyze transductive empirical risk minimization, removing both earlier sample-imbalance restrictions and an extra logarithmic confidence factor under bounded-loss and appropriate localization conditions, with applications to realizable VC classification and the empirical kernel spectrum.

Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers

The paper establishes the minimax rate of nonparametric in-context learning on heterogeneous manifold mixtures whose complexity grows with sample size, and proves that a structure-informed transformer approximates a tangent local-polynomial estimator attaining that rate; optimality of a trained predictor still requires sufficient training tasks and a sufficiently small empirical-risk optimization gap.

The Price of Locality: Why Forward-Forward Underperforms Backpropagation?

This mechanism-analysis paper combines conditional local convergence theory, spectral analysis of representation kernels and error signals, and interventions on training groups and update spectra to explain why the Forward-Forward Algorithm (FFA) often trails back-propagation (BP); relaxing gradient isolation between CNN12 layers raises accuracy from 53.02% to 78.74%, still below BP's 85.67%.

Two-Fidelity Best-Action Identification for Stochastic Minimax Tree

Under a known fast-evaluation bias envelope and slow evaluations unbiased for true node minimax values, 2FFS uses endpoint certificates and recursive budgets to choose adaptively between expansion and sampling, achieving high synthetic-tree accuracy with substantially fewer node visits, while stopping and cost guarantees require additional regularity conditions.