๐ Optimization & Theory¶
๐ง NeurIPS2026 ยท 16 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (13) ยท ๐ท CVPR2026 (22) ยท ๐ฌ ICLR2026 (222) ยท ๐งช ICML2026 (88) ยท ๐ค AAAI2026 (21) ยท ๐ง NeurIPS2025 (126)
๐ฅ Top topics: LLM ร2
- Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients
-
The paper analyzes finite-horizon calibrated Adam without bias correction under generalized smoothness and second-moment ABC noise, shows that the confidence exponent \(\delta^{-1/2}\) cannot generally be removed, and separates expectation upper bounds for \(p<1\) from worst-case lower bounds for a specified algorithm family when \(1\leq p<2\).
- Amortized Optimal Transport from Sliced Potentials
-
The paper uses inexpensive one-dimensional Kantorovich potentials as features, predicts original-space potentials with shared linear coefficients, and reconstructs approximate transport plans; RA-OT and OA-OT reduce training costs and support variable numbers of atoms, but are not universally fastest at inference or best in generation quality.
- Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods
-
The paper separates surrogate geometry from utility sensitivity through the pullback Fisher tensor of the posterior map, then replaces TuRBO's kernel-lengthscale weights with regularized local Fisher diagonal weights, achieving competitive SE-kernel benchmark performance without a regret or global convergence guarantee for outer Bayesian optimization.
- Bidirectional Information Flow (BIF) - A Sample Efficient Hierarchical Gaussian Process for Bayesian Optimization
-
BIF constructs a soft parent prior from child Gaussian processes' acquisition maps and allocates real parent responses to children using uncertainty-aware weights for continual learning, improving reconstruction and learning trajectories in low-budget composite tasks; this feedback comprises biased pseudo-responses, not true subtask decomposition.
- Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models
-
The paper decomposes the noncommutative residue of two training updates into a signed token readout, localizing and intervening on training-order differences under local SGD conditions; targeted interventions close the held-out loss gap by a median 32.0% on Qwen-3-4B, while paired-endpoint order assignment reaches 66/72 = 91.7%.
- Cumulative-Goodness Free-Riding in Forward-Forward Networks: Real, Repairable, but Not Accuracy-Dominant
-
The paper proves that cumulative goodness exactly attenuates deep blocks' local discrimination gradients on examples already separated upstream, and repairs block health through history removal, hardness gating, and gradient compensation, but finds no accuracy-dominant benefit from these repairs; MGC even reduces CIFAR-100 Stage-1 single-crop accuracy by 1.05 percentage points.
- Dynamic Regret in Online Convex Optimization with Indicator Switching Costs
-
The paper aggregates distributions from randomized lazy FTRL learners restarted at dyadic scales through a movement-aware meta-learner, uses maximal coupling to control actual action changes, and obtains near-optimal expected dynamic regret for piecewise-constant comparators alongside a separate guarantee for comparators with small path length.
- Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models
-
EoupCT uses a frozen pretrained model and learnable soft prompts to construct differentiable knowledge proxies vulnerable to new-task updates, then combines distillation with gradient projection to protect historical tasks and general capabilities; experiments on SuperNI/MMLU subsets across six models improve retention, but its โfirst-order-onlyโ and โabsolute zero forgettingโ claims require qualification.
- Finite-Sample Performance of Gradient Descent in Logistic Regression with Gaussian Design
-
Under well-specified Gaussian logistic regression, the paper combines population curvature analysis with approximate invertibility of the empirical gradient to prove linear convergence of gradient descent to a statistical-error neighborhood, and obtains a sharper high-dimensional error bound by estimating direction and norm separately; large-step acceleration is only local, and the stated near-optimal regime contains a condition discrepancy that must be retained.
- FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models
-
FluxLite jointly adjusts sparse jump rates and compensating weights without retraining a discrete diffusion model, preserving the target marginal-distribution path through graph divergence while mitigating particle-weight degeneracy with the local HEU rule or a small nonnegative quadratic program, D-VCG.
- Non-Linear Pricing Restores Tractability for a Data Seller
-
With known linear buyer valuations and budgets and separate pricing for each dataset, allowing nonlinear prices turns the previously APX-hard optimal linear-pricing problem into a polynomial-time LP and guarantees an optimum with at most as many total kinks as buyers; instances constructed from California Housing also exhibit revenue ratios of at most 1.1 and at most 10 total kinks.
- Online Learning via Learned Latent Bayesian Tracking
-
The paper learns a low-dimensional parameter-generating space and dynamical prior offline, performs one latent extended Kalman filtering update per labeled sample online, and reconstructs prediction parameters, improving adaptation under limited supervision and computational constraints in image classification and time-varying wireless reception.
- QuanVI: Score-based Variational Inference via Quantum Maximally Mixed States
-
QuanVI replaces EigenVI's single-eigenvector solution with a maximally mixed state on a low-energy subspace and compresses the density operator through local quantum tensor networks, scaling to 100-dimensional chain-structured synthetic targets while remaining limited by local-window capacity and cost.
- Simple Extensions of Single-Objective Acquisition Functions and Hedge Strategies for Multi-Objective Bayesian Optimization
-
The paper generates candidates through Pareto search over a vector of single-objective acquisitions, selects queries using posterior-mean hypervolume, and adaptively chooses acquisitions with MO-Hedge; results are competitive on nine benchmarks, but the runtime advantage is mainly evident in batch mode rather than uniformly across settings.
- SimpleEvol: An Agent-Loop Framework for LLM-Driven Automated Heuristic Design with Minimal Human Priors
-
SimpleEvol replaces elaborate population and evolutionary-operator orchestration with one โgenerate codeโexecute and evaluateโcompress memoryโ search trajectory, obtaining the highest TSP/CVRP intelligence-conversion regression slopes across ten LLMs and lower test gaps at every evaluated size with GPT-5-mini, without establishing that fewer priors necessarily improve performance.
- When One Leak Pays Forever: Context Binding and the Price of Deterring Collusion
-
The paper characterizes the exact deterrence threshold through the future value reachable by one terminal leak: genuine context isolation requires exposure covering only one round, whereas long-lived reuse can multiply the requirement by \(T\) over an undiscounted finite horizon.