Adaptive inference for functionals of M-estimands¶
Conference: NeurIPS2026 (accepted status self-reported by the authors in the paper list)
arXiv: 2609.39274
Area: Adaptive Statistical Inference / Statistical Learning Theory
Keywords: M-estimands, Neyman orthogonality, influence functions, conditional variance stabilization, self-normalization
TL;DR¶
This paper extends inference for smooth functionals of nonparametric M-estimands to adaptive data with known sampling policies, combining one-step bias correction with per-round conditional variance stabilization, or self-normalization when variance converges to a random nonzero limit, to construct asymptotically valid confidence intervals; dynamic pricing simulations show undercoverage for ordinary GLMs and GLMs using inverse propensity weighting alone.
Background & Motivation¶
Online experiments and contextual bandits change subsequent actions in response to previous rewards, so their final datasets are not independent and identically distributed samples from a fixed law. Even after weighting removes first-order bias, an estimator's variance can remain trajectory dependent: early observations favor different actions, changing the information available later. When optimal actions tie or have similar rewards, the sampling policy need not converge to a deterministic policy independent of the trajectory, undermining classical normal approximations and sandwich variance estimates. Here, inference means statistical inference, not faster model execution.
Existing methods separately handle arm means, average treatment effects, linear coefficients, or particular M-estimators through importance weighting, online debiasing, or influence function corrections. They often couple the target, working model, and variance estimation procedure. Correctly specified linear models simplify analysis but do not cover flexible nuisance functions and misspecified working models; methods allowing misspecification may instead require deterministic policy limits or separate estimation of conditional score moments. This paper seeks a reusable statistical interface: when the target is a smooth functional of a risk minimizer under a specified evaluation law, derive its correction from loss curvature and the functional derivative, then address adaptive variance separately.
This interface has clear boundaries. It requires known sampling probabilities, an identifiable and sufficiently smooth target, orthogonality, nuisance convergence rates, and exploration conditions; it does not guarantee inference for arbitrary black-box data collection. Core idea: use a Riesz representation in the Hessian metric to construct an orthogonal one-step correction, express target error through martingale increments, and choose per-round stabilization or self-normalization according to whether conditional variance has a stable random limit.
Method¶
Overall Architecture¶
The inputs are sequential contexts, actions, rewards, and the sampling policy used in each round. The outputs are a point estimate and an asymptotic confidence interval for a functional under a fixed evaluation law, not a new policy learning algorithm. The method has a shared target definition and one-step correction followed by two distinct variance-handling routes; its contributions are statistical constructions and limit theorems, not a neural network architecture.
Write the observation as \(Z_t=(C_t,A_t,Y_t)\) and the past information as \(\mathcal F_{t-1}\). The sampling policy \(\pi_t\) may depend on the past, but not on the current unobserved reward. Assumption 1 requires independent and identically distributed contexts independent of the past, and a time-invariant reward distribution conditional on context and action; dependence over time enters only through the policy. The results therefore do not directly guarantee inference for arbitrary stateful reinforcement learning trajectories, nonstationary environments, or drifting rewards.
The evaluation policy \(\pi_e\) is fixed and independent of history. It shares the context distribution and reward mechanism with the actual sampling law but changes action allocation. Denote its law by \(P_e\). The loss \(\ell\), nuisance function \(\eta_{P_e}\), and smooth map \(m\) define the target:
An M-estimand is the population risk-minimizing target, not the M-estimator computed from a sample. Under a misspecified linear working model, it can remain the best projection coefficient under the evaluation law, but this coefficient is not automatically a true structural parameter. Validity under misspecification means covering this explicitly defined risk-minimizing target, not dispensing with model, identification, or convergence conditions.
In each round, past data estimate the target function, nuisance function, and Riesz representer before a one-step corrected value is computed on the new observation. Known policies provide the importance weight \(w_t=\pi_e(A_t\mid C_t)/\pi_t(A_t\mid C_t)\). Assumption 2 also requires a common dominating measure and absolute continuity of the evaluation policy with respect to every sampling policy: actions used by the evaluation policy must have sampling support. Knowing probabilities cannot compensate for zero exploration.
After the shared correction, the first route reweights historical increments to estimate conditional standard deviation under the current policy, determines the stabilization weight before the new observation, and accumulates corrected values. The second route does not estimate each round's conditional variance; it cancels the random scale using observed squared increments, but its theorem requires average conditional variance to converge to a finite, almost surely nonzero random variable.
Key Designs¶
1. Orthogonal Riesz correction: convert target bias into a removable first-order term
Directly plugging an estimated regression function or parameter into the target functional typically leaves an error of the same order as the fitting error, preventing inference at the square-root sample-size scale. The paper defines an inner product through the population loss Hessian and finds a Riesz representer \(\alpha_{P_e}\) such that the functional derivative in any direction equals its Hessian inner product with that direction. This converts the direction most relevant to the target into one measurable by the loss gradient:
Assumption 3 requires an existing unique risk minimizer, twice Fréchet differentiability of the loss in the target parameter, once differentiability of the functional, population first-order conditions, and a positive lower bound for the Hessian. This lower bound ensures local identification; uninformative directions or nearly degenerate models cannot use the representation without further treatment. Assumption 4 additionally bounds local derivatives and their differences through an envelope with finite fourth moment and corresponding local Lipschitz bounds.
Neyman orthogonality requires the population first-order target derivative of the loss to have zero first-order response to nuisance perturbations. Nuisance estimation is not irrelevant; its errors instead enter second-order or product terms. Theorem 1 leaves products of target and Riesz errors, squared target error, and nuisance-related cross and squared errors. For a linear functional, quadratic loss, and no nuisance parameter, Remark 1 permits a weaker target–Riesz error product condition in place of the generic rates.
After defining the round-specific corrected value \(F_t\), the oracle influence function under the evaluation law is the corrected value minus the target. The mechanism of Equations 3 and 5 can be summarized as:
Subtracting the loss derivative along the Riesz direction cancels the leading plug-in bias. Importance weighting changes the conditional expectation under the actual policy into an expectation under the evaluation law, so the oracle increment \(w_t\varphi_{P_e}(Z_t)\) has conditional mean zero. These address different problems: one-step correction removes fitting bias, whereas importance weighting handles changing action distributions. Weighting alone does not resolve random variance.
2. Reweighted conditional variance stabilization: estimate the current policy's information scale from historical increments
An adaptive policy can change every round, leaving the oracle increment's conditional variance \(\sigma_t^2\) without a stable limit. Theorem 2 does not first require convergence of raw variance; it rescales increments using historical-information-measurable conditional standard deviation estimates. Its one-step estimator and normalization are:
Theorem 2 gives the asymptotically standard normal statistic \(\hat A_T(\hat\Psi_T^{\mathrm{os}}-\Psi_e(\theta_{P_e}))/\sqrt T\). Standard normal quantiles and standard error \(\sqrt T/\hat A_T\) therefore yield a confidence interval. The denominator is the realized sum of stabilized importance weights, not a quantity that can simply be replaced by sample size.
The challenge is estimating the expected squared increment under the current policy. Direct conditional-moment regression could require estimation of the context distribution, reward distribution, and multiple nuisance functions. Instead, the paper uses each historical observation's then-current influence function estimate \(\hat\varphi_s\) and transports its square from the historical sampling policy to the variance scale associated with the current policy. Equation 9 is:
Both the historical policy \(\pi_s\) and current policy \(\pi_t\) appear in the weight; this is not the ordinary sample variance of historical IPW increments. The historical policy cancels the observation's source law, while the current policy retains the information scale of the target conditional variance. Each historical influence function uses its contemporaneous fitted estimates, without fitting an additional reward-distribution model for variance estimation. The first-round variance is initialized to a fixed positive finite constant.
Theorem 2 requires all estimates to be measurable before observing the current sample; time-averaged fourth powers of the target, nuisance, and Riesz estimation errors must each be \(o_p(T^{-1})\); oracle weighted increments must satisfy conditional Lindeberg; and the maximum evaluation-to-sampling density ratio must be \(O(T^{1/3})\). Estimated variances also need a positive lower bound, and the time average of true-to-estimated variance ratios must converge to 1. These are sufficient conditions, not permission to plug in arbitrary nonparametric estimators.
Theorem 4 establishes when Equation 9 satisfies the final variance consistency requirement, under stronger conditions: actual policies lie in a deterministic pointwise-separable class; its log policies obey a covering-entropy bound in the uniform norm with exponent in \([0,2)\); the maximum importance ratio over the entire class is only \(O(T^{1/8})\); and both true and estimated standard deviations have positive lower bounds. The low-complexity restriction applies to the sampling policy class, not a requirement that target and nuisance estimates use only parametric models.
The estimated influence function also needs a historical-information-measurable center: its time-averaged squared error must be \(o_p(T^{-1/2})\), and centers must be uniformly bounded in probability. The variance formula itself avoids an additional conditional-moment model, but valid center and nuisance estimates still need construction. Appendix D recommends burn-in and holdout blocks, so “no additional data splitting” should not be read as saying the complete implementation never uses held-out data.
3. Random-limit self-normalization: omit per-round variance estimation only under a stable random scale
Between deterministic stability and persistent instability lies a regime where average conditional variance converges to a finite, almost surely nonzero random variable \(\kappa\) determined by the realized trajectory. Without normalization, the limit can be a Gaussian variance mixture rather than a Gaussian with fixed variance. Theorem 3 estimates this same random scale from observed squared increments and divides it out to obtain a standard normal statistic.
Its point estimator omits per-round standard deviation weights while retaining importance weighting and one-step correction. To center squared increments correctly, choose \(m_T\to\infty\) with \(m_T=o(T)\) and obtain \(\bar\Psi_T\) using early fitted estimates and a subsequent centering block. The realized square sum starts at round \(2m_T+1\). Equations 7–8 can be written as:
Under the conditions, \((\hat S_T-\Psi_e(\theta_{P_e}))/\sqrt{\hat V_T}\) is asymptotically standard normal. The square sum is realized quadratic variation rather than separately modeled per-round conditional variance; unweighted residual variance is not a substitute. The first \(2m_T\) observations support centering and are excluded from this square sum, while the original point estimator is still defined as a sum over all rounds.
Beyond Theorem 2's estimation-error, Lindeberg, and \(O(T^{1/3})\) overlap conditions, Theorem 3 requires the three early fitted objects to achieve \(o_p(m_T^{-1/4})\), a maximum importance ratio of \(o(m_T^{1/2})\), and convergence in probability of average oracle conditional variance to the stated \(\kappa\). The last requirement determines whether this route applies; it is not an optional technical detail.
Sublinear exploration followed by a frozen policy, or a single random constant-policy block occupying nearly the whole sample, are the paper's motivating examples. However, experiments redesigned for each terminal horizon \(T\) still need their variance limits checked. Appendix E.4 explicitly notes that terminal average-variance convergence alone is insufficient to establish negligibility of the discarded early block in an arbitrary horizon-dependent array. A policy frozen in the latter part of one trajectory does not verify all cross-horizon conditions.
If weights inflate with sample size, the centering block must also be sufficiently long. The strict little-\(o\) requirement should govern the choice of \(m_T\); the paper's verbal pairing of \(T^{1/3}\) weights with a \(T^{2/3}\) centering block is only an order-of-magnitude indication, and equal exponents at the boundary need not satisfy the condition. Policies updated every round without a finite nonzero random variance limit should use Theorems 2–4, not claim self-normalization is valid merely because a square sum can be computed.
A Worked Example¶
In dynamic pricing, the action is a common price, the reward indicates customer renewal, and customer attributes are the context. Renewal probability follows a logistic random-utility model; price sensitivity \(\beta\) is the inference target, while baseline demand heterogeneity is an unknown nuisance function. Prices do not depend on the current customer's attributes, so the example differs slightly from the general context-before-action description but fits the framework through context-independent policies.
The ordinary logistic score multiplies price by the renewal residual and is generally not orthogonal to first-order changes in unknown baseline demand. Appendix C.2 residualizes price using its conditional mean weighted by Bernoulli variance under the evaluation law, then multiplies residualized price by the renewal residual. The associated information is the population mean of squared residualized price times Bernoulli variance. This removes the first-order effect of demand heterogeneity from the price-sensitivity score.
Implementation first fits demand and price sensitivity on historical data, selects a greedy price using estimated revenue, and retains exploration probability. For each new customer, it records price, renewal, and actual sampling probability, then computes the orthogonal corrected value. With frequently changing policies, historical squared increments are transported to estimate current variance and stabilize each round. Direct accumulation of realized squared increments is justified only after the random-limit conditions are separately established. Both routes construct intervals for the same price-sensitivity target; they do not alter the sampling policy by comparing interval lengths.
Loss & Training¶
This is a statistical inference scheme without a generic neural training loss. Target fitting can use weighted risk under the evaluation law; the Riesz representer can be estimated by optimizing the Hessian quadratic form minus the functional derivative. The chosen learners must still satisfy the theorem's measurability, local regularity, and rate requirements.
The authors recommend a fixed-proportion initial burn-in to train pilot estimates, exclude those observations from inference, and optionally freeze fits or update them in batches. With frozen pilots, the generic time-averaged condition reduces to each of the three fitted objects achieving \(o_p(T^{-1/4})\). This is an implementation suggestion for controlling large early errors and computation, not permission to repeatedly treat burn-in data as new independent samples.
Dynamic pricing simulations fit ridge-penalized logistic regression every round with \(\lambda=0.1\) and features comprising low-order polynomials and pairwise interactions of six customer attributes. This particular experiment does not test every black-box nonparametric learner or empirically establish each generic theorem rate.
Key Experimental Results¶
Main Results¶
The paper primarily reports coverage and average interval length. Coverage is the proportion of repeated simulations in which the confidence interval contains the target. The cache preserves captions and qualitative descriptions rather than a table of exact readings, so the following table reports verifiable configurations and qualitative findings without inventing coverage decimals or percentage gains from plots.
| Experiment / Comparison | Verifiable configuration | Reported result | Evidence |
|---|---|---|---|
| Main dynamic pricing figure | Nominal coverage 0.95; caption specifies \(T=5000\) and 200 repetitions | Both proposed routes approach nominal coverage and have similar interval lengths | Figure 1, Section 4 |
| GLM and IPW-GLM | Both use standard sandwich variance estimates | Substantial undercoverage under adaptive sampling; IPW alone is insufficient | Section 4, Figure 1 |
| Additional pricing setup | \(\beta=1\), \(\tau=0.5\); 10 prices from 0.5 to 5 in increments of 0.5; uniform evaluation policy; \(\epsilon=0.1\) | Baselines fail without a unique optimal price; undercoverage is also observed with a strong margin | Appendix C.2, Figure 5 |
| No-margin two-arm bandit | Gaussian rewards with variance 1; equal arm means; uniform evaluation policy; \(T=10000\) | Theorem 2 intervals attain target coverage; OLS undercoverage increases with adaptivity | Appendix C.1, Figures 2–4 |
Pricing sample sizes conflict within the source: C.2 specifies \(T=2000\), the captions of Figures 1 and 5 specify \(T=5000\), and C.3 lists \(T=10000\) in its computational description. These cannot be silently merged into one definitive setup. C.1 uses trajectory length 10000, while C.3 separately gives 500 replications per cell in the bandit grid; trajectory length is not a replication count.
Ablation Study¶
The paper does not report conventional module-removal ablations. The following sampling-mechanism and theoretical-boundary analysis uses actual Appendix C configurations, without fabricating numerical drops from removing correction or variance components.
| Sampling configuration / Analysis | Source setting | Theoretical and empirical boundary |
|---|---|---|
| Sublinear exploration then commitment | \(T_0=T^{1/2}\) | Representative stable-random-variance regime; all Theorem 3 conditions still need verification |
| Linear exploration then commitment | \(T_0=\lfloor0.5T\rfloor\) | The initial phase is not a negligible sublinear block, so that example's self-normalization guarantee does not apply automatically |
| Piecewise-constant policies with random switching | Update probability \(p\in\{0.1,0.2,0.3,0.7\}\) | Local stability does not imply the required global random limit; per-round stabilization applies under its conditions |
| Fully adaptive policy | Policy updates every round | Theorem 2 route has conditional guarantees; favorable Theorem 3 simulations are not an additional theorem |
| Algorithm comparison in the no-margin two-arm problem | \(\epsilon\)-greedy: \(\epsilon=0.1\); Thompson sampling: Gaussian prior scale \(\tau=10\); also UCB | OLS severely undercovers with Thompson sampling; Theorem 3 intervals are wider with UCB |
Key Findings¶
- Variance stabilization and bias correction are not interchangeable: the failure of IPW-GLM shows that correcting action distributions alone does not validate classical intervals.
- No margin deliberately creates difficulty, but pricing baselines also undercovered in the strong-margin setting; a unique optimal action is not sufficient to establish finite-sample reliability.
- Theorem 3 approaches nominal coverage in fully adaptive pricing simulations and is more conservative in some bandit settings. The authors explicitly identify this as an empirical phenomenon without a general theoretical explanation.
Highlights & Insights¶
- Separating target construction from variance handling unifies corrections for average treatment effects, parameter projections, and partially linear targets. Applications should first define what the evaluation law's interval covers, avoiding confusion between misspecified projection coefficients and structural truths.
- Transport weights for conditional variance use both current and historical policies. They reuse stored increments rather than retraining a complex conditional-moment model after every change.
- A random variance limit does not imply that every central limit theorem fails. With stable convergence and consistent quadratic variation, self-normalization can remove the random scale.
Limitations & Future Work¶
- Sampling policies must be known; estimated or missing propensities are not directly covered. Preserving asymptotic normality with unknown policies is an open question identified by the authors.
- These are asymptotic intervals for a fixed target, not finite-sample guarantees or anytime-valid confidence sequences, and they do not automatically justify arbitrary optional stopping.
- Misspecification is allowed, but target uniqueness, curvature bounds, orthogonality, envelope moments, estimation rates, and sufficient overlap remain required. Black-box learners and degenerating exploration need separate checks.
- Evaluation is simulated, without real deployments or extensive nonparametric learner comparisons; inconsistent sample-size descriptions limit precise replication. Attainment of semiparametric efficiency bounds remains unresolved.
- Appendix D's cached centering definition omits normalization whereas its proof uses an average scale. Lemma 7 states weighted fourth moments are \(O_p(1)\), while its proof retains potentially growing weight factors; Lemma 8's statement also differs from the subsequent squared-difference argument. No author formulas are reconstructed here, and the appendix advice is not treated as a fully verified executable recipe.
Related Work & Insights¶
- vs van der Laan et al. (2026): The paper inherits the Hessian–Riesz debiasing representation for nonparametric M-estimands, replaces independent-sample analysis with known-policy martingale inference, and adds two adaptive variance-handling routes.
- vs Hadad et al. (2021) / Bibaut et al. (2021): Both lines rely on weighted influence functions and variance control, but this paper provides a generic construction for smooth M-functionals beyond arm means and policy values.
- vs Guo and Xu (2025) / Leiner et al. (2026): It extends misspecification-tolerant finite-dimensional estimation ideas to broader targets, while retaining policy complexity, overlap, and nuisance estimation requirements.
- vs online linear debiasing methods: The unknown-policy approaches of Deshpande, Ying, and Khamaru differ from this paper; the latter is not a replacement that dispenses with logging probabilities.
Rating¶
- Novelty: 4/5 — Unifies adaptive inference for nonparametric M-functionals and distinguishes persistent change from stable random variance.
- Experimental Thoroughness: 3/5 — Simulations span multiple sampling regimes, but numerical readings and replication settings are insufficiently clear.
- Writing Quality: 3/5 — The main argument is clear, with appendix formulas and sample sizes requiring clarification.
- Value: 4/5 — A useful methodological reference for online experiments, bandit logs, and inference under model misspecification.