Skip to content

FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models

Conference: NeurIPS2026
arXiv: 2609.35947
Area: Discrete Diffusion Models / Optimization & Theory
Keywords: proposal control, graph flux, Feynman–Kac, sequential Monte Carlo, variance control

TL;DR

FluxLite jointly adjusts sparse jump rates and compensating weights without retraining a discrete diffusion model, preserving the target marginal-distribution path through graph divergence while mitigating particle-weight degeneracy with the local HEU rule or a small nonnegative quadratic program, D-VCG.

Background & Motivation

Reward alignment, temperature adjustment, and posterior sampling require turning a pretrained generative distribution into a tilted distribution, rather than merely drawing more independent samples. Feynman–Kac sequential Monte Carlo (SMC) splits this process into particle propagation and importance weighting: particles follow proposal dynamics, and weights compensate for probability changes that propagation does not realize. D-FKC already provides this correction for discrete diffusion, but when propagation is poorly aligned with the target tilt, a few particles absorb most of the weight, and additional particles do not necessarily yield more effective exploration.

The question is therefore not only how often to resample, but whether particles can first move in a more target-aligned direction, leaving less change for weights to carry. For continuous diffusion, DriftLite exploits an equivalence between drift control and weighted divergence; discrete diffusion can only jump along directed edges allowed by the pretrained reverse process, with nonnegative rates on every edge. Directly transplanting continuous vector-field control can introduce unavailable edges or negative rates.

FluxLite instead controls probability flux on a graph, distinguishing the algebraic identity that preserves the target from the approximate optimization that selects a low-variance representative. Core Idea: reallocate target probability changes between sparse jumps and residual weights, preserve the population target with exact graph-divergence compensation, and improve finite-particle sampling through training-free local variance control.

Method

Overall Architecture

The inputs are pretrained reverse jump rates, a reward or guidance potential, a time grid, and the current weighted particles; the outputs are weighted samples approximating the terminal tilted distribution. Reverse time advances from the noisy end to the data end, typically with the reward gradually activated; the temperature exponent and reward jointly define the target path:

\[ q_t(x)=\frac{(p_t^{\leftarrow}(x))^{\gamma}e^{r_t(x)}}{Z_t}. \]

Here \(Z_t\) is the normalizing constant, the main-text tilted construction assumes \(\gamma>0\), and the reward can use a linear ramp from zero at the noisy end to the target strength at the data end. The algorithm must initialize from \(q_0\): the base-model prior cannot arbitrarily stand in for the tilted prior. They coincide when the noisy-end distribution is uniform and the reward is still inactive. “Path preservation” means preserving the entire path of time marginals, not preserving the joint distribution of random trajectories under the controlled CTMC.

Each step first formulates control through flux-equivalence compensation, then selects either HEU or D-VCG to obtain effective rates and a residual potential. Algorithm 1 orders the operations as control construction, weighting, ESS-based resampling, and propagation; there is no newly trained neural network or training-supervision branch. The diagram represents this inference data flow. The two controllers are alternatives, not a sequence that applies HEU before D-VCG.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400, 'subGraphTitleMargin': {'top': 8, 'bottom': 16}}}%%
flowchart TD
    A["Reverse rates, reward<br/>weighted particles"] --> B["Flux-equivalence<br/>compensation"]
    B -->|Select HEU| C["HEU local<br/>reallocation"]
    B -->|Select D-VCG| D["D-VCG nonnegative<br/>basis control"]
    C --> E["Effective rates<br/>and residual potential"]
    D --> E
    subgraph S["Controlled SMC recursion"]
        direction TB
        F["Residual-potential weighting"] --> G["ESS-triggered resampling"] --> H["Effective-rate propagation"]
    end
    E --> F
    H -->|Next time step| A
    H --> I["Terminal weighted samples"]

Key Designs

1. Flux-equivalence compensation: change particle motion without changing target marginals

The paper uses the column-source convention: \(Q_t(y,x)\) is the rate from source state \(x\) to destination state \(y\). Increasing outflow on an edge reduces probability at its source, so the compensating potential must use graph divergence defined as net outflow divided by current mass, not the opposite sign. For an admissible perturbation, the core identity is:

\[ Q'_t=Q_t+R_t,\qquad g'_t(x)=g_t(x)+\operatorname{div}_{q_t}R_t(x),\qquad \operatorname{div}_{q_t}R_t(x)=\frac{\sum_{y\ne x}\left[R_t(y,x)q_t(x)-R_t(x,y)q_t(y)\right]}{q_t(x)}. \]

The added transport flux and the mass change from the added potential cancel term by term, leaving the continuous-time population equation for the same \(q_t\); the divergence also has zero mean under \(q_t\). The identity requires \(q_t(x)>0\) on the support being considered and \(Q'_t(y,x)\ge0\); practical control additionally restricts edges to the sparse support of the pretrained reverse graph. This is a population equivalence using exact target quantities, not a claim that density estimation, numerical integration, and finite-particle sampling are all error-free.

To give different controllers a common starting point, the authors first turn off all propagation, obtaining the pure-reweighting potential \(g_t^0=g_t-\operatorname{div}_{q_t}Q_t\). The residual potential for any candidate effective rate is this pure-reweighting potential plus the candidate rate's graph divergence; control minimizes its variance under the target distribution. In flux variables, this becomes a sum of squared residual mass terms weighted by \(1/q_t(x)\), subject to nonnegativity and sparsity: a convex weighted least-squares problem, rather than an arbitrary modification of denoising logits.

On a complete directed graph, DEN can explicitly eliminate the residual potential, but it requires full-state densities and dense jumps, so it is not directly implementable for large-scale discrete diffusion. PR occupies the opposite extreme: no propagation, with every change carried by weights, making it unable to create states absent from the initial particle support. These extremes explain why lower residual variance and lower terminal error under a finite budget are not the same optimization problem.

2. HEU local reallocation: move source-term differences within a one-hop neighborhood

HEU extracts a simple rule from the dense zero-residual construction: define the source term \(\mu_t(x)=q_t(x)g_t^0(x)\) and move flux from smaller to larger source terms only along admissible outgoing edges. Positive-part truncation ensures nonnegative rates, a local neighborhood size replaces the total state count in the normalization, and damping limits reallocation strength. With reciprocal neighborhoods and matching bidirectional normalizations, this approximately replaces the source term with a damped neighborhood average, reducing local fluctuations without automatically removing differences between neighborhoods.

HEU is therefore a one-hop heuristic, not an exact sparse least-squares solution, and it has no guarantee of monotonically reducing global variance on arbitrary directed graphs. Mask diffusion can normalize using outgoing neighbors alone or the combined outgoing and incoming neighbor count, producing different control strengths. The appendix gives an analytical expansion of neighboring source terms intended to reuse already evaluated edge rates and densities without extra score-network evaluations; the necessary local quantities must still be available. HEU is omitted from the Ising experiments because a direct particle histogram spuriously assigns zero mass to most unvisited single-flip neighbors; adopting the analytical expansion is left for a subsequent implementation.

3. D-VCG nonnegative basis control: reduce full-graph optimization to a small particle-cloud QP

Instead of optimizing exponentially many edges separately, D-VCG prepares a few nonnegative rate bases: the original reverse rates and a target-aligned basis reweighted by density ratios and reward differences. Each basis uses the admissible graph structure, so nonnegative combinations remain valid rates; the control parameters are a few coefficients for the current time step, not model parameters. The population optimization is:

\[ Q_t^{\mathrm{eff}}=\sum_{j=1}^{J}\theta_jQ_t^{(j)},\qquad \theta_t^\star\in\arg\min_{\theta\ge0}\operatorname{Var}_{q_t}\!\left[g_t^0+\sum_{j=1}^{J}\theta_j\operatorname{div}_{q_t}Q_t^{(j)}\right]. \]

The implementation evaluates basis divergences on the current weighted particles, centers them, assembles the quadratic term from weighted covariances, and forms the linear term from covariances between the pure-reweighting potential and the divergences. The coefficients thus reflect which propagation direction best offsets current weight variation, rather than imposing a manually fixed guidance strength throughout sampling. Unconstrained optima satisfy the normal equations \(A_t\theta=-c_t\); a unique unconstrained solution exists only when the covariance matrix is nonsingular, and the nonnegative problem cannot be reduced to direct matrix inversion.

The appendix implementation uses active-set candidate enumeration, a small diagonal ridge, and candidate scoring to select coefficients; candidate generation and scoring use different regularizers, making this an approximate solve. The joint annealing-plus-reward Ising experiment additionally uses a target-basis anchor penalty that grows with the reward ramp; pure annealing and pure reward experiments omit it. The augmented library includes an intermediate-temperature basis and an energy-difference flux basis; the latter uses a forward-rate prefactor rather than strictly retaining the main-text learned reverse-rate multiplier form. The “2-basis” and “4-basis” labels identify minimal and augmented variants. Actual basis counts change under joint guidance or degenerate temperature settings, so the labels are not fixed counts for every configuration.

4. Controlled SMC recursion: propagation control and weight correction must be paired

After constructing effective rates, the algorithm applies exponential incremental weights from the corresponding residual potential, normalizes them, and checks effective sample size (ESS); for normalized weights, ESS is \(1/\sum_n w_n^2\). When ESS divided by particle count falls below the threshold, particles are resampled and reset to equal weights, then propagated to the next time step using the effective rates. Changing rates without the compensating potential leaves the target equivalence class; resampling alone without better propagation does not resolve inadequate exploration of target regions.

The theoretical fixed-grid recursion uses a bootstrap version with weighting, every-step multinomial resampling, and propagation in that order. The CTMC experiments instead use midpoint rates and potentials with half-weight/propagation/half-weight splitting, plus ESS-adaptive systematic resampling; the default threshold is 0.5. This distinction matters: Algorithm 1 is a first-order skeleton, the experimental splitting is not a literal implementation of each skeleton line, and the particle theorem does not fully guarantee the experimental adaptive feedback loop.

A Worked Example

For a masked CTMC with vocabulary size 5 and length 3, the state space contains 216 states, while the terminal unmasked target contains only 125 states. Starting near the all-mask state, PR cannot produce complete sequences absent from the original particle cloud by reweighting alone, so it is not reported in this setting. D-VCG instead mixes the original reverse basis and the target-aligned basis on admissible reveal edges, adjusts particle weights using the residual potential, and then resamples and propagates particles along controlled edges to reveal tokens. At the next step it re-estimates coefficients; explored states provide support for subsequent weight correction. The gain comes from jointly changing movement and residual weights, not from introducing inadmissible jumps. This is a mechanism-level example based on the paper's benchmark, without invented per-particle numerical values.

Loss & Training

FluxLite itself does not retrain the model; the Ising base model is a U-Net pretrained on 200,000 Swendsen–Wang samples with denoising score-entropy loss. Theorem 4.1 analyzes population bias with a known forward kernel and a learned local density ratio: training error uses \(\ell(u)=-\log u-1+u\), where \(u\) is the estimated ratio divided by the true ratio. It also requires bounded guided potentials, a positive bounded ratio window, and a finite coverage factor for the tilted path under the training-edge measure. The coverage factor weights tilt-to-base density ratios at both edge endpoints and aggregates them over time; it is not simply a score-accuracy condition. Strong tilting can amplify bias despite low average training error.

Theorem 4.2 only bounds particle error relative to the same fixed-grid population recursion with deterministic controls and every-step bootstrap resampling. Its particle-count dependence is \(N^{-1/2}\), with constants depending on residual-potential oscillation. The appendix's one-step weight-variance lemma explains why reducing residual-potential variance can reduce exponential-weight variance when potential magnitude is also controlled. These results support the local objective, but do not prove that minimum instantaneous variance yields optimal terminal quality for arbitrary finite particle counts, nor cover feedback error from estimating controls on the same particle cloud.

Key Experimental Results

Main Results

The finite-state CTMC benchmark uses 4,000 particles, 80 steps, and 10 seeds; the uniform-state space has 125 states, and the masked state space has 216. D-VCG improves terminal KL over D-FKC by up to 114.8×, referring to a geometric-mean KL ratio across seeds, not a 114.8× improvement in every condition. KL compares the exact target with the weighted histogram after clipping probabilities at \(10^{-15}\) and renormalizing; figure error bands use the standard deviation of log KL.

Music infilling uses Lakh monophonic sequences of length 256, vocabulary size 129, and a pretrained SEDD model. All methods match 1,024 function evaluations: SGDD uses 32 outer iterations times 32 inner steps, whereas both SMC methods use 128 steps times 8 particles. This matches NFE, not runtime.

Method ρ=40%: Hellinger ↓ ρ=40%: Meas. Error ↓ ρ=60%: Hellinger ↓ ρ=60%: Meas. Error ↓
SGDD 0.076 4.93 0.151 6.11
D-FKC 0.033 0.29 0.081 0.88
D-VCG 0.035 0.00 0.079 0.05

The table preserves the values and column order of the paper's Table 2. The observation model defines \(\rho\) as the fraction of revealed positions, although later prose calls it a “masking ratio”; the interpretation here follows the observation model, not the unrevealed fraction. Measurement error is defined as a mismatch fraction on revealed positions, but the source table does not specify percentage scaling and includes 4.93 and 6.11. The displayed scale is therefore preserved without adding percent signs or converting values. D-VCG has the best observation consistency, but its Hellinger distance at \(\rho=40\%\) is 0.035 versus D-FKC's slightly better 0.033, so it does not lead on every metric.

The Meissonic text-to-image experiment uses 100 prompts, with 25 in each of 4 style categories, 64 denoising steps, 8 particles, and CFG scale 9.

Method MPS ↑ HPSv2 ↑
D-FKC 15.282 ± 0.032 0.2980 ± 0.0003
D-VCG 15.422 ± 0.032 0.3033 ± 0.0003

The ± values are standard errors across 100 prompts, not standard deviations; the scores assess preference/alignment rather than directly measuring distance from a true image distribution. The appendix's CFG parameter identification gives \(\gamma=1-s=-8\), making the main-text Proposition 2.1 rates with a \(\gamma\) prefactor invalid. The authors instead construct effective propagation from nonnegative CFG rate bases; this extension cannot directly invoke the main-text population-stability theorem requiring a positive exponent.

Ablation Study

The table selects configuration analyses from Ising rather than estimating unreported absolute errors from plotted curves. All improvement factors divide D-FKC row-correlation MSE by the corresponding augmented D-VCG MSE; geometric means and peaks aggregate over each row's stated sweep.

Config Geometric-mean improvement Peak improvement Sweep and scope
Pure annealing, βtrain=0.4 5.33× 24.5× βtarget=0.20–0.55; no anchor penalty
Joint annealing and reward, βtrain=0.4 6.77× 55.4× βtarget=0.45; βr=0.02–0.40; anchor penalty
Pure annealing, βtrain=0.3 1.69× 4.84× βtarget=0.20–0.60; ESS threshold 0.25
Pure reward, βtrain=βtarget=0.4 5.21× 17.2× γ=1; βr=0.02–0.40; no anchor penalty
Particle-count sweep, βr=0.20 15.21× 34.7× N=100–5000; peak at N=1000

Ising uses a \(16\times16\) periodic lattice, with each target reference estimated from 2,000 Swendsen–Wang samples and a ghost spin for nonzero external fields. Row-correlation MSE averages squared differences between generated and reference mean uncentered row correlations over separations 1–13; the correlation estimator excludes a one-site boundary margin and does not subtract mean magnetization. Ising curves average 3 seeds, with bands equal to seed standard deviation divided by \(\sqrt{3}\); annealing uses 5,000 steps, while reward and joint configurations use 2,000 steps.

Key Findings

  • Pure reward still improves after removing joint annealing and the anchor penalty, so gains do not depend entirely on added regularization; the joint configuration is not an isolated ablation of the unregularized variance objective.
  • Increasing D-FKC particle count does not eliminate the gap in the tested range; the appendix reports 5,000-particle D-FKC still having 5.4× worse row-correlation MSE than 500-particle D-VCG.
  • Sparse D-VCG sometimes outperforms zero-residual DEN, showing that propagation randomness, damping, and fixed discretization also affect terminal error; this does not invalidate DEN's population zero-residual identity.
  • On a single A100 with 500 Ising particles and the augmented basis setting, D-VCG per-step runtime is within roughly 5% of D-FKC; this is not a universal claim about every task or end-to-end runtime.

Highlights & Insights

  • The method separates whether proposal modification preserves the target from which proposal to choose: algebraic compensation answers the first, approximate optimization the second, clarifying where approximation enters.
  • Nonnegative rate bases make validity a structural property of the parameterization rather than an after-the-fact negative-rate truncation; a few weighted covariances feed current particle needs back into propagation.
  • PR and DEN are extremes of weight-carried and propagation-carried changes, not simply an ordered pair of weak and strong baselines; finite-particle optimization must consider both exploration coverage and propagation noise.

Limitations & Future Work

  • The exact identity requires positive support and local target-density information; a finite-particle histogram cannot replace high-dimensional neighborhood densities, as illustrated by omitting HEU from Ising.
  • The objective is instantaneous residual-potential variance, not terminal KL, preference scores, or every observable; strong guidance, sparse-graph bottlenecks, and insufficient basis libraries can still limit performance.
  • Theorem 4.1 does not cover forward-kernel misspecification, arbitrary independently learned rates, or numerical generator error; Theorem 4.2 excludes score error, time discretization, adaptive ESS, numerical propagation, and particle-dependent controls.
  • The joint-mode anchor penalty, toy damping, and approximate active-set scoring show that practical results involve more than the bare quadratic objective; future analysis should jointly examine these stabilizers and propagation variance.
  • Large-scale validation is currently concentrated on music infilling and image CFG with 100 prompts, without a large-scale text-reasoning benchmark; extension to DLM reasoning is a direction, not an established result.
  • Further directions include end-to-end adaptive-control error analysis, hybrid discrete–continuous control, and distillation into lower-latency generation; these are not completed compression or distillation methods in this paper.
  • vs D-FKC: both use Feynman–Kac correction, but FluxLite additionally changes propagation rates within an equivalence class and compensates the residual potential, targeting weight degeneracy that remains after correction.
  • vs DriftLite: both use control plus divergence compensation, but the control object changes from continuous drift to discrete graph flux, adding sparse-support and nonnegative-rate constraints.
  • vs twisted / controlled SMC: both seek better proposals to reduce degeneracy; FluxLite uses pretrained rate bases and small online optimization instead of training additional proposal networks, without guaranteeing a globally optimal twisted proposal.
  • vs SGDD: SGDD is a split-Gibbs posterior-sampling baseline in music infilling, where D-VCG improves observation consistency; the result concerns this discrete inverse problem at matched NFE, not a general theorem comparing Gibbs and SMC.
  • Research direction: dynamically select control bases using their marginal benefit and covariance conditioning, then separately quantify propagation noise from additional exploration; this is a mechanism-inspired proposal, not a verified result of the paper.

Rating

  • Novelty: 4/5 — Systematizes continuous proposal control as a discrete flux equivalence class with sparsity and nonnegativity constraints.
  • Experimental Thoroughness: 4/5 — Covers exact CTMCs, learned Ising, and two large-scale applications, but lacks text reasoning and broad runtime evaluation.
  • Writing Quality: 4/5 — Provides detailed proof scopes and implementation choices; basis labels, music-ratio terminology, and table units require careful interpretation.
  • Value: 4/5 — Offers a reusable training-free sampling-control principle, not model-parameter compression.