Diversity Combining for Multi-Path LLM Reasoning¶
Conference: NeurIPS2026
arXiv: 2609.38829
Area: LLM Reasoning
Keywords: self-consistency, multi-path reasoning, correctness correlation, effective sample size, adaptive compute budget
TL;DR¶
The paper explains diminishing returns in multi-path reasoning through correctness correlation and the design effect, then estimates a fixed deployment budget from a labeled offline four-path pilot; across five modelโtask configurations, 4โ10 paths retain 96%โ103% of the binary majority-vote accuracy at 32 paths, although that evaluation metric is not the actual plurality-vote accuracy for open answers.
Background & Motivation¶
Self-consistency (SC) repeatedly samples chain-of-thought paths for the same question and selects an answer by vote count. Adding paths is straightforward, but does not necessarily add an equal amount of independent evidence: questions within a model's competence often succeed on every path, while questions outside it repeatedly fail. Substituting aggregate single-path accuracy into an independent Bernoulli voting model can therefore be overly optimistic for strong models and overly pessimistic for models whose accuracy is below one half.
The statistical level of this correlation matters. Paths can be sampled independently within a question, yet their correctness remains positively correlated after pooling questions of different difficulty. The paper does not establish unavoidable dependence between sampling random seeds; instead, it compresses across-question competence heterogeneity into a correctness correlation coefficient and uses that coefficient to diagnose the marginal value of repeated sampling. Diversity combining in wireless communications supplies analytical language, not a claim that an LLM internally implements a Gaussian channel.
The paper consequently places the question of whether more paths are worthwhile upstream of aggregation design: estimate the correlation structure of a model, prompt, and task before choosing its budget, and consider weighting only when branch quality differs substantially. Core idea: use correctness correlation to quantify effective-sample-size saturation, select a deployment path budget through offline calibration, and use prompt-template changes as a structural diagnostic rather than equating answer diversity with improved accuracy.
Method¶
Overall Architecture¶
The inputs are labeled calibration questions and subsequent unlabeled queries. The implemented procedure performs โCorrelation calibrationโ followed by โBudget selectionโ; deployment uses only โUniform aggregationโ at the fixed budget, without accessing correctness labels for new queries. The โTemplate probeโ is a separate controlled experiment examining whether prompt changes alter correlation, not a mandatory stage of Adaptive-K deployment.
Three objects serve distinct roles: a latent-embedding channel provides the theoretical abstraction, correctness indicators support offline statistics and evaluation, and discrete answers support actual deployment. Default deployment neither decodes latent embeddings nor runs GLS; empirical GLS-R requires hidden states and covariance calibration and is a separate optional construction.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Calibration questions and gold answers"] --> B["Correlation calibration"]
B --> C["Budget selection"]
C -->|Fixed budget, no new labels| D["Uniform aggregation"]
E["New query"] --> D
D --> F["Plurality answer"]
B -.->|Separate SC versus PT comparison| G["Template probe"]
G -.-> H["Correlation-structure diagnosis"]
Key Designs¶
1. Correlation calibration: turn shared success and failure across questions into a budget diagnostic
Each calibration question receives four sampled paths. Extracted answers are compared with gold answers to obtain binary correctness. Open QA and DROP use a token F1 threshold of at least 0.5; code tasks use execution pass/fail. Mean single-path accuracy is denoted by \(\bar p\), and the correlation estimate averages over all path pairs:
This is neither the fraction of identical answer strings nor a label-free confidence measure. It measures whether different paths jointly succeed or fail on the same question set. Identically prompted SC slots can be treated as exchangeable; independent sampling conditional on a question does not remove the marginal correlation induced by pooling questions.
Under common accuracy and equicorrelation, the variance of average correctness equals the independent variance multiplied by a design effect. Effective sample size therefore means the number of independent samples with the same variance of their mean, not the number of distinct solution strategies. The paper separately defines a participation-ratio effective rank for latent covariance; these quantities must remain distinct:
At the voting level, \(c\) is correctness correlation; at the latent level, \(\rho\) is channel-gain correlation. The former cannot be replaced by the square of the latter. Appendix G derives the participation ratio from one shared eigenvector direction and the remaining difference directions of an equicorrelated matrix; Appendix D directly sums Bernoulli variances and covariances to obtain the voting formula. The Gaussian-threshold surrogate in Appendix V establishes only that correctness correlation increases monotonically with decision-layer correlation at fixed single-path accuracy, not numerical equivalence between the levels.
Effective sample size alone does not uniquely determine majority-vote accuracy. To predict an accuracy curve, the authors additionally assume that each question has a Beta-distributed success probability and paths are independent conditional on that probability, producing a beta-binomial (BB) model:
Method-of-moments estimates determine the parameters, after which the model sums the probability that more than half the paths are correct; an exact tie at even budgets contributes half its probability. This mixture incorporates question-difficulty heterogeneity, but knowing the mean and correlation does not prove that the true distribution belongs to the BB family.
2. Budget selection: stop expanding the budget when marginal effective-sample-size gain becomes small
Adaptive-K does not check answer agreement query by query. It selects one global budget for a model, prompt, and task configuration. The authors compare the derivative of continuously extended effective sample size with a prescribed threshold and solve for the operating point:
The default threshold is \(\varepsilon=0.025\), and the maximum budget is 32. The algorithm first checks whether pilot mean accuracy degenerates to all correct or all incorrect; a degenerate pilot falls back to the maximum budget. Otherwise, it clips the correlation estimate to 0.05โ0.99 and restricts the selected budget to between 1 and the maximum. Clipping is part of the implementation: a near-zero estimate should not simply be substituted into the expression containing division.
A โfour-path pilotโ means four paths per calibration question, not merely four answers to one question. The recommended calibration set contains 50โ100 questions, with labels used only at this stage; new deployment queries require no labels. Changes to the model, prompt, or question distribution fall outside the claimed transfer regime for the old operating point and require recalibration.
The rule controls marginal diversity in a variance-based sense. It is neither a strict guarantee on accuracy loss nor an optimal controller for a latency constraint. Its advantage is a small fixed calibration requirement; its limitation is that it cannot distinguish easy and difficult questions within a task.
3. Uniform aggregation: default to equal weights for symmetric branches without extending GLS to universal voting optimality
Deployment generates the selected number of paths, extracts discrete answers, and returns the most frequent answer, i.e., the plurality vote. The theoretical support for equal weighting comes from a separate linear embedding-estimation model: each branch representation contains a shared target scaled by a quality gain plus zero-mean error. Error covariance is \(\boldsymbol\Sigma\), and the quality-gain vector is \(\mathbf g\). Under a linear unbiasedness constraint, minimum mean-squared-error weights are:
Appendix H derives this expression using a Lagrangian constraint. With equal gains, variances, and correlations, inverse covariance maps the all-ones vector back onto its own direction, yielding uniform optimal linear weights. Appendix Q further shows that, with equal gains and feasible uniform weights, strict improvement requires covariance row sums that are not all identical. Correlation does not automatically call for weighting: if every branch is equally redundant, none deserves extra emphasis.
Embedding mean-squared error and discrete classification accuracy are different objectives. They may align only when decoding preserves the relevant symmetry; this is not a theorem that majority vote is Bayes-optimal among all aggregation rules. Actual equal-weight discrete voting also does not execute GLS before decoding.
Evaluation introduces another important distinction: MV@K is a binary majority vote over gold-scored correctness indicators, scoring 1 if more than half the paths are correct and 0.5 at an exact tie. Deployment with open answers compares the counts of individual answers; when wrong answers fragment, the correct answer can win without receiving half of all votes. Binary majority-vote curves therefore cannot directly serve as deployment plurality accuracy for open QA. Appendix W's argument under uniform wrong-answer fragmentation is heuristic and does not prove stochastic dominance for the full multinomial distribution.
4. Template probe: change branch distributions to test whether the saturation diagnostic responds to configuration changes
The controlled experiment assigns different templates to eight path slots. Math templates include algebraic solving, estimation before precise calculation, backward reasoning, and substitution-based verification; QA templates include key-fact extraction, direct answers, and answering followed by verification. These changes break identical branch distributions and test whether prompting can alter correlation structure, rather than presenting template variation as a new training method.
Cross-task comparisons report mean pairwise Pearson correlation between path-correctness vectors, denoted in the source table by \(\hat\rho_{\mathrm{SC}}\) and \(\hat\rho_{\mathrm{PT}}\). These are observed statistics, not latent channel \(\rho\). Relative change is defined as \(\Delta\rho=(\hat\rho_{\mathrm{PT}}-\hat\rho_{\mathrm{SC}})/|\hat\rho_{\mathrm{SC}}|\). Under SC this statistic coincides with the earlier correlation estimator; under heterogeneous templates it can differ slightly. Equicorrelated effective sample size should not unconditionally be treated as an exact summary of the full covariance structure.
Templates can reduce correlation while also reducing single-path accuracy, so decorrelation does not guarantee more accurate majority voting. Appendix I observes larger weighting gains in PT cells with high slot-accuracy variation, but weights are estimated from accuracy on the same evaluation set. This is a label-informed diagnostic reference, and deployment still requires held-out calibration. The authors call it an oracle upper bound; it is not a universal accuracy bound that no deployable aggregator can exceed.
A Worked Example¶
For the Qwen2.5-7B GSM8K configuration, generate four paths on each of 100 labeled questions and estimate correlation at approximately 0.60 from correctness patterns. The default threshold selects six paths; each subsequent query generates six answers and votes by answer identity, without consulting gold answers. This illustrates the configuration in Table IV, not an additional experiment.
In evaluation, binary majority-vote accuracy is 81.6% with six paths and 81.9% with 32. Calibration is expensive if deployment serves just one new query, but can be amortized over many queries under a fixed configuration. Table IV uses the accounting expression \((K^{*}+4)/32\), giving 31%. Under the algorithm's actual one-time calibration, the path-count ratio for \(N\) new queries is \((4n+K^{*}N)/(32N)\); deployment does not require four additional labeled paths for every new query.
Key Experimental Results¶
Main Results¶
The main GSM8K saturation experiment covers three model families and five seeds, with 100 questions per seed. Cross-task evaluation covers five instruction-tuned models, 12 benchmarks, and eight paths, with 50 questions per seed. The main sampling temperature is 0.7. The following table reproduces selected results from Table IV: all accuracies are binary majority-vote scores, and retention and cost retain the source's integer rounding.
| Task / model | Four-path correlation | Selected paths | Selected-budget accuracy | 32-path accuracy | Retention | Table IV cost |
|---|---|---|---|---|---|---|
| GSM8K / Qwen-7B | 0.60 | 6 | 81.6% | 81.9% | 100% | 31% |
| GSM8K / Llama-8B | 0.53 | 8 | 78.1% | 79.3% | 98% | 38% |
| GSM8K / Mistral-7B | 0.45 | 10 | 43.6% | 42.4% | 103% | 44% |
| HotpotQA / Llama-8B | 0.61 | 6 | 50.2% | 52.2% | 96% | 31% |
| BoolQ / Llama-8B | 0.79 | 4 | 80.2% | 80.3% | 100% | 25% |
Paired bootstrap 95% intervals for the accuracy difference include zero in all five cells; 103% retention is not reliable evidence that fewer samples are superior. Table IV cost counts paths, not measured GPU time or token expenditure. Appendix J gives 400 paths for a one-time 100-question calibration; serving 1000 queries costs 13.8%โ32.5% of fixed 32-path sampling across the five configurations, including calibration. Some cost descriptions in Appendices R/S mix one-time calibration with Table IV's simplified accounting; the note distinguishes these conventions instead of silently making them identical.
Ablation Study¶
The following saturation and calibration analysis comes from Tables I and II, rather than network-module removal ablations. BB and the independent binomial model are fitted on the same evaluation data, so these are in-sample predictions.
| GSM8K model | Effective sample size at 32 paths | Observed MV@32 | BB prediction | Independent binomial prediction |
|---|---|---|---|---|
| Qwen2.5-7B | 1.67 | 81.9% | 81.1% | 100.0% |
| Llama-3.1-8B | 1.94 | 79.3% | 77.8% | 99.9% |
| Mistral-7B | 2.13 | 42.4% | 41.3% | 20.6% |
For Qwen, Table I reports correctness correlation 0.586, effective sample size 1.67, and ceiling 1.71 at 32 paths, reaching 98% of the ceiling. Mean single-path accuracy is 79.2%, not the majority-vote accuracy of 81.9% in Table II. BB in-sample error is 0.8โ1.5 percentage points. When a four-path pilot on half the questions predicts 32-path outcomes on the other half, error increases to 3.3โ4.8 percentage points, versus 18โ24 for the independent model. These are different levels of validation and should not be conflated.
The following table selects four benchmarks from Table III. Correlations are observed Pearson correlations. Percentage reductions average model-level relative changes and need not equal the ratio computed from the displayed mean correlations.
| Benchmark | Valid models | Mean SC correlation | Mean PT correlation | Mean relative correlation change | Mean effective-sample-size change |
|---|---|---|---|---|---|
| TriviaQA | 5 | 0.68 | 0.19 | -71.5% | +2.1 |
| HotpotQA | 5 | 0.55 | 0.17 | -62.5% | +2.0 |
| GSM8K | 5 | 0.55 | 0.38 | -29.4% | +0.5 |
| MATH | 5 | 0.60 | 0.55 | -9.2% | +0.2 |
Key Findings¶
- Of 60 modelโbenchmark cells, three DROP cells are excluded because single-path accuracy is below 2%. Correlation decreases in 55 of the remaining 57, but only 43 have 95% intervals excluding zero. This does not mean that 55 cells significantly improve accuracy.
- PT accuracy changes in Appendix L have mixed signs across six domains: mean changes are -0.4 percentage points for QA, -1.3 for Math, -3.9 for Code, and +2.3 for NLU. Expanding diversity and maintaining path quality are separate objectives.
- In Appendix P, a 50-question pilot gives a maximum correlation-estimate coefficient of variation of 14.2% and budget standard deviations of approximately 1.5โ2.0. Some budget estimates are substantially less stable at 25 questions. A low-cost four-path pilot still needs enough questions.
- Qwen3.5-9B in thinking mode selects 14 paths on GSM8K in Appendix M and retains 99.8% of MV@32. Generation is capped at 16,384 tokens, and 5.4% of paths are scored incorrect for reaching that cap; these results cannot be extrapolated to harder questions or unlimited reasoning lengths.
Highlights & Insights¶
- Separating sampled paths from effective evidence is more diagnostic than plotting accuracy curves alone. The same 32 paths can correspond to a very small design-effect effective sample size when correlation is high.
- Symmetric redundancy does not automatically justify complex weighting; checking branch heterogeneity first is more useful. GLS reveals this condition rather than endorsing arbitrary majority-vote rules.
- Template variation as a structural probe exposes task differences better than reporting a single mean accuracy gain. Open-answer tasks decorrelate more strongly, but this remains an association rather than causal proof about answer-space openness.
Limitations & Future Work¶
- Binary scoring discards the wrong-answer distribution, so deployment plurality outcomes for open answers can differ from the evaluated curves. Actual deployment-answer accuracy and wrong-answer concentration need separate reporting.
- A global budget cannot allocate compute by question difficulty, and matched-compute comparisons with online stopping remain open. Per-query stopping could be layered onto the offline budget prior, but requires independent validation.
- Once PT breaks exchangeability, a single mean correlation cannot fully describe slot quality and covariance structure. Appendix J reports a mean ratio of 1.16 between block-variance effective sample size and the equicorrelated approximation, reaching 2.5 in an individual open-QA cell.
- BB is a parametric mixture assumption, and the saturation ceiling bounds effective sample size rather than arbitrary task accuracy. Difficulty stratification, non-Beta mixture checks, and held-out prediction intervals would strengthen the analysis.
- Calibration still requires 50โ100 labeled questions and must be repeated when the configuration changes. Without labels, under ongoing distribution shift, or with very few questions, the rule is not a calibration-free solution.
Related Work & Insights¶
- vs Self-Consistency: Repeated sampling and answer voting remain unchanged; the addition is an offline diagnosis of budget saturation, not a new chain-of-thought generator.
- vs Adaptive-Consistency / ESC / RASC: These methods stop per query based on agreement or quality during sampling; this paper selects a fixed budget from calibration. They may be complementary, but no matched-compute superiority claim is established.
- vs CISC and other confidence-weighted methods: The paper identifies heterogeneous slots as the more promising regime for weighting. Same-set label-derived weights in Appendix I do not replace deployment confidence-calibration experiments.
- vs compound-inference scaling: Both address across-question difficulty heterogeneity; this paper summarizes it with one correlation statistic without separating its sources. A research direction is to jointly examine conditional correctness correlation and wrong-answer concentration, distinguishing saturation due to question mixing from voting failures due to specific error modes.
Rating¶
- Novelty: 4/5. Classical statistics are organized into an actionable LLM multi-path budget diagnostic; the central formulas are not new theory.
- Experimental Thoroughness: 4/5. Multiple families, tasks, and held-out BB tests provide useful coverage, but deployment plurality metrics and online-stopping comparisons are missing.
- Writing Quality: 4/5. Assumption mapping and theoretical boundaries are explicit, although some cost descriptions and observed-correlation notation can confuse readers.
- Value: 4/5. Useful for stable inference configurations with labeled calibration sets, not a universal voting-optimality or label-free stopping guarantee.