Skip to content

Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

Conference: NeurIPS2026 โ€” Evaluations and Datasets Track (preserved from the supplied task metadata; official acceptance or main-track status was not independently verified)
arXiv: 2609.35279v1
Area: Multi-Agent Systems / LLM Evaluation
Keywords: homogeneous-panel debate, transition ledger, conditional collapse, correction preservation, signed replay

TL;DR

The paper decomposes debate among three copies of one model into preservation, collapse, correction, and unrepaired transitions, using pre-debate screening and round-level traces to locate risk; however, an offline freeze replay prevents 29 collapses while losing 108 corrections, showing that a risk signal is not an effective controller.

Background & Motivation

Multi-agent debate typically asks models to answer independently, inspect their peersโ€™ reasoning, and revise their answers before evaluating the final majority. Final accuracy, however, reveals only the net of two opposing flows: discussion can rescue an initially wrong panel or derail an initially correct one. Answer-change frequency also fails to distinguish direction. The paper illustrates this with Sonnet 4.5 and Llama-3.1-8B: their debate flip rates are 0.705 and 0.726, but their conditional-collapse rates are 2.15% and 8.46%. Willingness to revise is not the same as beneficial revision.

This distinction also changes how mitigation is evaluated. If a model readily revises, or the majority changes in the first round, stopping debate might appear to reduce risk. Yet the same signals can precede useful corrections. Reporting only prevented collapses can portray freezing as a safety improvement while omitting the recoveries it discards. Rather than introducing a new debate architecture, the paper develops an auditable evaluation protocol that separates selection-time screening, runtime diagnosis of individual trajectories, and intervention replay.

The authors deliberately study homogeneous, closed-book multiple-choice questions, using gold answers and rule-based coding to reduce evaluation ambiguity. This does not make parsing and voting choices immaterial. Core idea: express debate gains through a directional transition ledger with fixed denominators, then audit interventions using collapses prevented minus corrections lost, rather than treating a risk ranking as a controller.

Method

Overall Architecture

Inputs comprise the model, question pool, gold labels, probe outputs, and saved round-level debate records. Outputs comprise screening statistics, a four-cell transition ledger, collapse-onset rounds, and intervention prevented/lost/net counts. The stages are pre-debate screening, transition accounting, trajectory localization, and signed replay: measurement and offline analysis stages, not a deployed sequential control network.

Standard debate uses three agents instantiated from the same model. They answer independently at Round 0, then inspect the panelโ€™s current responses and revise through Rounds 1โ€“3; the Round 3 majority is the final response. Agent temperatures are 0.5, 0.7, and 1.0, so homogeneous refers to model identity, not identical sampling settings. The paperโ€™s majority is implemented by coding rules, potentially including removal of unparsed answers and tie-breaking; it does not necessarily mean unanimous three-agent consensus.

Probes separately measure an individual modelโ€™s revisability, whereas round-level records capture actual debate. They are linked through modelโ€“scaffold rows and question identities, but their analysis units differ. The 14-row association cohort, 6,925-debate onset cohort, and 6,525-debate freeze-replay cohort must not be collapsed into one unified experimental sample.

Because the contribution is primarily an evaluation protocol and mechanism diagnosis, measurement dependencies are explained in prose rather than depicted as a network that could be mistaken for an online control architecture.

Key Designs

1. Pre-debate screening: measure revisability rather than produce a safety ranking

Each modelโ€“question pair first receives an initial answer and then eight measurement conditions, crossing four argument-strength levels with two social-channel settings. This note retains only experimental factors and aggregate statistics, not prompt text that could be used to induce incorrect answers. Each condition records whether the answer changes; averaging produces the probe flip rate. The model-level total flip rate is then aggregated over the question pool to prioritize modelโ€“scaffold rows for more expensive trace auditing.

\[ \alpha(M,q)=\frac{1}{|\mathcal{P}|}\sum_{p\in\mathcal{P}}r_p. \]

A revision event is 1 when the answer changes and 0 otherwise. Total flip rate includes both harmful and beneficial changes and is not a collapse probability. The authors additionally condition on initial correctness: adversarial flip rate is the proportion of initially correct answers that subsequently become wrong, while corrective flip rate is the proportion of initially wrong answers that become correct. Moving between wrong options is also a flip, but not a correction. These conditional rates have different denominators and cannot simply be added to recover total flip rate.

Channel diagnosis is not a new controller either. Social sensitivity is the mean flip rate under social conditions minus that under solo conditions. Argument sensitivity is the mean under strong and very-strong conditions minus that under weak and moderate conditions, averaging over the other factor in each case. S/A divides social sensitivity by the sum of the absolute values of both sensitivities plus 0.000001. It preserves the sign of the social effect and is not an ordinary ratio of two positive quantities. Channel differences vary across models; the total-flip-rate association does not establish a uniform independent-channel mechanism within Qwen.

The high-FR pool is selected using probe behavior, not subsequent debate outcomes. From a fixed 1,000-question pool, the procedure takes each modelโ€™s top 200 questions by solo-condition flip rate. This constrains outcome-driven leakage, but different models can still receive different questions. Shared-pool and full-pool checks bound this selection effect; they do not automatically make cross-model comparisons strictly paired.

Screening is intended for the stage before debate traces are collected. Once the three initial agent answers exist, disagreement and answer diversity are directly available, and an additional full probe pass may not justify its cost. The appendix also notes that the full psychometric configuration can require more generations than a small direct debate audit. A single-model probe is therefore not necessarily cheaper.

2. Transition accounting: decompose accuracy into verifiable bidirectional changes

Initial and final panel correctness define four cells: correct to correct is preservation, correct to wrong is collapse, wrong to correct is correction, and wrong to wrong is unrepaired. The operational definition of collapse requires only an incorrect final majority, not unanimity or the absence of intermediate recovery. Conditional collapse uses initially correct majorities as the at-risk denominator, rather than all debates.

\[ C^{\mathrm{cond}}=\Pr(\text{final wrong}\mid\text{initial majority correct}). \]

Conditional correction analogously uses initially wrong majorities as its denominator. When every initial and final majority is defined, corrections minus collapses divided by all debates exactly equals final-minus-initial accuracy. The ledger thus decomposes accuracy rather than replacing it, exposing different risk and recovery structures behind identical net accuracy changes. When majorities are undefined, their treatment must be stated before applying this identity.

Answer extraction prioritizes explicit final-answer markers, answer/choose variants, standalone option letters, and finally a last-valid-option-letter fallback. An unparsed probe answer is treated as no revision; debate voting filters unparsed answers and breaks ties alphabetically. These rules avoid an LLM judge, but they can alter labels on low-signal cases. The protocol therefore also requires parser-failure reporting, recoding with alternative parsers, and voting-stability audits.

3. Trajectory localization: identify when risk appears without presenting hindsight as prediction

For a debate that ultimately collapses, onset is the first round in which the running majority switches from correct to wrong. It can occur in the first, second, or third round. If an early wrong majority is later repaired, onset still should not be equated with the final switch to an incorrect state. Saved round-level answers make timing labels recomputable rather than dependent on subjective judgments of full transcripts.

The authors compare out-of-fold AUC for pre-debate features, features augmented with Round 1, and features augmented with the full trajectory. Round 1 features include majority change, agent-flip count, and agreement, capturing per-question runtime states more directly than static screening. However, the baseline in Appendix E.1 also includes initial correctness, a gold-label audit variable. Full trajectories additionally use later rounds. These results therefore cannot all be described as prospective warnings available in deployment.

AUC measures diagnostic discrimination, not whether a particular action improves utility. The paper also states that its AUC-delta intervals resample debate indices, while the open materials cannot recover the preregistered question/model-clustered intervals. This limits statistical interpretation and reinforces the separation between trajectory associations and causal-mechanism claims.

4. Signed replay: price avoided errors alongside sacrificed recoveries

Offline replay reads completed debates and substitutes the saved initial majority for the saved final majority whenever a gate fires. On collapse rows, this preserves the originally correct answer; on correction rows, it erases the correct answer obtained through discussion. The other two cells generally leave accuracy unchanged. Both prevented collapses and lost corrections must therefore be counted.

\[ U_w(g)=w_{\mathrm{coll}}\,\mathrm{prevented\_collapses}(g)-w_{\mathrm{corr}}\,\mathrm{lost\_corrections}(g). \]

Both weights default to 1, making net utility divided by cohort size the accuracy change relative to saved standard-debate outcomes. With unequal weights, this quotient is normalized weighted utility, not literal accuracy gain. The break-even ratio is the minimum collapse-value/correction-value ratio needed for nonnegative utility. It should be fixed before policy selection, not chosen after observing the results to redefine safety gains.

Probe gating uses leave-one-model-out validation (LOMO). The risk classifier, freeze threshold, and freeze mode are selected using the other models and fixed before evaluation on the held-out model. The authors also test a fixed Round 1 majority-change rule and a strict-LOMO decision stump. These are replay baselines for auditing the ledger, not effective deployed policies delivered by the paper. Their cohorts differ, so their accuracy changes cannot directly rank controllers.

Loss & Training

The paper does not train a new base LLM or fine-tune debate behavior. Its primary statistical analysis averages related variants within seven families and applies an exact two-sided Spearman test. The 14 model rows are a sensitivity view, not 14 independent units of model evidence.

The preregistered plan targeted 18 model rows; missingness and identity/question-pool checks left 14, including six Qwen rows. Family aggregation corrects non-independence in the realized cohort rather than indicating unchanged completion of the original target. Initial probe answers actually use temperatures 0.5, 0.7, and 1.0; only post-challenge replies use temperature 0 where supported. Appendix G.2 explicitly corrects the earlier description that all probe calls used temperature 0.

Key Experimental Results

Main Results

The following selection from the paperโ€™s Table 3 retains its family-mean row order. Conditional collapse is the family mean of model-specific conditional rates, not a ratio recomputed by pooling all family questions.

Family Models Probe total FR Conditional collapse (%)
DeepSeek 1 0.113 0.00
Google 2 0.288 0.29
OpenAI 2 0.346 1.52
Anthropic 1 0.450 2.15
Phi 1 0.469 10.24
Qwen 6 0.699 19.75
Meta 1 0.774 8.46

The primary analysis covers seven families, with Spearman correlation 0.8929 and exact two-sided p=0.0123. The free capability-pressure proxy, one minus initial-majority accuracy, has family correlation 0.821. The family partial after controlling initial-majority accuracy is 0.767 with p=0.0877, below conventional significance requirements. The evidence therefore supports protocol-bound audit triage, not a calibrated capability-adjusted risk predictor.

The next table comes from the paperโ€™s Table 5. Column labels and group sizes are preserved to distinguish per-question diagnosis from family ranking.

Predictor Gemini (N=990) Haiku (N=500) GPT (N=494) Pooled (N=994)
Mean probe FR 0.455 0.626 0.735 0.658
S/A ratio 0.394 0.473 0.576 0.529
Initial disagreement 0.765 0.550 0.858 0.573
Combined (FR+Diff+Disagree) 0.845 0.719 0.931 0.762
Combined (+S/A) 0.843 0.719 0.933 0.765

These AUCs broadly illustrate the value of runtime disagreement and combined features, but disagreement does not outperform probe FR in every column: it is lower for Haiku and Pooled. The paperโ€™s stronger-disagreement summary must retain this within-table boundary. Pooled N=994 is also not the sum of the first three column sizes and is not rewritten here as an all-sample aggregate.

Ablation Study

The following table comes from the paperโ€™s Table 7, with cohort and freeze semantics explained in Appendices E.7 and E.8. Delta acc is in percentage points (pp), relative to each cohortโ€™s saved standard-debate outcomes.

Policy Cohort Validation Delta acc (pp) Prevented Lost Net Break-even ratio
Probe-gated freeze 6,525 OSS debates LOMO -1.21 29 108 -79 3.72
Round 1 majority change 1,255 traces Fixed replay -1.04 56 69 -13 1.23
Learned Round 1 stump 1,255 traces Strict LOMO -2.31 2 31 -29 15.50

The probe-replay cohort contains 251 collapses and 804 corrections. Its gate prevents 29 of the former while losing 108 of the latter; presenting only the first count would reverse the interpretation. The fixed Round 1 rule can have positive weighted utility when collapse is valued sufficiently highly, but remains negative at equal weights, and this is not evidence of an online deployment benefit.

Key Findings

  • Onset analysis covers 6,925 debates and 253 collapses: Round 1 contributes 149 (58.9%), Round 2 contributes 47 (18.6%), and Round 3 contributes 57 (22.5%). The first round is the largest risk window, but a substantial late-onset tail remains.
  • Table 2 and Appendix E.1 report out-of-fold AUC increasing from 0.669 to 0.768 with Round 1 features and to 0.846 with full trajectories. The final result includes subsequent states and is a diagnostic reference only.
  • In the open-subset audit of Appendix A.3.3, alphabetic tie-breaking fires in 298/3,240 initial/final voting reductions (9.20%). Removing the last-letter fallback changes four-cell membership in 51 of 966 rows that remain strictly labelable (5.28%). Rule-based labels are not inherently error-free.
  • In the GPQA stress check, Mistral Small 4 produces 20 collapses and 21 corrections, appearing nearly accuracy-neutral. Excluding four rows without a valid initial majority changes the net from +1 to -1. This supports transition accounting, not cross-benchmark transfer of screening.

Highlights & Insights

  • Separate risk identification from intervention value. High correlations and better AUC can coexist with a negative-utility policy. Intervention evaluation must count sacrificed recoveries as a cost rather than crediting only errors avoided.
  • Retain conditional denominators and direction. Identical net accuracy gains can hide very different collapse burdens. The four-cell ledger distinguishes stable but uncorrective behavior from active bidirectional transitions; it enriches evaluation reports rather than certifying general safety.
  • Make negative results reusable benchmarks. Known equal-weight negative freeze baselines help test whether new controllers genuinely preserve corrections. Comparisons across cohorts or weights still require matched held-out auditing.

Limitations & Future Work

  • Small and uneven family coverage. Primary inference uses only seven families, and conditional collapse spans 5.11% to 43.90% within Qwen. Averaging reduces duplicate counting of related variants but does not remove heterogeneity or capability confounding.
  • Explicit open-rebuild boundaries. The 6,525-debate probe-LOMO decision matrix is represented only by frozen aggregate results, while exact per-debate Round 1 features remain access-restricted. The open materials do not establish reproduction of every row-level control comparison.
  • Parsing and ties are evaluation choices. Main results use alphabetic tie-breaking, whereas the shared-initial-answer follow-up in Appendix G.5 retains ties as undecided and counts them as wrong. Their collapse and accuracy results cannot be mixed directly.
  • Limited mechanism and transfer evidence. Homogeneous, closed-book, static-choice tasks do not represent heterogeneous oversight, tool use, retrieval, or unrestricted generation. Round-level associations also do not establish causal mediation.
  • The self-revision comparator is not a compute-matched solo-reasoning ceiling. Follow-up arms share Round 0; across 12 multiple-choice settings, private revision is 1.44โ€“4.81 percentage points below peer debate under binary accuracy, but differences become -1.62 to +0.38 under expected tie scoring. All 11 reasoning-mode rows still contain both transition directions. These are post-hoc boundary checks, not proof that debate necessarily beats longer single-model reasoning.
  • Preserve source-consistency issues. Text extraction duplicates mathematical renderings such as โ€œ88โ€ and โ€œ253253โ€; this note resolves the probe count to eight using the explicit 4ร—2 design and tables. The shorthand sum in Appendix B.4 must not be read as an unweighted sum of conditional rates. Aggregate AUC descriptions and individual table columns should likewise not be rewritten to manufacture agreement.
  • vs Du et al.โ€™s multi-agent debate. Earlier work emphasizes reasoning and factuality gains. This paper retains the standard scaffold and decomposes average gains into collapse and correction, rather than presenting its measurement protocol as a new collaboration algorithm.
  • vs confidence, diversity, and adaptive-stopping research. Those approaches design aggregation or control mechanisms. This paper supplies a rescoring target: report prevented/lost/net counts on fixed held-out cohorts under declared utility weights, comparing actions rather than diagnostic scores alone.
  • vs sycophancy and conformity research. Behavioral measurement distinguishes argument and social channels, but there are no activation caches or controlled training experiments. The results are neither a circuit replication nor evidence that reluctance to revise implies safety.
  • Potential research direction. Evaluate abstention or external verification that preserves correction opportunities on uncertain cases, comparing parsing rules, action coverage, and weighted transition utility on a unified cohort. This is an implication of the note, not a method validated by the paper.
  • Resources. Paper; DebateLedger project page. This note uses only the supplied local full text; project-release status was not checked online.

Rating

  • Novelty: 4/5 โ€” The contribution is directional accounting and an audit protocol, not a new debate architecture.
  • Experimental Thoroughness: 4/5 โ€” Held-out negative results, parser checks, and scope extensions are substantial, but primary family coverage is limited and some row-level materials are restricted.
  • Writing Quality: 3/5 โ€” Claims are carefully bounded, but cohort versions and main-text/appendix conventions require close reading.
  • Value: 4/5 โ€” Provides a more transparent reporting framework for debate controllers than collapse rate or final accuracy alone.