Skip to content

The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining

Conference: NeurIPS2026
arXiv: 2609.32964
Code: https://github.com/vnht/commit-abstain-circuit
Area: Hallucination / Interpretability
Keywords: epistemic abstention, unsupported commitment, causal gating, sparse circuit, policy readout

TL;DR

The paper localises a sparse Commit-Abstain Circuit (CAC) using the pre-generation commit–abstain logit margin, identifies early commitment accumulation with insufficient late abstention correction, and trains a lightweight policy on component contributions that raises reported mean decision accuracy from 0.692 to 0.814, rather than demonstrating improved answer-content accuracy.

Background & Motivation

A language model can produce a substantive answer despite insufficient information, but this does not necessarily mean that its internal states contain no signal to abstain. Existing hidden-state probes recover correctness, truthfulness, or unanswerability information, whereas semantic entropy, multi-model agreement, and abstention training primarily address outputs or behaviour. A component-level explanation is missing between these perspectives: if relevant information is already present, why does the final output still cross the commitment threshold? The paper narrows its target to unsupported commitment—failure to abstain on an unanswerable input—rather than all factual errors.

Abstention here is epistemic abstention: deciding whether available knowledge or context is sufficient to answer. It differs from safety refusal based on the nature of a request and from giving an incorrect answer to an answerable question. KUQ concerns questions without definitive answers, SQuAD 2.0 concerns questions unsupported by a supplied passage, and MuSiQue removes supporting documents to create gaps in reasoning chains. These settings make whether to answer a separate binary prediction target.

The authors measure preference before the first output token to avoid mixing the initial decision with path dependence introduced by autoregressive generation. They then jointly intervene on attention heads and MLP sublayers rather than merely asking whether a direction is detectable. Core Idea: preserve individual component contributions to the commit–abstain decision instead of reading only their scalar sum, thereby recovering abstention information masked by the final commitment preference.

Method

Overall Architecture

The input is a question with optional context, and the supervision is an answerability label. The analysis proceeds through Margin Attribution, Causal Gating, and Functional Intervention, producing a model-specific CAC. The policy stage extracts CAC features from an ungated pre-generation forward pass and predicts whether to commit or abstain.

These stages use different computations. Margin Attribution uses a linear decomposition with the final normalisation factor frozen; gating and functional interventions modify the actual forward pass and retain downstream nonlinearities. The final policy is an external lightweight classifier: it neither updates the base language model nor requires the localisation gates to remain active at deployment.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    X["Question and optional context"] --> A["Margin Attribution"]
    A -->|Training data and answerability labels| B["Causal Gating"]
    B -->|Localised components| C["Functional Intervention"]
    C -->|Analytical validation, not a required inference step| D["Component-Level Policy Readout"]
    A -->|Inference: ungated forward contributions| D
    B -->|Fixed CAC component set| D
    Y["Training supervision: answerability labels"] -.-> D
    D --> O["Commit or abstain"]

Key Designs

1. Margin Attribution: decompose first-step preference into traceable component contributions

The authors construct an abstention-start token set \(\mathcal{A}\), with the remaining vocabulary forming the commitment set \(\mathcal{C}\). They collect the first tokens of abstaining responses on 1,000 unanswerable instances disjoint from training and testing, pool them across models, disambiguate them, and obtain annotator validation. The main setting retains 17 tokens. This is a lexical operationalisation, not a complete semantic judgement of arbitrary responses; a token frequently starting abstention does not imply that every continuation is an abstention.

At the final input position, the highest-logit token is selected from each set, defining the preference margin:

\[ \Delta(x)=z_{c^{*}(x)}(x)-z_{a^{*}(x)}(x),\qquad c^{*}(x)=\arg\max_{t\in\mathcal{C}}z_t(x),\quad a^{*}(x)=\arg\max_{t\in\mathcal{A}}z_t(x). \]

\(\Delta(x)\geq0\) denotes a preference to commit, while a negative margin denotes a preference to abstain. It compares the strongest first-step candidate in each set, not their total probability mass or the probability that an answer is correct. Mean agreement with full greedy-generation behaviour is 97.6%, supporting an approximate readout while leaving exceptions such as mid-generation reversals.

The residual stream sums attention-head and MLP-sublayer outputs together with embedding and other terms. The authors freeze the final normalisation factor at its ungated-forward value and project each component output onto the difference between the two selected token readout directions:

\[ \Delta(x)=\kappa(x)+\sum_{k=1}^{K}v_k(x),\qquad v_k(x)=d(x)^{\top}c_k(x),\quad d(x)=u_{c^{*}(x)}-u_{a^{*}(x)}. \]

Here \(c_k(x)\) is a component output, \(u_t\) is a readout vector incorporating the example-specific frozen normalisation factor, and \(\kappa(x)\) collects embedding, bias, and related terms. Under this attribution convention, positive contributions raise the margin and negative contributions lower it. The decomposition explains the final ungated margin; it is not an exact prediction of behaviour after removing a component, which changes downstream activations and normalisation.

2. Causal Gating: identify decision-relevant components under opposing regularisation pressures

Removing every component independently is expensive and can miss interactions. The authors first screen candidates using the training-set difference in mean contribution between answerable and unanswerable inputs, then jointly train sigmoid gates on candidate attention heads and MLP sublayers. Head gates act before the output projection; MLP gates act before residual addition. Non-candidate components remain open, and base-model weights stay frozen.

The gating objective raises the margin for answerable inputs and lowers it for unanswerable inputs. The highest commitment and abstention tokens used by the loss are obtained from the ungated model and held fixed throughout optimisation, rather than reselected after each gated forward pass:

\[ \mathcal{L}_{\mathrm{task}}(G) =-\frac{1}{B}\sum_{x\in\mathrm{batch}}(2y-1)\Delta(x\mid G),\qquad \mathcal{L}(G;\lambda) =\mathcal{L}_{\mathrm{task}}(G)-\lambda\sum_{g\in\mathcal{S}}\mathrm{clip}(g,-C,C). \]

\(y=1\) means answerable, and \(G=\sigma(g)\). Task optimisation with \(\lambda=0\) provides a shared initialisation. Retention pressure then pushes gates open, while removal pressure pushes them closed, yielding \(G^{+}\) and \(G^{-}\). A component remaining at \(G^{-}>0.5\) against removal pressure is a c-component; one remaining at \(G^{+}<0.5\) against retention pressure is an a-component. Role assignment also uses cross-seed mean gate values, with final membership requiring the same role in at least 8/10 seeds.

The a/c names describe differential roles relative to answerability supervision, not a rule that every a-component contribution is negative and every c-component contribution positive. Appendix D.1 gives an example with mean contributions of +0.3 on answerable inputs and +1.5 on unanswerable inputs. Both are positive, yet the difference of −1.2 can produce an a-role because the component raises commitment more strongly where the model should abstain. The differential-gradient explanation is a local attribution intuition: joint optimisation changes downstream computation, so the full model cannot be treated as independent linear components.

3. Functional Intervention: distinguish threshold correction from ranking structure

After localisation, the authors progressively remove a- or c-components in causal-score order, comparing against 20 size-matched random non-CAC ablations. They also amplify component outputs to examine false abstention and false commitment. These experiments use actual intervened forward passes, not simply subtraction or scaling of terms in the frozen attribution equation.

The central result contradicts a literal reading of the role names: removing c-components also causes more abstain-to-commit flips. The authors interpret c-components as supporting both commitment and the ranking separation between answerable and unanswerable inputs; their removal damages the discriminative structure. By contrast, a-components more often act as corrections that move borderline inputs across zero, with smaller ranking changes after removal. Directional flips and AUROC must therefore be reported together rather than assigning roles from the overall abstention rate alone.

Depth analysis shows both trajectories accumulating positive contributions until approximately 80% depth. True-abstention trajectories then turn sharply negative, whereas unsupported-commitment trajectories are not sufficiently corrected. This motivates the accumulate-yet-undercorrect pattern; it does not establish an identical temporal circuit for every example and model.

Amplification also has nonlinear limits. Moderate c-component amplification increases commitment and reduces false abstention; moderate a-component amplification increases abstention and reduces false commitment, each introducing the opposite error trade-off. Under extreme scaling, RMSNorm makes the normalised representation approach the component direction instead of increasing its magnitude without bound. False commitment under a-component amplification still reaches a nonzero plateau, so amplifying an abstention component sufficiently is not equivalent to a reliable abstention policy.

4. Component-Level Policy Readout: retain answerability information before summation

The final policy reads each CAC component's margin contribution at the last input position, appending the mean, standard deviation, minimum, maximum, and summed contributions of a- and c-components. Model-specific CACs contain 16–64 components, with these statistics added to the feature dimension. Rather than thresholding another total sum, the classifier can distinguish a large positive outlier masking several negative signals from broad agreement in favour of commitment.

The MLP architecture is \(d_{\mathrm{in}}\to32\to128\to64\to32\to1\). A 32-dimensional bottleneck standardises different model-specific feature sizes before nonlinear mixing. The output is \(P(\mathrm{answerable})\), with abstention triggered below a validation-calibrated threshold. This is a supervised answerability prediction, not a guarantee that the language model's particular answer is correct.

Logistic regression on the same features also improves mean decision accuracy by +10.6 percentage points, compared with +12.2 for the MLP. Their 1.6-point difference suggests that component-level information and a revised readout account for most of the gain, rather than classifier complexity alone.

A Worked Example

The KUQ example in Appendix F.1 is labelled unanswerable and evaluated on Llama 3B. Its overall margin is +7.94, so Zero-Threshold commits. L19.MLP contributes +8.27, while several other CAC components contribute negatively.

The authors report a negative-contribution sum of −8.09, a positive-contribution sum of +12.62, and a net CAC contribution of +4.53. Reading only the positive final margin loses the structure of one strong positive term overwhelming other signals; the component-level policy produces a readout of 0.089 and correctly abstains. This illustrates the use of contribution patterns, not a rule that any large positive contribution implies unanswerability or a judgement of the generated answer's medical correctness.

Loss & Training

Each main dataset supplies 1,000 sampled instances. Supervised experiments use a 50/50 training–test split, with a validation subset held out from training to calibrate the policy threshold. The main analysis covers ten instruction-tuned models from five families at 3B–14B. Full-generation behaviour is checked with greedy decoding and at most 256 new tokens.

Gating uses Adam with batch size 16 and 100 steps per phase. The learning rate decays from 0.1 to 0.01 in the first phase and from 0.5 to 0.1 in the two pressure phases; the clipping bound is 4. Regularisation strength follows the scale of the ungated training-set margin, and results aggregate ten independent seeds. Candidate screening is correlational, while subsequent joint gating supplies intervention evidence; these are not interchangeable.

The policy MLP uses batch normalisation, ReLU, and dropout 0.15 in hidden layers. It is trained with binary cross-entropy and AdamW, with learning rate and weight decay both \(10^{-3}\), for up to 500 epochs and early-stopping patience 120. Localisation and policy training are repeated for each model, rather than sharing gates or classifier weights across models.

Key Experimental Results

Main Results

The table below selects results from the paper's Table 2. Each model row averages KUQ, SQuAD 2.0, and MuSiQue test results. Acc is commit/abstain decision accuracy, not answer-content accuracy; AUROC measures answerability ranking.

Model Zero-Threshold AUROC CAC MLP AUROC Zero-Threshold Acc CAC MLP Acc
Qwen 3.5 4B 0.725 0.879 0.655 0.829
Qwen 3.5 9B 0.798 0.911 0.697 0.855
Gemma 3 4B 0.722 0.807 0.644 0.765
Gemma 3 12B 0.854 0.889 0.770 0.845
Llama 3B 0.689 0.807 0.623 0.755
Llama 8B 0.784 0.874 0.651 0.820
Ministral 3B 0.790 0.874 0.714 0.815
Ministral 14B 0.848 0.877 0.753 0.829
Phi-4 mini 0.780 0.837 0.705 0.785
Phi-4 14B 0.830 0.888 0.711 0.840
All-configuration mean, §5.2 0.782 0.864 0.692 0.814

Decision accuracy is the overall proportion of answerable instances assigned commitment and unanswerable instances assigned abstention. False abstention (FA) is the proportion of answerable inputs assigned abstention; false commitment (FC) is the proportion of unanswerable inputs assigned commitment. Neither is an exact-match error rate for generated answers. Policy evaluation reports mean FA decreasing from 0.320 to 0.127, approximately a 2.5-fold reduction.

The margin analysis in §4.2 and Appendix C, Table 8 reports mean AUROC 0.811, whereas the policy comparison in §5.2 reports 0.782 for Zero-Threshold. These belong to different analysis/evaluation summaries, and the paper does not fully explain the numerical discrepancy. Substituting 0.811 into the policy table and recomputing the reported +8.2 AUROC-point gain would be inappropriate.

Ablation Study

The table summarises §4.5 and Appendix D.5, Table 13. Flip percentages use all 1,500 held-out instances as the denominator. A→C means abstain-to-commit and C→A means commit-to-abstain; AUROC changes are cross-model means reported in §4.5.

Intervention A→C C→A Mean AUROC change Interpretation
Remove all a-components 17.1% 7.8% −0.03 More borderline decisions cross zero; ranking remains relatively stable
Remove all c-components 31.7% 3.4% −0.17 Also biases flips towards commitment, but substantially damages discrimination

The paper also reports Qwen 4B AUROC falling from 0.73 to 0.40 after c-component ablation; size-matched random ablation changes AUROC by −0.069 on average. This supports CAC specificity but does not mean that every CAC component's complete semantic function has been established.

The following selection from Appendix E.3, Table 18 tests whether the policy architecture is essential. Means aggregate ten models.

Dataset Zero-Threshold Acc CAC Logistic Regression Acc CAC MLP Acc
KUQ 0.707 0.887 0.923
SQuAD 2.0 0.665 0.743 0.745
MuSiQue 0.704 0.764 0.774
All datasets 0.692 0.798 0.814

Key Findings

  • CACs are localised in all ten main models, covering 2.0%–11.8% of components with median 5.2%; MLP sublayers account for approximately 25% on average. This supports extending gating beyond attention heads rather than attributing abstention to a single head.
  • The MLP's advantage over logistic regression is concentrated on KUQ, with a mean gap of 3.6 percentage points; the gap is 0.2 points on SQuAD 2.0 and 1.0 on MuSiQue. Component-level linear readout captures much of the improvement on context-based tasks.
  • On unseen HotpotQA and SelfAware, the policy outperforms Zero-Threshold in all 24 model–dataset configurations, with mean decision-accuracy gain 6.9 points. Gemma 3 27B and Qwen 3.5 35B are first re-localised and trained on in-distribution data, then transferred zero-shot across datasets. The four larger-model OOD configurations gain 6.3 points on average; small-model weights are not transferred zero-shot.
  • The depth claim should be understood as an aggregate tendency. The main text says that every model has more than two c-components per a-component at 60%–70% depth, but Table 9 gives 1.5/1.4 for Qwen 9B and 1.4/1.3 for Gemma 12B. The table supports cumulative ratios above 1 in every model, not above 2 in every model.

Highlights & Insights

  • Separating information availability from its use in decision control is the central insight. Answerability signals recoverable from component contributions can be masked by residual aggregation, so optimising only the total margin may not recover the correct decision.
  • The analysis preserves the counterintuitive finding that removing commitment-labelled components can increase commitment. Distinguishing ranking structure from zero-threshold calibration avoids treating role names as causal evidence.
  • The policy reads scalar projections from a small component set rather than full hidden vectors or multiple generations. This suggests lightweight deployment, although offline localisation, the base-model forward pass, and activation hooks still incur costs.
  • Amplification reveals a control boundary imposed by normalisation. Moderate intervention gains do not imply unlimited correction at extreme magnitudes; training a policy has stronger support than blindly amplifying abstention signals.

Limitations & Future Work

  • The scope is unsupported commitment on unanswerable inputs, not incorrect answers to answerable questions. KUQ's open problems, SQuAD 2.0's passage support, and MuSiQue's missing documents represent particular unanswerability mechanisms rather than comprehensive real-world fact checking.
  • Both the abstention vocabulary and full-generation labels depend on lexical expression, potentially missing indirect abstention or mixed responses. Robustness to the 17-token set and its expansions does not remove this operationalisation limit; semantic counterexamples remain necessary.
  • Fixed top tokens and frozen final-normalisation factors make attribution tractable but do not fully characterise generation dynamics. Behavioural agreement of 97.6% is not 100% and does not replace mechanistic analysis of multi-turn or mid-generation decisions.
  • The CAC is a group-level causal localisation obtained after correlational screening, not a complete functional map of every head and MLP. Component interactions, downstream changes induced by gates, and screening omissions require finer mechanistic tests.
  • Reporting includes unclear boundaries and small conflicts: analysis AUROC 0.811 should not be mixed with policy-baseline 0.782; Appendix F.2 gives margin +1.37 in prose but +1.38 in Table 20. Appendix D.4 also describes a shared remainder across inputs, not fully consistent with §3's \(\kappa(x)\) notation. These should not be silently unified into an allegedly established exact model.
  • Further work could test lexical variation, different abstention costs, and long contexts while jointly reporting answer coverage, answer correctness, and calibration error. Cross-dataset transfer has evidence here; shared policy weights across models do not.
  • vs Causal Head Gating: the paper retains joint component identification under opposing regularisation pressures, extends gates to MLP sublayers, and targets the commit–abstain margin. It is decision-specific mechanism localisation, not a complete replacement for general circuit discovery.
  • vs HaMI / HaloScope: these methods use token hidden states or hidden subspaces to detect hallucination-related information. CAC first localises components with intervention evidence and then uses their contributions as features. The MLP classifier is not the main novelty; mechanistically grounded feature selection is.
  • vs Semantic Entropy / Multi-LLM: output-level approaches use multiple samples or models, whereas the CAC policy reads internal signals before generation. Reduced generation requirements come with white-box access and model-specific localisation, preventing unconditional claims of overall superiority.
  • Research implication: future experiments could examine whether training brings corrective components earlier or routes their information more effectively into the final readout. This is a hypothesis motivated by the depth pattern, not proof that changing timing reliably reduces all hallucinations.

Rating

  • Novelty: 4/5 — Connects answerability representations with component-level decision control; the policy itself is lightweight.
  • Experimental Thoroughness: 4/5 — Broad model families, intervention controls, and dataset transfer, but limited scope and reporting ambiguities remain.
  • Writing Quality: 3/5 — Clear mechanistic narrative, with component naming, depth ratios, and some numbers requiring careful interpretation.
  • Value: 4/5 — Provides an actionable mechanistic readout for epistemic abstention, not a general guarantee of factual correctness.