Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning¶
Conference: NeurIPS2026 (task-queue assignment; this note follows arXiv v2)
arXiv: 2605.21127
Code: https://github.com/itsluketwist/thinkpack/
Area: LLM Reasoning
Keywords: reasoning-trace collapse, structural validity, supervised fine-tuning, loss masking, conditional evaluation
TL;DR¶
By evaluating answer correctness separately from the continued production of complete explicit reasoning, this paper shows that ordinary fine-tuning without reasoning traces can improve Chemistry accuracy while reducing valid reasoning to approximately zero, and that reasoning-region loss masking mitigates this structural degradation, with model- and task-dependent effects.
Background & Motivation¶
Explicit reasoning models typically learn to produce a chain-of-thought (CoT) before a final answer through reinforcement learning, distillation, or dedicated format supervision. Yet downstream adaptation of small models often uses ordinary instruction–response data: examples contain answers, sometimes with explanations, but lack the separate reasoning region expected by the model. The mismatch is not that the data contains no reasoning-related content whatsoever; rather, its training targets do not preserve the previous reasoning-output protocol. A model can therefore adapt to the new task while learning to bypass that protocol.
Answer-only pass@1 can miss this change. An explanation without reasoning delimiters may still provide a correct answer; conversely, a long but unclosed reasoning trace may leave no usable final answer. Compressing both outcomes into accuracy conflates declining task performance among complete traces with a declining frequency of complete traces. Existing step verification, logical evaluation, and faithfulness analyses concern trace content and usually assume that a trace exists. This paper adds an earlier structural check, not another metric proving that the model genuinely reasons.
The authors combine model-specific chat templates, trace parsing, and control over training loss into a reusable evaluation pipeline. Core idea: jointly report trace structure, final-answer accuracy, and accuracy conditioned on valid traces, while masking the training loss on empty reasoning regions to avoid directly teaching reasoning omission as the target behavior.
Method¶
Overall Architecture¶
This is an evaluation and controlled fine-tuning study, not a new reasoning network. Its inputs are an existing explicit reasoning model and instruction–response data without separate reasoning traces. Missing-Reasoning Controls first vary data formatting and supervised regions; checkpoints then generate responses periodically, Cross-Format Structural Parsing separates traces from answers, and Conditional Joint Evaluation measures both structure and task success.
ThinkPack connects chat templating, parsing, statistics, and masking to Hugging Face transformers workflows. Training masks determine only which tokens contribute to the loss; they do not force trace generation at inference time. Inference still uses each model's template, greedy decoding, and a shared budget. Trace retention is therefore an observed post-training behavior, not content supplied by an external generator.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Model and trace-free data"] --> B["Missing-Reasoning Controls"]
B -->|Training: formatting and loss supervision| C["Fine-tuning checkpoints"]
C -->|Inference: fixed examples and greedy decoding| D["Generated responses"]
D --> E["Cross-Format Structural Parsing"]
E --> F["Conditional Joint Evaluation"]
Key Designs¶
1. Missing-Reasoning Controls: include both data representation and supervision in the experiment
When ordinary responses lack explicit traces, the template must still represent that absence. The empty-think format inserts an empty <think>...</think> block before the answer; no-think omits the reasoning block entirely. Qwen3-8B defaults to the former, whereas Olmo-3-7B, Llama-R1-8B, and Nemotron-7B default to the latter. Every model is tested under both formats rather than assuming that its default is safe. An empty block demonstrates immediate reasoning closure, while omission demonstrates bypassing the reasoning region; these signals may affect differently post-trained models differently.
Mitigation experiments retain empty-think data but change the support of the training loss. Masked-think excludes the empty reasoning region while retaining the remaining supervision; response-only additionally excludes the prompt, so only the final response contributes to the loss. Here, “final response” means the dataset's explanation plus answer, not merely its last option or number. Masking does not remove the input: prompts and template structure still condition prediction, but masked tokens are not directly included as prediction targets. This preserves existing behavior by removing an opposing training signal, not by supplying correct CoT examples.
Distillation provides a more expensive control. GPT-5-mini receives both the question and target response and produces a concise reasoning summary, with up to two attempts to extract a non-empty trace. Of 2,674 training examples, 2,408 receive traces, or 90.1%; the remaining 266 are retained using each model's default missing-reasoning format. This is neither an idealized dataset with teacher traces for every example nor trace supervision independently verified step by step.
2. Cross-Format Structural Parsing: separate failures into actionable categories
ThinkPack uses the tokenizer and chat template to identify inline reasoning, reasoning whose opening delimiter is prefixed by the template, and reasoning stored in a separate field; users can override the format definition when automatic detection fails. This adaptation matters for Olmo: its opening delimiter is already supplied by the template, so a generation lacking another opening delimiter should not automatically be classified as missing. Parsing must account for the expected format when deciding whether reasoning can be separated from the final answer.
A structurally valid trace must be complete, non-empty, and reliably extractable. Empty means delimiters exist but contain no content; missing means no trace can be extracted; truncated means reasoning begins but remains incomplete, for example because its closing delimiter is absent or generation reaches the limit. VR, ER, MR, and TR are the proportions of all generations in these four categories. Structural validity indicates compliance with an output protocol, not step correctness, faithful reflection of internal computation, or a causal role in producing the answer.
Truncation is further decomposed into overflow and early stopping: the former exhausts the available generation allowance, while the latter emits an end-of-sequence token before closure. Both break the structure but suggest different interventions. A larger budget may help overflow, but it does not complete traces that have already terminated early. High TR should therefore not automatically be attributed to excessively long reasoning.
The authors manually inspect 100 generations across models, tasks, formats, and checkpoints, covering all four structural categories, and obtain 100% agreement with the parser. This supports parsing reliability for the formats studied, not error-free operation on arbitrary new models, and not semantic correctness of the reasoning.
3. Conditional Joint Evaluation: expose structural retention and task performance separately
Each checkpoint reports overall pass@1 and Rpass@1 conditioned on valid traces. Let \(V_i\) indicate structural validity for response \(i\) and \(C_i\) indicate answer correctness. The following expresses the metric definitions mathematically; it is not a new optimization objective:
ER, MR, and TR use the same denominator \(N\), replacing the validity indicator with the indicator for the corresponding failure category. The authors suppress Rpass@1 when there are 10 or fewer valid responses; “—” does not mean zero accuracy. The two accuracy metrics have different denominators, and overall pass@1 is not simply VR multiplied by Rpass@1, because structurally invalid responses may still answer correctly.
When both VR and Rpass@1 fall, structural retention and task performance among valid traces are both deteriorating. When VR falls sharply but Rpass@1 remains high, the remaining complete-trace responses are still frequently correct. The paper treats the latter as a characteristic signature of collapse, but it does not show that latent reasoning ability remains unchanged on all examples: the conditioning set changes during fine-tuning and may preferentially retain easier or more familiar problems.
The appendix additionally tracks Qwen3-8B's valid-and-correct GSM8K set. Across the final five checkpoints of seed 42, VR falls from 65.2% to 60.5%, the set shrinks from 162 to 152 examples, and Rpass@1 stays at 96.8%–98.1%; consecutive-checkpoint Jaccard overlap is 0.87–0.90. High conditional accuracy in this case is therefore not explained by dramatic set reshuffling, but conditional selection remains, and the analysis does not establish the traces' causal contribution to answers.
Loss & Training¶
The following masked cross-entropy summarizes the training mechanism. It explains ordinary autoregressive supervision with the paper's masking operation, rather than presenting a numbered equation quoted from the source:
Masked positions have \(m_t=0\) and contribute no direct token loss; the others have \(m_t=1\). Masked-think and response-only differ in whether the prompt region is also excluded, not in online sampling of CoT during training. Because the final response is still supervised in the context of an empty block, masking cannot guarantee trace preservation; effectiveness must be measured.
Main experiments use non-quantized LoRA for 3 epochs, with learning rate \(10^{-5}\), AdamW, weight decay 0.01, cosine scheduling, and a warm-up ratio of 0.1. LoRA rank and alpha are both 16, dropout is 0, and adapters target attention and MLP projections. Training uses bf16, per-device batch size 2, gradient accumulation 2, and seeds 42 and 67.
Chemistry contains 2,674 training examples and 507 held-out examples. Evaluation occurs every 100 steps, with the final evaluated checkpoint at step 2000; each task uses a fixed 256-example subset reused across settings. The EvalPlus subset contains 83 HumanEval and 173 MBPP problems, rather than averaging separate full-benchmark HumanEval+ and MBPP+ evaluations.
Evaluation uses no system prompt and greedy decoding, with a minimum generation length of 32 tokens. The shared context limit is 32,768 tokens, and the effective generation allowance subtracts the prompt length; this is not a full 32,768-token output budget for every example. Chemistry and GSM8K request boxed answers, but the checker accepts common equivalent formats; generated code is tested with the EvalPlus harness.
Key Experimental Results¶
Main Results¶
The following extracts Chemistry results from Tables 2 and 3 in Appendix F. Base denotes the pre-fine-tuning model; Final denotes step 2000 averaged over two training seeds. Values are percentages, retaining the source's uncertainty. Default formatting does not necessarily protect reasoning.
| Model | Status / default strategy | pass@1 | VR | Rpass@1 |
|---|---|---|---|---|
| Qwen3-8B | Base | 28.9 ± 5.8 | 73.0 ± 5.7 | 39.6 ± 7.2 |
| Qwen3-8B | Final / empty-think | 54.1 ± 7.0 | 0.0 ± 1.9 | — |
| Olmo-3-7B | Base | 23.8 ± 5.6 | 83.6 ± 5.0 | 28.5 ± 6.4 |
| Olmo-3-7B | Final / no-think | 19.9 ± 6.1 | 59.2 ± 7.0 | 33.6 ± 9.0 |
| Llama-R1-8B | Base | 32.4 ± 6.0 | 57.4 ± 6.1 | 51.7 ± 8.0 |
| Llama-R1-8B | Final / no-think | 54.5 ± 7.0 | 0.2 ± 2.1 | — |
| Nemotron-7B | Base | 34.0 ± 6.0 | 73.0 ± 5.7 | 46.5 ± 7.1 |
| Nemotron-7B | Final / no-think | 53.5 ± 7.0 | 0.0 ± 1.9 | — |
Base intervals are 95% Wilson intervals over 256 examples. Fine-tuned results construct approximate 95% uncertainty intervals by averaging the endpoints of per-seed 97.5% Wilson intervals. They mainly describe evaluation-sample uncertainty, not training-seed standard deviations or full training randomness.
Qwen, Llama, and Nemotron separate successful task adaptation from nearly vanished explicit traces; Olmo instead exhibits declining task performance and increased truncation. The main text summarizes Llama's final Chemistry VR under its default setting as 0%, while Appendix Table 3 gives 0.2%. The precise table value is retained here rather than turning that summary into an exact zero.
Ablation Study¶
The following extracts final GSM8K results from Appendix F, Table 3. It demonstrates both formatting effects and the absence of universally effective masking or distillation; all values are percentages.
| Model | Strategy | pass@1 | VR | Rpass@1 |
|---|---|---|---|---|
| Qwen3-8B | empty-think | 79.5 ± 6.2 | 60.2 ± 7.0 | 98.4 ± 4.2 |
| Qwen3-8B | no-think | 96.3 ± 3.7 | 96.3 ± 3.7 | 97.2 ± 3.5 |
| Qwen3-8B | masked-think | 96.5 ± 3.6 | 99.8 ± 2.1 | 96.7 ± 3.5 |
| Qwen3-8B | distillation | 96.7 ± 3.5 | 99.6 ± 2.2 | 97.1 ± 3.4 |
| Olmo-3-7B | no-think | 41.4 ± 7.0 | 44.9 ± 7.0 | 92.2 ± 7.5 |
| Olmo-3-7B | empty-think | 74.8 ± 6.5 | 79.7 ± 6.2 | 93.9 ± 4.9 |
| Olmo-3-7B | masked-think | 64.8 ± 6.9 | 69.3 ± 6.8 | 93.5 ± 5.4 |
| Olmo-3-7B | distillation | 21.1 ± 6.2 | 21.9 ± 6.3 | 96.4 ± 10.4 |
Qwen's final Chemistry VR also rises from 0.0% under empty-think to 81.4% under masked-think, 83.0% under response-only, and 99.2% under distillation. Corresponding final pass@1 values are 42.8%, 45.1%, and 50.8%. The main text's masking scores of 44%–47% and distillation score of approximately 55% refer to peak performance: exact Appendix Table 4 values are 44.1%, 46.7%, and 54.9%, and should not be confused with final results.
The following comes from Appendix E.4, Table 1, diagnosing Olmo truncation at the final checkpoint using seed 42 only. Overflow is the budget-exhaustion portion of TR; the two columns must not be added.
| Dataset / strategy | TR (%) | Overflow (%) |
|---|---|---|
| Chemistry / no-think | 41.8 | 39.8 |
| GSM8K / no-think | 55.1 | 0.4 |
| GSM8K / distillation | 78.9 | 0.0 |
Key Findings¶
- Similar trace loss can have different mechanisms. Qwen, Llama, and Nemotron mainly produce empty or missing traces; Olmo mainly produces truncated traces. Olmo's GSM8K distillation setting has 78.9% TR but 0.0% overflow, so increasing the budget is not a universal fix.
- Structural retention is not task-performance retention. Llama's final GSM8K VR under empty-think is 100.0%, but pass@1 is 70.1%, below its Base score of 85.9%. VR alone cannot establish that reasoning capability has been fully preserved.
- Collapse is not confined to Chemistry. Single-seed appendix experiments show collapse under standard fine-tuning for Qwen3-14B. When Qwen3-8B is fine-tuned on 2,674-example subsets of OrcaMath or SelfOSS, collapse is often stronger on the corresponding math or code task. These are targeted robustness checks, not full multi-seed replications.
- Seed and subset checks support the trend within limited scope. Across 24 non-distillation combinations for Qwen and Olmo, mean absolute VR differences over training never exceed 3.1 percentage points, and final differences never exceed 6.6 points. Qwen's GSM8K VR decline is 34.0 points on a replacement subset versus 38.3 points on the original subset.
Highlights & Insights¶
- Ask whether traces exist before asking whether they are trustworthy. Structural checks provide a prerequisite for content evaluation and prevent describing a system as preserving explicit reasoning after traces disappear. They are necessary protocol checks, not substitutes for faithfulness or step-correctness evaluation.
- Training format is an experimental variable. Empty blocks and omitted blocks both express missing reasoning but can cause very different model responses. Practical adaptation should jointly validate templates, supervised regions, and models rather than recording only learning rates and LoRA settings.
- Diagnosis is more actionable than a single leaderboard score. ER, MR, and TR suggest different interventions; separating truncation mechanisms further distinguishes budget, termination, and supervision problems. For this setting, masking offers a low-cost starting point without teacher calls.
Limitations & Future Work¶
- Main experiments cover only four 7–8B models, one training domain, non-quantized LoRA, two training seeds, and fixed 256-example evaluation subsets. They do not establish prevalence across all scales, full-parameter fine-tuning, RL, or real deployment traffic.
- Rpass@1 conditions on a post-training selection of valid traces. Set-stability analysis covers only one model–task combination; common-example, difficulty-stratified, and paired comparisons would reduce optimistic interpretations caused by conditional selection.
- GPT-5-mini judges 94% / 92% of 100 valid Base / Final Qwen GSM8K traces correct, respectively, but this is a small model-assisted audit. It establishes neither unchanged semantic quality nor faithful or causally necessary reasoning.
- Parsing depends on chat templates and format conventions, and 100 manually checked examples cannot cover every unusual boundary case. New models, serving frameworks, separate reasoning fields, and irregular delimiters require additional adaptation and validation.
- Lower learning rates delay collapse but weaken in-domain adaptation; a single-seed learning-rate sweep does not establish a universally optimal setting. Future work could jointly evaluate structural retention, answer quality, trace length, and compute cost.
Related Work & Insights¶
- vs CoT prompting and DeepSeek-R1 post-training: these primarily induce or establish explicit reasoning behavior; this paper asks whether existing behavior survives downstream adaptation. Its contribution is retention evaluation and diagnosis, not a stronger reasoning algorithm.
- vs On the Impact of Fine-Tuning on Chain-of-Thought Reasoning: related work examines post-fine-tuning reasoning performance and faithfulness; this study explicitly separates missing, empty, and incomplete traces, adding a failure dimension more basic than semantic quality.
- vs Distilling Step-by-Step! and SCOTT: distillation adds positive supervision through teacher reasoning, while masking removes supervision that encourages empty reasoning. The former incurs generation costs and model-dependent outcomes; the latter is lighter but cannot teach new reliable reasoning.
- vs DSPy and LLM Reasoners: these frameworks emphasize reasoning-pipeline composition and optimization; ThinkPack addresses model-level chat formatting, trace parsing, statistics, and masking. It is infrastructure for adaptation experiments, not a complete serving-orchestration system.
Rating¶
- Novelty: 4/5 — Separating structural trace loss from answer degradation exposes an easily overlooked evaluation blind spot.
- Experimental Thoroughness: 3/5 — Models, formats, mitigations, and diagnostics support one another, but subset size and seed count limit generalization.
- Writing Quality: 4/5 — Definitions and failure categories are clear; rounded main-text values, peaks, and final checkpoints require careful distinction.
- Value: 4/5 — Provides practical acceptance metrics for low-cost adaptation of explicit reasoning models, not certification of trustworthy reasoning.