๐ก LLM Reasoning¶
๐ง NeurIPS2026 ยท 8 paper notes
๐ Same area in other venues: ๐ท CVPR2026 (16) ยท ๐ฌ ICLR2026 (241) ยท ๐ฌ ACL2026 (82) ยท ๐งช ICML2026 (78) ยท ๐ค AAAI2026 (37) ยท ๐ง NeurIPS2025 (82)
๐ฅ Top topics: Reasoning ร6 ยท LLM ร3
- Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization
-
PEPO weights token updates in successful reasoning trajectories by uncertainty relative to their neighborhoods while preserving each trajectory's total weight, outperforming the compared GRPO and global-entropy baselines in the main mathematical reasoning experiments, with explicit limits on trend robustness and cross-task generalization.
- Diversity Combining for Multi-Path LLM Reasoning
-
The paper explains diminishing returns in multi-path reasoning through correctness correlation and the design effect, then estimates a fixed deployment budget from a labeled offline four-path pilot; across five modelโtask configurations, 4โ10 paths retain 96%โ103% of the binary majority-vote accuracy at 32 paths, although that evaluation metric is not the actual plurality-vote accuracy for open answers.
- Externalized CPDAG Summaries Improve LLM Causal Deduction
-
Structured Thinking asks the same large language model to produce a format-constrained CPDAG summary before checking whether a causal hypothesis holds in every compatible DAG, raising Qwen3.5-27B primary-seed F1(YES) on Corr2Cause from 73.01 to 86.36 without formally guaranteeing graph validity or reasoning over the entire equivalence class.
- Provable Test-Time Scaling for Beam Search in LLM Reasoning
-
With sampling access to next tokens but no full logits, this paper filters low-frequency tokens before beam pruning and, under alignment between prefix likelihood and correctness, reduces the worst-step coverage dependence of search samples from worst-case quadratic to nearly linear; accuracy on an LLM matrix-multiplication task increases from 37.2% to 38.8%.
- Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning
-
By evaluating answer correctness separately from the continued production of complete explicit reasoning, this paper shows that ordinary fine-tuning without reasoning traces can improve Chemistry accuracy while reducing valid reasoning to approximately zero, and that reasoning-region loss masking mitigates this structural degradation, with model- and task-dependent effects.
- SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
-
SAGE uses two soft potentialsโwhether an operation addresses unresolved constraints and whether its resulting state approaches a target structureโto guide both candidate sampling and rewards during post-training, improving long-horizon reasoning while retaining ordinary inference-time decoding; Qwen3.5-35B mathematical average accuracy rises from 61.92% with GRPO to 64.86%.
- Structured Sparse Memory for Recurrent Reasoning
-
CoSE replaces a large task table with a factorized FiLM conditioning branch and rank-32 instance residuals, reducing task-memory parameters to roughly 1/15 while improving pass@2 in controlled ARC experiments; CHARM combines this memory with synthetic data, recurrent computation, and extended training to reach 84.0% / 46.7% on ARC-AGI-1 / 2 public evaluation, which cannot be attributed entirely to CoSE.
- Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning
-
This diagnostic study charges diffusion generation and reward scoring to a shared inference budget, finds that deterministic top-1 PRM guidance on Dream-7B trails independent sampling with a task-matched ORM, and locates the failures through candidate-pool, terminal-scoring, and readout controls.