Skip to content

๐Ÿ’ก LLM Reasoning

๐Ÿง  NeurIPS2026 ยท 8 paper notes

๐Ÿ“Œ Same area in other venues: ๐Ÿ“ท CVPR2026 (16) ยท ๐Ÿ”ฌ ICLR2026 (241) ยท ๐Ÿ’ฌ ACL2026 (82) ยท ๐Ÿงช ICML2026 (78) ยท ๐Ÿค– AAAI2026 (37) ยท ๐Ÿง  NeurIPS2025 (82)

๐Ÿ”ฅ Top topics: Reasoning ร—6 ยท LLM ร—3

Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization

PEPO weights token updates in successful reasoning trajectories by uncertainty relative to their neighborhoods while preserving each trajectory's total weight, outperforming the compared GRPO and global-entropy baselines in the main mathematical reasoning experiments, with explicit limits on trend robustness and cross-task generalization.

Diversity Combining for Multi-Path LLM Reasoning

The paper explains diminishing returns in multi-path reasoning through correctness correlation and the design effect, then estimates a fixed deployment budget from a labeled offline four-path pilot; across five modelโ€“task configurations, 4โ€“10 paths retain 96%โ€“103% of the binary majority-vote accuracy at 32 paths, although that evaluation metric is not the actual plurality-vote accuracy for open answers.

Externalized CPDAG Summaries Improve LLM Causal Deduction

Structured Thinking asks the same large language model to produce a format-constrained CPDAG summary before checking whether a causal hypothesis holds in every compatible DAG, raising Qwen3.5-27B primary-seed F1(YES) on Corr2Cause from 73.01 to 86.36 without formally guaranteeing graph validity or reasoning over the entire equivalence class.

Provable Test-Time Scaling for Beam Search in LLM Reasoning

With sampling access to next tokens but no full logits, this paper filters low-frequency tokens before beam pruning and, under alignment between prefix likelihood and correctness, reduces the worst-step coverage dependence of search samples from worst-case quadratic to nearly linear; accuracy on an LLM matrix-multiplication task increases from 37.2% to 38.8%.

Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning

By evaluating answer correctness separately from the continued production of complete explicit reasoning, this paper shows that ordinary fine-tuning without reasoning traces can improve Chemistry accuracy while reducing valid reasoning to approximately zero, and that reasoning-region loss masking mitigates this structural degradation, with model- and task-dependent effects.

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

SAGE uses two soft potentialsโ€”whether an operation addresses unresolved constraints and whether its resulting state approaches a target structureโ€”to guide both candidate sampling and rewards during post-training, improving long-horizon reasoning while retaining ordinary inference-time decoding; Qwen3.5-35B mathematical average accuracy rises from 61.92% with GRPO to 64.86%.

Structured Sparse Memory for Recurrent Reasoning

CoSE replaces a large task table with a factorized FiLM conditioning branch and rank-32 instance residuals, reducing task-memory parameters to roughly 1/15 while improving pass@2 in controlled ARC experiments; CHARM combines this memory with synthetic data, recurrent computation, and extended training to reach 84.0% / 46.7% on ARC-AGI-1 / 2 public evaluation, which cannot be attributed entirely to CoSE.

Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning

This diagnostic study charges diffusion generation and reward scoring to a shared inference budget, finds that deterministic top-1 PRM guidance on Dream-7B trails independent sampling with a task-matched ORM, and locates the failures through candidate-pool, terminal-scoring, and readout controls.