Skip to content

๐Ÿ’ป Code Intelligence

๐Ÿง  NeurIPS2026 ยท 5 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (2) ยท ๐Ÿ”ฌ ICLR2026 (59) ยท ๐Ÿ’ฌ ACL2026 (49) ยท ๐Ÿงช ICML2026 (22) ยท ๐Ÿค– AAAI2026 (10) ยท ๐Ÿง  NeurIPS2025 (19)

CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

CodeScaler distills high-quality execution feedback into a code reward model, uses strict code extraction and reward shaping to support reinforcement learning and candidate reranking without online execution, and exceeds RLVR by 1.55/4.23 percentage points in average Avg@8 across four benchmarks for 8B/14B policies trained on DeepCoder.

Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents

CRR restores a small number of decision states during software-agent training and scores realized actions using the terminal-return difference between the original trajectory and an alternative-action continuation, raising SWE-bench Verified pass@1 from extended GRPO's 36.7% to 41.7% within the same 40.125-hour training window on an eight-H100 node, without eliminating fork execution costs.

CTE-Bench: Counterfactual Trace Evaluation for Stateful Software Simulators

CTE-Bench executes software with and without an intervention to evaluate sustained response prediction on fixed future calls: effect-step accuracy reaches 54.3%โ€“61.5% with correct earlier answers, falls to 24.8%โ€“33.2% in free rollout, and whole-trace exact match peaks at just 1.2%.

Execution Guided Line-by-Line Code Generation

At inference time, EG-CFG previews and executes several short continuations of a code prefix, injects runtime traces into the prompt, and generates code token by token with dual-distribution classifier-free guidance while refreshing feedback at line boundaries; with DeepSeek-V3-0324, the paper reports 96.6% / 99.4% / 69.9% accuracy on MBPP / HumanEval / DS-1000, respectively, but requires executable provided tests and additional search computation.

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation

MemDocAgent combines a traversal order constrained by dependencies and directory hierarchy with a readable, writable, and verifiable RepoMemory to let one agent continuously produce component-, module-, and repository-level documentation; across 20 Python repositories, its Qwen3-Coder configuration achieves 0.979 completeness, 0.916 truthfulness, and 0.690 helpfulness, but its coverage-preserving fallback does not guarantee that every document is trustworthy.