๐ LLM Evaluation¶
๐ง NeurIPS2026 ยท 5 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (3) ยท ๐ฌ ICLR2026 (131) ยท ๐ฌ ACL2026 (96) ยท ๐งช ICML2026 (40) ยท ๐ค AAAI2026 (16) ยท ๐ง NeurIPS2025 (38)
๐ฅ Top topics: LLM ร4 ยท Reasoning ร2
- LLM Judge Validation Under Sparse Overlap: From Inference to Design
-
This paper treats the quantity and allocation of overlapping annotations for LLM-judge validation as a statistical design problem: variance analysis guides overlap planning, while stratified allocation improves representativeness; low overlap substantially changes deployment decisions and judge rankings on real benchmarks, but stratification can increase false approval even as it reduces false rejection.
- OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models
-
OTROPE uses optimal transport in semantic embedding space to reweight human-labeled residuals from an older model and correct a newer model's mean proxy evaluation, reducing cross-model preference-win-rate estimation error without response likelihoods, but its effectiveness depends on representation coverage and residual transferability, and it does not outperform PPI under every shift.
- SciR: A Controllable Benchmark for Scientific Reasoning in LLMs
-
SciR generates verifiable deduction, induction, and causal-discovery problems before rewriting their premises as multi-document scientific discourse, separately varying inference complexity and premise obfuscation to diagnose LLMs; document formalisation remains a substantial bottleneck even with symbolic solvers.
- Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning
-
OracleLadder fixes teacher-written mathematical sub-goals and crosses isolated milestone tests with increasing roadmap and answer assistance on the parent problem: across six models, 33โ48% of a screened 354-problem NuminaMath set falls into a composition gap, where all milestones are solvable but the parent remains unrecovered; the share is 24โ37% after excluding grading issues flagged by automated review, but this is not a pure causal measurement of an internal composition ability.
- SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation
-
SpanUQ distills claim consistency across offline sampled responses into a probe over frozen LLM hidden states, using set prediction to locate semantic spans and assign continuous uncertainty; five separately trained backbone probes achieve 0.908โ0.944 AUROC and support finer-grained risk filtering than rejecting entire responses.