Skip to content

๐Ÿ“Š LLM Evaluation

๐Ÿง  NeurIPS2026 ยท 5 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (3) ยท ๐Ÿ”ฌ ICLR2026 (131) ยท ๐Ÿ’ฌ ACL2026 (96) ยท ๐Ÿงช ICML2026 (40) ยท ๐Ÿค– AAAI2026 (16) ยท ๐Ÿง  NeurIPS2025 (38)

๐Ÿ”ฅ Top topics: LLM ร—4 ยท Reasoning ร—2

LLM Judge Validation Under Sparse Overlap: From Inference to Design

This paper treats the quantity and allocation of overlapping annotations for LLM-judge validation as a statistical design problem: variance analysis guides overlap planning, while stratified allocation improves representativeness; low overlap substantially changes deployment decisions and judge rankings on real benchmarks, but stratification can increase false approval even as it reduces false rejection.

OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

OTROPE uses optimal transport in semantic embedding space to reweight human-labeled residuals from an older model and correct a newer model's mean proxy evaluation, reducing cross-model preference-win-rate estimation error without response likelihoods, but its effectiveness depends on representation coverage and residual transferability, and it does not outperform PPI under every shift.

SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

SciR generates verifiable deduction, induction, and causal-discovery problems before rewriting their premises as multi-document scientific discourse, separately varying inference complexity and premise obfuscation to diagnose LLMs; document formalisation remains a substantial bottleneck even with symbolic solvers.

Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning

OracleLadder fixes teacher-written mathematical sub-goals and crosses isolated milestone tests with increasing roadmap and answer assistance on the parent problem: across six models, 33โ€“48% of a screened 354-problem NuminaMath set falls into a composition gap, where all milestones are solvable but the parent remains unrecovered; the share is 24โ€“37% after excluding grading issues flagged by automated review, but this is not a pure causal measurement of an internal composition ability.

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

SpanUQ distills claim consistency across offline sampled responses into a probe over frozen LLM hidden states, using set prediction to locate semantic spans and assign continuous uncertainty; five separately trained backbone probes achieve 0.908โ€“0.944 AUROC and support finer-grained risk filtering than rejecting entire responses.