Skip to content

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

Conference: NeurIPS2026
arXiv: 2607.05721
Paper: Project page
Area: LLM Evaluation
Keywords: Span-level uncertainty, hidden-state probing, multi-sample distillation, Mixture of Beta, selective prediction

TL;DR

SpanUQ distills claim consistency across offline sampled responses into a probe over frozen LLM hidden states, using set prediction to locate semantic spans and assign continuous uncertainty; five separately trained backbone probes achieve 0.908–0.944 AUROC and support finer-grained risk filtering than rejecting entire responses.

Background & Motivation

Generated responses often place facts with different reliability levels in the same sentence. Token entropy can identify unpredictable words, but cannot distinguish a rare, correct name from a fluent, incorrect assertion; sequence-level Semantic Entropy or SelfCheckGPT captures semantic disagreement through repeated generation but ultimately produces a single response score. Broadcasting that score to local claims gives correct and incorrect parts of the same response identical rankings, preventing precise highlighting or selective verification of the most questionable information.

The paper places the evaluation unit at a contiguous text span: each span carries an independently assessable semantic unit, its boundaries need not coincide with sentence boundaries, and claims can overlap. This granularity does not merely average token scores; it jointly identifies where an assessable claim occurs and how unstable it is. Fixed windows and BIO segmentation are therefore imperfect preprocessing choices: windows can misalign with semantic boundaries, while BIO struggles to represent overlapping spans. Accurate scoring cannot compensate for an unsuitable initial segmentation.

Offline multi-sample comparison provides fine-grained supervision, but resampling and verifying every claim after each online generation multiplies deployment costs. Frozen middle-to-late-layer hidden states provide another entry point: a small probe learns how internal representations relate to cross-sample inconsistency, moving expensive operations into dataset construction. Core Idea: jointly learn overlapping semantic span detection and continuous consistency-risk prediction so that the internal representations of one generated response replace online multi-sample comparison, rather than forcing a sequence score onto local text.

Method

Overall Architecture

The inputs are a prompt, one generated response, and its selected-layer hidden states from a frozen LLM; the outputs are start-position, end-position, and uncertainty triples, optionally aggregated into a sequence score. Training first constructs spans and soft labels through multi-sample distillation; deployment does not call the labeling judge, sample additional responses, or retrieve external knowledge.

The main stages are set-based span detection, content enrichment and Beta mixture prediction, and uncertainty-conditioned iterative refinement. Detection internally fuses selected layers, compresses and encodes tokens, and uses a query set to locate spans; prediction reads content inside those spans; refinement runs the shared decoder again using the first-round scores. “Single forward pass” refers to extracting backbone representations for one response and invoking the probe once, not generating an entire autoregressive response with one backbone operation or running the UCIR decoder only once.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Offline: primary response<br/>and 20 sampled responses"] --> B["Multi-sample distillation"]
    B -.->|Training: span supervision| D["Set-based span detection"]
    C["Inference: one response<br/>frozen LLM hidden states"] --> D
    D --> E["Content enrichment and<br/>Beta mixture prediction"]
    B -.->|Training: consistency soft labels| E
    E --> F["Uncertainty-conditioned<br/>iterative refinement"]
    F --> G["Valid spans and scores<br/>optional sequence aggregation"]

Key Designs

1. Multi-sample distillation: supervision is a conditional contradiction rate, not a factual error probability

For each prompt, a greedy primary response supplies the text positions and hidden states used as probe inputs; another 20 responses are generated at temperature 1.0 with nucleus sampling parameter 0.95. The judge decomposes only the primary response into atomic claims with character offsets aligned to the original text, then reads the sampled responses for each claim and classifies them as supported, contradicted, or not mentioned. Because the claims come from one primary response, no additional cross-generation claim matching is required. The judge is Claude Opus 4.6, but its verification reference is the model’s sampled text, not a database or retrieved evidence.

For span \(k\), the label is the fraction of contradictions among samples that explicitly support or contradict the claim:

\[ u_k^*=\frac{n_k^{\mathrm{con}}}{n_k^{\mathrm{sup}}+n_k^{\mathrm{con}}},\qquad n_k^{\mathrm{sup}}+n_k^{\mathrm{con}}>0. \]

Not mentioning a claim is not counterevidence and must be excluded from the denominator; if all samples are silent, the span has no defined label and is dropped. The soft label therefore measures how inconsistent the model is when addressing the proposition, not the probability that the proposition is objectively false. An identically repeated error can receive a low score, and a true proposition can receive a high score when the model vacillates. The initial method description abbreviates the target as an unsupported fraction; the equation and dataset-construction section provide the precise definition.

Soft labels preserve intermediate risk levels, but the effective evidence count is not always 20: when few samples take a position, one judge error can substantially change the ratio. The appendix claims that a single misjudged sample changes the label by at most 0.05, which holds only under particular conditions such as a denominator of 20; excluding silent samples removes that universal bound. This suggests retaining the effective evidence count or supervising with label confidence intervals instead of treating every ratio as equally precise.

2. Set-based span detection: queries learn overlapping boundaries rather than imposing a segmentation first

The main Qwen3-14B setting averages hidden states from layers 22, 24, and 26 element-wise, projects the 5120-dimensional states to 512 dimensions, adds sinusoidal positional encodings, and applies a two-layer Transformer encoder. Layer selection follows single-layer probing: uncertainty signals exhibit an inverted-U pattern in middle-to-late layers, and the final layer is not optimal. The resulting token feature pool feeds both decoder cross-attention and subsequent content enrichment, giving localization and scoring a shared representation space.

The main configuration has 32 learnable queries, permitting up to 32 candidate spans; a three-layer Transformer decoder uses query self-attention to coordinate competing and overlapping claims, followed by cross-attention to locate content in the token feature pool. Queries do not initially bind to particular positions and develop soft specialization during training. Each query independently regresses normalized start and end coordinates, allowing two queries to cover the same token without assigning a unique BIO label to every token. The count of 32 belongs to the main setting rather than all backbones: the final Mistral-7B configuration uses 16 queries.

Boundary regression must read the query before content enrichment; otherwise, using boundaries to extract content and then using that content to predict boundaries would create a circular dependency. Training uses Hungarian one-to-one assignment between predictions and annotations, combining boundary L1/GIoU, absolute uncertainty error, and validity classification in the matching cost; unmatched queries learn suppression as empty slots. Inference removes empty slots at a validity probability threshold of 0.5. This probability indicates whether a span was detected, not whether its content is factually correct.

3. Content enrichment and Beta mixture prediction: turn localization representations into calibrated local scores

A query knows roughly where to look but may not retain the most informative factual content there. SpanUQ constructs a soft boundary mask from two sigmoids and performs query-guided attention pooling within that region; differentiability lets scoring errors propagate back to the boundaries. A gated residual adds pooled content to the query, allowing dates or entity relations to receive more weight than articles while introducing less redundant content for short spans. The appendix uses boundary sharpness 10, which controls the transition at boundaries rather than introducing another discrete segmentation.

The enriched query predicts validity and Beta mixture parameters. Beta distributions naturally cover the 0–1 range, and a three-component mixture can represent different shapes near endpoints and in intermediate regions more flexibly than one component; a shape-parameter floor of 0.5 permits U-shaped components. The point estimate is the weighted sum of component means:

\[ p(u_k\mid\tilde{\mathbf q}_k)=\sum_{j=1}^{3}\pi_{k,j}\operatorname{Beta}(u_k;\alpha_{k,j},\beta_{k,j}),\qquad \hat u_k=\sum_{j=1}^{3}\pi_{k,j}\frac{\alpha_{k,j}}{\alpha_{k,j}+\beta_{k,j}}. \]

This distribution fits the conditional soft-label distribution given internal representations; training does not directly supervise a claim’s full true probability distribution. A mixture-weighted sum of component shape-parameter sums gives the effective precision used in refinement. It is a learned concentration proxy, not a validated decomposition of epistemic uncertainty. In particular, differences between component means also affect total variance, so component precision alone is not equivalent to a full predictive confidence interval.

4. Uncertainty-conditioned iterative refinement: a second decoding pass with shared weights, not another sample

After the first pass predicts uncertainty mean, effective precision, and validity probability, a small MLP maps these three scalars into a conditioning vector, adds it to the query, and runs the same three-layer decoder and prediction path again. The second pass revisits the same token feature pool conditioned on its initial judgment, without invoking a second LLM or generating another response. The final score is a convex combination with second-round weight 0.7 and first-round weight 0.3; both rounds receive training supervision, with reduced loss weight for the first round.

This reuses parameters rather than eliminating second-pass computation. The paper separately reports UCIR overhead below 15% relative to the probe’s own base decoder and total probe latency below 3% relative to the backbone forward pass; these claims have different denominators. Sharing weights alone also does not establish small computational overhead, so the cost claims should remain scoped to the authors’ reported implementation and hardware measurements.

When a sequence score is required, the model learns span importance, applies softmax to validity-modulated importance scores, and takes a weighted average of span uncertainties. Sequence supervision is the mean annotated span soft label, so the experiment tests composability for that particular target rather than proving that every form of sequence uncertainty can be fully decomposed into local claims. The main equation specifies a weighted mean, while some experimental descriptions call it importance-weighted top-k; the source does not clearly reconcile these formulations.

A Worked Example

In the appendix’s John Derek biography example, the generated birth name, birth date, and parentage information form separate spans, alongside other spans describing his career. Offline labeling assigns the birth-date span a risk of 0.95; at deployment, the probe does not read the 20 reference responses, instead detecting the span and predicting 0.89 from the current biography’s hidden states. The appendix also reports 0.94 for the birth name, 0.95 for parentage, and 0.03 for reliable career information.

The output therefore does not declare the entire biography unreliable; it identifies local claims to prioritize for verification. These numbers belong to a qualitative example rather than an average over biographies. A low score indicates low estimated instability, not external factual verification.

Loss & Training

The backbone remains frozen and only the probe is trained. The main experiment uses 17,494 training prompts, with 15 epochs of detection-only warmup followed by up to 25 joint-training epochs; learning useful boundaries first avoids initially learning risk from the wrong regions. AdamW uses learning rate \(10^{-4}\), weight decay 0.01, batch size 16, cosine annealing, and early stopping on dev AUROC with patience 5.

The joint objective combines boundary L1/GIoU, all-query validity BCE, Beta mixture negative log-likelihood on matched spans, risk ranking, sequence consistency, and intermediate decoder-layer auxiliary supervision. NLL fits continuous labels, while ranking emphasizes risk order; neither replaces the other. Endpoint labels are clipped to \([10^{-4},1-10^{-4}]\) before NLL computation to avoid numerical instability.

Ranking uses margin 0.1: the appendix groups high-risk spans above 0.3 and low-risk spans below 0.1, sampling up to 256 pairs per batch and falling back to upper and lower quartiles when necessary. The main text describes within-sequence pairs, whereas the appendix describes batch-level stratified sampling. The pairing scope differs in wording and should not be silently asserted to be exclusively within-sequence.

Sequence consistency detaches gradients from span predictions to avoid pulling all spans toward one average score, primarily adapting the aggregation path to the sequence target. Appendix D.1 lists default loss weights of 5.0, 2.0, 4.0, 1.0, 0.5, and 0.4 in the stated order; Appendix A.3 instead gives auxiliary weight 0.5. Reproduction should check this discrepancy rather than silently reconcile it.

Key Experimental Results

Main Results

SpanUQ-Bench covers long-form QA, TriviaQA, ELI5, Biography, and FELM. The retained dataset has 19,994 prompts: 17,494 training, 500 development, and 2,000 test prompts. The raw pipeline emits 311,385 spans; removing spans addressed by no sample leaves approximately 293K overall and 256K in training. Raw and filtered counts must not be conflated.

The following selection from Table 1 reports default seed 42 results, with a separate probe trained on each backbone’s own responses, labels, and hidden states. AUROC uses the binary target of whether consistency labels are at least 0.5, MAE evaluates continuous consistency labels, and correlation is span-level Spearman; these are not AUROC measurements directly against human factual labels. The MLP baseline has ground-truth span boundaries and is an oracle comparison.

Backbone SpanUQ AUROC SpanUQ MAE SpanUQ Span Spearman Oracle MLP AUROC
Qwen3-14B 0.939 0.110 0.685 0.881
Qwen3-8B 0.930 0.129 0.692 0.884
Qwen3-4B 0.944 0.126 0.754 0.873
Qwen3-30B-A3B 0.936 0.110 0.647 0.893
Mistral-7B 0.908 0.126 0.637 0.863

On Qwen3-14B, SelfCheckGPT-NLI obtains AUROC/span Spearman of 0.808/0.471, and Token Entropy obtains 0.603/0.097. Broadcasting sequence scores inherently prevents within-response ranking, so this comparison exposes a granularity-adaptation advantage rather than demonstrating universal failure on their native sequence-level tasks.

The independent human study uses five non-author annotators, each judging 400 non-overlapping spans for factual correctness and submitting external evidence, for 2,000 spans overall. The following Table 2 results evaluate actual error after rejecting uncertain spans using those human labels, rather than comparing again with the training pipeline labels.

Method Risk at 100% Coverage Risk at 80% Coverage Risk at 60% Coverage Risk at 40% Coverage AURC PRR
SpanUQ 18.6% 10.5% 7.9% 6.3% 0.076 0.587
Oracle MLP Probe 18.6% 12.8% 9.4% 7.1% 0.098 0.470
SelfCheckGPT-NLI 18.6% 16.1% 14.5% 12.9% 0.151 0.184

Risk is the human-judged error rate among retained spans, coverage is the retained fraction, and AURC is the area under the risk–coverage curve, where lower is better. PRR measures the fraction of the random-to-oracle gap closed relative to random rejection. Rejecting the highest-risk 20% reduces SpanUQ’s risk from 18.6% to 10.5%, a 43.9% relative reduction. Thus, “approximately 80% coverage for a 10% risk budget” is only approximate: the table’s risk at 80% coverage remains above 10%, not a strict budget guarantee.

Ablation Study

The following progressive component additions preserve Table 4’s results without presenting them as numerically identical to Table 1’s evaluation.

Config AUROC ECE MAE Span Spearman Sequence Spearman
Base DETR + Single Beta 0.907 0.058 0.154 0.673 0.765
+ Enrichment Gate 0.920 0.052 0.144 0.725 0.729
+ MoB (three components) 0.933 0.020 0.116 0.760 0.846
+ UCIR (one refinement round) 0.939 0.013 0.106 0.790 0.839

MoB reduces MAE from 0.144 to 0.116 and ECE from 0.052 to 0.020, making it an important contributor to continuous scoring and calibration. UCIR further improves local ranking but slightly decreases sequence correlation. The results should not be summarized as every metric improving after every module addition.

The source’s full Qwen3-14B MAE/span Spearman is 0.110/0.685 in Table 1 but 0.106/0.790 in Table 4, without a clear explanation of the difference; both are reported separately. Sequence Tables A1/A3 give Spearman 0.851 and Pearson 0.839, whereas the abstract and Table 4 also use 0.839 for sequence correlation. These occurrences of 0.839 should not all be treated as the same correlation coefficient.

Key Findings

  • The detection experiment controls scoring with the same MLP: DETR precision/recall/F1 is 0.857/0.970/0.910, versus the best window F1 of 0.653, a 39.4% relative improvement. DETR+MLP AUROC of 0.927 is not the full MoB+UCIR result of 0.939.
  • Detection F1 uses IoU of at least 0.3 for matching; uncertainty metrics use Hungarian-matched pairs without that IoU cutoff. Scoring quality and boundary quality are separate dimensions requiring separate verification.
  • Human labels agree with thresholded pipeline labels on 89.4% of spans, with Cohen’s kappa 0.618; 22.3% of high-label spans remain factually correct according to humans. A single-person audit of 45 such cases attributes 40 to genuine cross-sample conflict, so uncertainty cannot simply be renamed factual error.
  • Across five training seeds, Qwen3-14B AUROC is 0.930±0.007 and Mistral-7B AUROC is 0.884±0.014. The main table is not a multi-seed mean, and architectural portability does not mean zero-shot transfer of one probe across models.

Highlights & Insights

  • Interpretability comes from the output unit, not merely a more elaborate distribution head. Binding scores to independently verifiable semantic spans makes selective verification actionable.
  • Differentiable content enrichment connects boundary learning with score learning. Localization becomes more than preprocessing because local risk supervision can influence which tokens the model reads.
  • Separating detection evidence from risk-filtering evidence is valuable. Independent human risk–coverage results support practical filtering utility without automatically establishing that outputs are objective error probabilities.

Limitations & Future Work

  • White-box hidden states are required, excluding direct use with closed APIs that return only text; data construction also requires sampling and judge calls for each backbone. Other languages, reasoning steps, and code execution risk remain proposed extensions rather than experimental conclusions.
  • Consistency supervision has a structural blind spot for identically repeated errors, while finite evidence counts and judge noise lack component-wise quantification. Future training could combine external factual verification with explicit distinctions between semantic instability, insufficient evidence, and actual error.
  • Parameter counts conflict: the main experiment states 24.4M, while Appendix A14 lists 27.04M for Qwen3-14B; other probes vary, including 17.18–26.51M. “Approximately 25M” must not universally become less than 0.2% of the backbone: even 17.18M on the 4B backbone is approximately 0.43%.
  • The authors report probe overhead below 3% of backbone forward latency and 10–20-fold lower total time than methods requiring 10–20 samples. These have different denominators, exclude expensive offline labeling, and are not universal guarantees for every deployment platform.
  • Appendix C.9 calls the output an error probability, whereas the stricter dataset definition and limitations explicitly identify a cross-sample consistency score. Deployment should select thresholds using external factual labels from the target distribution and warn that low-risk, unhighlighted regions have not been verified.
  • vs Semantic Entropy / SelfCheckGPT: Sampling methods directly observe generation differences; SpanUQ distills related signals offline and learns local spans. The trade-off is supervised data, internal access, and per-backbone retraining.
  • vs Semantic Entropy Probes / HaloScope / HaMI: These also exploit internal representations to reduce extra generation costs. SpanUQ primarily differs by jointly modeling boundaries and continuous risk rather than broadcasting one sequence score.
  • vs FActScore / SAFE: External evidence checking targets factual correctness, whereas SpanUQ training compares only the model’s own samples. The main text’s FActScore baseline also uses an adapted self-verification protocol, so its results should not be generalized to standard retrieval-based factual verification.
  • vs DETR / BIO: Set prediction is borrowed for one-dimensional semantic span localization, not a visual object detection task. Promising extensions include supervision for effective sample coverage and budget allocation for external checking, testing which local scores actually save verification costs.

Rating

  • Novelty: 4/5 — Combines semantic span localization, continuous risk prediction, and shared-decoder refinement, with the main contribution in task granularity and system integration.
  • Experimental Thoroughness: 4/5 — Five backbones, component ablations, and independent human filtering evaluation are substantial; label-noise decomposition and cross-task validation remain limited.
  • Writing Quality: 3/5 — Mechanisms are clear, but parameter counts, pairing scope, aggregation, and some metric protocols have unexplained differences.
  • Value: 4/5 — Useful for local verification and inexpensive risk highlighting, provided consistency scores are not mistaken for factual certification.