Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets¶
Conference: NeurIPS2026
arXiv: 2606.25760
Area: Multimodal VLM
Keywords: GUI grounding, uncertainty quantification, ranking transfer, selective execution, conformal prediction
TL;DR¶
Argus compares uncertainty methods using unified single-step GUI click records under different observable interfaces, finding stronger cross-dataset method-ranking transfer at a fixed model than across models or from open weights to API-only systems, while error discrimination cannot replace probability calibration or spatial coverage checks.
Background & Motivation¶
A vision-language model (VLM) can predict a click coordinate from a screenshot and instruction, but the coordinate is intended for execution by the operating system rather than presentation as an answer. Whether it falls inside the target box is only the first question: deployment also requires deciding which predictions to defer, estimating how far errors stray, and checking whether a region around the predicted point contains the actual target. Differences in difficulty and target geometry across ScreenSpot-v2, ScreenSpot-Pro, OSWorld-G, and UI-Vision-EG make confidence findings from a single model–dataset pair difficult to generalize.
Existing methods use token probabilities, repeated sampling, hidden states, attention, or verbalised confidence, but these signals are not accessible through every deployment interface. Open-weight models expose internal tensors, whereas closed-source interfaces generally return text and coordinates. Replacing unavailable internals with proxies confounds interface changes with implementation changes. Earlier GUI uncertainty studies mainly emphasize individual scores or intervention mechanisms, leaving limited evidence about whether method selection transfers under a common evaluation protocol.
Rather than introducing a new uncertainty estimator, this paper defines evaluation regimes through the model, dataset, and observable interface, then separately evaluates error ranking, rejection, probability calibration, and spatial regions. Core Idea: unify records and calibration protocols, compare uncertainty-method rankings across regimes, and separately validate which score to choose and whether that score supports the intended intervention.
Method¶
Overall Architecture¶
The input is a screenshot and instruction, and the base model generates one proposed executable click; additional sampled or perturbed responses estimate uncertainty rather than execute a multi-step task. Argus first constructs interface-consistent per-item records, fits supervised scores and mappings on a separate calibration split, and then evaluates multiple objectives, method-ranking transfer, and conformal click disks on held-out test records.
The open-weight panel contains 27 methods from 7 families, evaluated on 4 models and 4 datasets for 16 model×dataset cells. The models are Qwen2.5-VL-7B (Q7), Qwen2.5-VL-72B-AWQ (Q72), UI-TARS-1.5-7B (UI), and POINTS-GUI-G-8B (PT). The API-only panel retains 8 faithfully computable methods across GPT-5.4, Claude Sonnet 4.6, Gemini 3.1 Pro, and the same 4 datasets, yielding 12 cells. Strings such as “2727,” “44,” and “88” in the cached text are duplicated mathematical rendering, not thousands of methods or dozens of models.
The diagram describes evaluation and calibration, not a newly trained agent network. Dashed arrows denote calibration supervision or fitted artifacts; solid arrows denote record and test-data flow.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["Screenshot and instruction"] --> A["Interface-consistent records"]
A --> B["Calibration-isolated scoring"]
C["Calibration records and labels"] -.-> B
B --> D["Multi-objective ranking transfer"]
B --> E["Conformal click disks"]
C -.-> E
T["Held-out test records"] --> D
T --> E
D --> O["Method rankings and metrics"]
E --> P["Center coverage and radius"]
Key Designs¶
1. Interface-consistent records: compare the same click task using executable formulas
Each item includes one greedy click, 5 stochastic samples, responses to 3 instruction paraphrases, and responses to 3 image perturbations; coordinate clustering uses a 50-pixel tolerance. Correctness is defined by whether the predicted point lies inside the target bounding box. Stochastic samples measure repeated-generation consistency, paraphrases measure stability under semantically equivalent wording, and image perturbations measure visual-input stability. These sources cannot simply be pooled as 11 identically distributed stochastic samples.
The seven families use different information sources: logit methods read generated-token probabilities or entropy; sampling methods measure coordinate-cluster disagreement or response similarity; hybrid methods combine confidence and consistency; density/probe methods use last-layer hidden states; attention methods track attention chains; verbalised methods ask for confidence; and VLM-native methods compare clicks under paraphrases or image perturbations. HEDGE uses cluster entropy over paraphrase responses. IMGHEDGE uses three perturbations: saturation increased by 30%, Gaussian noise with standard deviation 8, and JPEG quality 50.
Across interfaces, missing logprobs, hidden states, or attention maps are not replaced by same-name proxies. The API-only panel retains SelfCons, SE, LexSim, CCP, Verb-1S, Verb-2S, HEDGE, and IMGHEDGE, requiring the same formulas as on the open-weight side. CCP uses the fraction of stochastic clicks agreeing with the greedy click within tolerance; SE uses the entropy of coordinate-cluster masses. Both can be computed from responses alone, but that does not justify silently retaining token-probability-weighted variants.
Closed-source inference is single-shot, without tools or a Python interpreter, rather than iterative crop-and-zoom grounding used in some leaderboards. Interfaces are not otherwise identical: Sonnet requires image resizing, and coordinate-format drift or missing confidence is handled by strict parsing into missing values. Formula fidelity makes method definitions comparable but does not remove differences in resolution, response compliance, or effective sample counts.
2. Calibration-isolated scoring: post-hoc does not mean no training or labels
Uncertainty outputs are oriented so that larger values indicate greater predicted error risk. Simple token or consistency scores can be computed directly from records, but density/probe methods require calibration data. Mahalanobis measures distance from the correct-class distribution in hidden-state space; Mahal-RMD subtracts a background-distribution distance; SAPLMA uses an MLP probe; and SEP uses a logistic-regression probe. RMD reduces interference from shared representation-scale effects rather than merely changing a classification threshold.
These estimators are fitted on the current cell’s calibration split and evaluated on that cell’s test split. The headline analysis tests cross-cell method rankings, not direct transfer of a probe trained in one cell to another. “Post-hoc” therefore means that the base GUI model is not updated; it does not make every method training-free or establish cross-domain probe-parameter validity from ranking stability.
To interpret scores as error probabilities, the authors fit isotonic regression on calibration records, with a preceding quantile transform for heavy-tailed scores, and measure ECE and Brier on test records. A score can rank most errors above correct predictions without having an interpretable probability scale, which motivates separate discrimination and calibration evaluation. The main text notes ECE above 0.40 for several density methods, but Appendix A15 qualifies this: among 80 cell×density-method combinations, none simultaneously ranks first within the family by AUROC and has ECE above 0.40. This does not support claiming that all winning density methods are severely miscalibrated.
3. Multi-objective ranking transfer: errors, severity, and method recommendations are different objects
AUROC treats incorrect clicks as positives and tests whether high scores rank them above correct clicks. For selective execution, PRR measures oracle-normalized rejection quality up to 50% rejection; AURC integrates error risk over accepted coverage, with lower values being better. Neither is the base model’s point-in-bbox accuracy, nor can either directly guarantee the risk of actual interface-action consequences.
To distinguish near misses from severe deviations, AUSE is evaluated only on incorrect clicks. Severity divides click-to-target-center distance by target scale and then applies a logarithmic compression. The legible source definition is:
Here \(\mathbf{c}^{\star}_i\) is the target center, and \(w_i,h_i\) are target-box width and height. AUSE compares the remaining-error curve obtained by removing errors in uncertainty order with the oracle curve obtained by removing them in true-severity order; lower is better. The normalization distinguishes deviations relative to small buttons from those relative to large panels, but does not model button function, action cost, or task recoverability.
For transfer, the authors rank methods within each cell using test AUROC and compute Spearman \(\rho\) between two rankings. The object is whether method A still outperforms method B—not correlation between per-item scores of one method, and not portability of raw scores or calibrated probabilities. The full open-weight matrix yields 120 cell pairs. The cross-tier conclusion compares only Q72-AWQ against the 12 closed-source cells on matched datasets, restricted to the shared 8 methods.
4. Conformal click disks: convert residuals into regions while specifying the coverage event
The fixed disk uses calibration click-to-target-center Euclidean distances as residuals and chooses a corresponding quantile as radius. At test time, the disk is centered at the predicted click, and coverage means that the target-box center lies inside it. The main text gives the following quantile relation; finite-sample conformal quantile conventions still need checking against the implementation rather than treating this shorthand as complete detail.
Disk-Normalized uses plug-in uncertainty as a local error scale, producing different radii for different test clicks. Disk-CQR uses conformalized quantile regression to construct more conservative regions. The cache does not supply complete training objectives sufficient to reproduce both variants; this note therefore explains their effect on radius without inventing exact losses or repairing corrupted formulas.
Marginal conformal coverage depends on exchangeability between calibration and test records. Changes in API behavior or sample composition require fresh empirical coverage checks. Even a disk containing the target center may contain several unrelated controls: it is not a region where every point is safe to click, and a large radius does not imply greater operational safety.
A Worked Example¶
For a screenshot instruction asking to click a button, the model first produces a proposed coordinate. The additional 5 stochastic responses estimate disagreement rather than constitute a 5-step trajectory. If the sampled clicks form two clusters under the 50-pixel tolerance, SE reflects dispersion of cluster mass, while CCP asks whether sampled clicks remain near the original greedy click. Consistency may still be high if the model repeatedly selects the wrong button.
A deployment exposing hidden states can fit SAPLMA or SEP using calibration labels, then compare AUROC, AUSE, and calibration error on records excluded from fitting. An API-only deployment cannot reuse these probes and must rerank its 8 available methods. This screenshot-and-cluster walkthrough illustrates the mechanism; it is not an additional empirical example reported by the paper.
A reported spatial example is PT×OSWorld-G: at \(\alpha=0.10\), the main text gives a fixed-disk radius of 743 pixels and a normalized mean radius of 481 pixels, approximately 35% smaller. This indicates a tighter region at the same coverage target, not a 35% improvement in click accuracy, and does not authorize clicking other controls inside the disk.
Loss & Training¶
The base GUI models remain frozen. Training or fitting occurs in post-processing: hidden-state probes, density estimators, isotonic probability mappings, and conformal thresholds. Panel and perturbation selection should likewise use calibration data, not test labels followed by deployment-performance claims.
The main protocol uses test/calibration = 80/20 across 50 stratified split seeds. These are repeated splits and fits over fixed per-item inference records, not 50 fresh base-model training runs. The calibration/test ratio sensitivity study fixes seed 0 and changes only the cut point; its values should not be conflated with the main table’s 50-split means.
The paper reports 500 bootstrap resamples and paired comparisons, but the main text describes stratified resampling while Appendix A29 specifies resampling over 50 seed values. The statistical unit is thus described differently, and shared underlying records mean these splits are not 50 independent datasets. The inclusion criterion of “30 minutes per cell” applies only to scoring after per-item records exist; generating Q72×OSWorld-G records still takes approximately 6.7 hours.
Key Experimental Results¶
Main Results¶
The following representative cells come from Table 4 and Appendix A29. Accuracy is the base click hit rate; AUROC and PRR are mean values for the cell’s AUROC-best method. These columns are not interchangeable accuracy measures. SP denotes ScreenSpot-Pro, V2 ScreenSpot-v2, and OSG OSWorld-G.
| Cell | Accuracy | AUROC-best method | AUROC | PRR |
|---|---|---|---|---|
| Q7×V2 | 0.883 | SAPLMA | 0.817 | 0.599 |
| Q72×SP | 0.447 | SAPLMA | 0.889 | 0.802 |
| UI×V2 | 0.878 | CoCoA-1MCA | 0.842 | 0.630 |
| PT×OSG | 0.659 | Mahal-RMD | 0.820 | 0.568 |
| GPT-5.4×SP | 0.367 | CCP | 0.735 | 0.574 |
| Sonnet 4.6×SP | 0.341 | CCP | 0.731 | 0.679 |
| Gemini 3.1 Pro×V2 | 0.447 | Verb-1S | 0.966 | 0.951 |
| Gemini 3.1 Pro×SP | 0.313 | Verb-2S | 0.856 | 0.721 |
The high AUROC on Gemini×V2 does not imply nearly perfect clicks: the base hit rate is only 0.447. Appendix A17 analyzes 273 valid confidence records from a locked 300-item subset, finding a sharply bimodal score distribution with distinct correctness rates across clusters. The finding remains restricted to valid records and the single-shot, no-tools protocol; missing responses must also be considered.
Ablation Study¶
These are transfer and objective-difference analyses, not module-removal ablations of a new network. Spearman measures method-ranking transfer rather than transfer of calibrated probabilities.
| Analysis scope | Count | Result | Interpretation boundary |
|---|---|---|---|
| All open-weight cell pairs | 120 pairs | Mean \(\rho=0.705\) | Mixes dataset and model changes |
| Fixed model, different datasets | 24 pairs | Mean \(\rho=0.79\) | Rankings are relatively stable at a fixed model |
| Fixed dataset, different models | 24 pairs | Mean \(\rho=0.69\) | Does not isolate scale, quantization, or fine-tuning |
| PT×OSG versus PT×SP | 1 pair | \(\rho=0.969\) | Strongest representative fixed-model transfer |
| Q72-AWQ versus matched closed-source cells | 12 pairs | Mean \(\rho=0.08\); 95% CI [-0.219, 0.373] | Shared 8 methods only; interval includes zero |
| Closed-source cell pairs | 66 pairs | Mean \(\rho=0.127\) | Strong vendor and dataset dependence |
| Agreement of AUROC/AUSE winners | 16 open-weight; 12 closed-source cells | 2/16; 9/12 | Different panel sizes; not a controlled causal comparison |
The main-text conformal examples below target 0.90 coverage of the target center. Radii are in pixels, and percentage reductions are approximate values reported in the text.
| Cell | Disk-Fixed radius | Disk-Normalized radius | Reduction |
|---|---|---|---|
| PT×OSG | 743 | 481 | 35% |
| Q72×OSG | 837 | 620 | 26% |
| UI×SP | 1469 | 1076 | 27% |
| Q72×SP | 1124 | 900 | 20% |
The abstract and conclusion summarize normalized-radius reductions of 40–60%, whereas these specific Section 7 examples give 20–35%. The former range should not be attributed to the cells above; the cache does not sufficiently explain how the two ranges relate. For closed-source Gemini×SP, fixed-disk empirical coverage is 0.926/0.884/0.795 at targets 0.95/0.90/0.80, reinforcing the need to check actual coverage.
Key Findings¶
- Density/probe strength primarily means that strong members often win and remain stable under model transitions, not that every density method performs well. Mahal-RMD outperforms naive Mahalanobis on all 16 cells, while the naive method itself is often near or below random discrimination.
- On ScreenSpot-Pro, the Q7→Q72 sampling/hybrid family AUROC changes are -0.277/-0.257, versus -0.039/-0.010 for Q7→UI. Degradation cannot uniformly be attributed to GUI fine-tuning. Q72 also uses AWQ, so the nominal scale-only comparison retains a quantization confound.
- Source inconsistencies remain explicit: the A11 caption says density-family means win most cells, which its displayed means do not support; A28 says changed cells remain within ±0.01 of baseline, yet UI×UIV reports 0.840 at 20/80 and 0.891 at 50/50. This note does not silently modify the original values.
Highlights & Insights¶
- Treating observability as part of method applicability matters more than simply adding baselines. Comparing same-name proxy formulas obscures deployment constraints; a faithfully shared intersection is easier to interpret.
- Answer accuracy and awareness of errors are separable. Metrics should match rejection, severity ranking, or probability interpretation rather than selecting only the highest AUROC.
- A recommended panel can serve as a transfer prior, but probes, probability mappings, and spatial radii need refitting or validation in the target regime. High rank correlation does not eliminate the need for target-domain labels.
Limitations & Future Work¶
- The study covers single-step click grounding, not multi-step trajectories, recovery, changing state, or accumulated risk. It does not establish multi-agent task-success conclusions.
- Only 5 pure stochastic samples limit reliable dispersion estimation; paraphrases and perturbations do not substitute for SafeGround’s larger pure stochastic budget. Increasing the budget can also change method cost rankings.
- Model-class and interface changes coincide with differences in capability, alignment, backbone, scale, quantization, resizing, and format compliance. Results are descriptive comparisons, not isolated causal effects.
- Ranking stability across repeated splits does not establish cross-cell probe transfer or conditional coverage for every UI subgroup. Target-size and scene stratification, followed by actual cross-domain deployment, remain useful extensions.
- Center distance does not distinguish the consequence of clicking a deletion control from clicking empty space. Future evaluation could incorporate semantic action costs, restrictions on irreversible operations, and actual clickable-region constraints rather than treating large disks as reliability.
Related Work & Insights¶
- vs LM-Polygraph / CoCoA: These supply text-generation UQ methods and confidence–consistency combinations. Argus faithfully adapts them to click coordinates and evaluates method rankings after regime changes, without claiming a new universal estimator.
- vs SafeGround: SafeGround uses stochastic click dispersion and Learn-Then-Test for rejection. Argus compares multiple families and observable interfaces, but its main sampling budget cannot fully replicate the former’s pure stochastic setting.
- vs HyperClick / UI-Zoomer / V2P: They respectively learn a spatial-confidence head, trigger zoom-and-regrounding with uncertainty, or calibrate visual attention. Argus adds none of these intervention modules; it supplies cross-regime evidence for selecting post-hoc signals.
- vs split conformal / CQR: Conformal methods address region coverage under assumptions; Argus observes radius and empirical coverage using click disks. Region coverage, click accuracy, and actual operation risk should remain separate audit targets.
Rating¶
- Novelty: 4/5 — The main contribution is a cross-regime GUI uncertainty-selection benchmark rather than a new estimator.
- Experimental Thoroughness: 4/5 — Open-weight and closed-source panels, multiple metrics, and sensitivity analyses are broad, but causal controls and cross-domain deployment remain limited.
- Writing Quality: 3/5 — The central question is clear, but radius-improvement ranges, family-mean summaries, and some appendix descriptions conflict.
- Value: 4/5 — Directly useful for GUI-grounding method selection, provided target calibration and coverage checks are retained.