Skip to content

Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning

Conference: ECCV 2026
Paper: ECCV Official
Area: LLM Reasoning
Keywords: multimodal agent, reinforcement learning, self-verification, tool use, test-time compute

TL;DR

Addressing poor tool-calling calibration and severe evidence noise in multimodal reasoning agents, this paper proposes Self-Verification via Reinforcement Learning (SVRL), an RL-only framework that calibrates search timing with weak competence priors, rewards query diversity, and embeds structured self-verification into reasoning traces, eliminating external verifiers at inference time and enabling test-time search compute scaling.

Background & Motivation

Multimodal large language models (MLLMs) are transitioning from static visual perception systems into interactive agents capable of invoking external tools such as web search and code execution. In complex multi-hop visual question answering (VQA) scenarios, visual evidence alone is frequently inadequate to deduce the final answer; models must accurately identify visual entities and retrieve missing factual knowledge from external repositories. However, under current reinforcement learning finetuning recipes such as GRPO, optimization relies almost exclusively on sparse outcome-level accuracy rewards. Consequently, agents receive virtually no fine-grained process feedback regarding intermediate decisionsโ€”specifically when to search, what to search for, and which retrieved evidence to trustโ€”which leads to frequent tool-calling failures.

In practice, existing tool-augmented multimodal agents exhibit two primary bottlenecks. First, their tool invocation calibration is severely deficient: agents either omit necessary searches due to over-confidence or trigger gratuitous queries when internal parametric knowledge is fully sufficient. On the MMSearch-R1 baseline, unnecessary tool calls exceed 30% of cases, while required searches are skipped in over 10% of instances. Second, retrieved web snippets are noisy and lack intrinsic filtering mechanisms: returned documents are frequently irrelevant, outdated, or contradictory (in MMSearch-R1, nearly 45% of retrieved content is irrelevant noise), causing models to credulously adopt misleading snippets and fail on over 28% of queries even when valid evidence is returned. Although deploying external verifiers (e.g., GPT-5) at test time filters distractors and improves accuracy, it incurs substantial inference latency and deployment costs.

This paper's core angle is that verification and evidence filtering can be internalized directly into the agent's own reasoning traces via reinforcement learning, completely removing the dependency on external verifiers at test time. Core idea: introduce Self-Verification via Reinforcement Learning (SVRL), an RL-only finetuning framework that leverages a single-pass forward competence prior from the pretrained model to penalize avoidable searches, rewards diverse multi-query proposals, and generates structured usefulness scores and justifications directly within reasoning trajectories, optimized stably via Dr. GRPO for effective inference-time search scaling.

Method

Overall Architecture

SVRL formalizes multimodal multi-hop reasoning as a policy optimization problem over decision trajectories. Given an input image \(I\) and a natural language question \(q\), the policy network \(\pi_\theta\) produces a trajectory \(\tau = (s_{1:T}, y)\) comprising intermediate reasoning steps \(s_{1:T}\) and a final answer \(y\). At each discrete step \(t\), the agent outputs natural language chain-of-thought enclosed within <reason>...</reason> tags, action decisions (such as text search or image search), and explicit structured evidence verification tags <verify>...</verify>.

During finetuning, the pipeline first constructs a search competence prior based on greedy zero-tool model predictions. During rollout generation, the agent independently proposes multiple candidate queries and assigns binary usefulness vectors to returned search snippets. An oracle verifier conditioned on ground-truth answers provides supervision during training by aligning query quality and snippet relevance, with the resulting reward factors multiplicatively gated by answer correctness under Dr. GRPO. At inference time, the oracle verifier is discarded, and the agent executes autonomous evidence filtering using its internalized self-verification capability.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Multimodal Query<br/>(Image I, Question q)"] --> B["Weak Competence Prior Tagging<br/>Pre-determine search_free attribute"]
    B --> C["Tool Calibration<br/>Identify needed search & penalize redundant calls"]
    C --> D["Candidate Query Sampling & Diversity<br/>Generate k candidate queries and reward diversity"]
    D --> E["Tool Retrieval Execution<br/>Retrieve external text or image snippets"]
    E --> F["Trace-Internal Structured Self-Verification<br/>Generate binary verification vector & justification"]
    F --> G["Dr. GRPO Policy Optimization<br/>Remove normalization distortion to stabilize training"]
    G --> H["Final Answer Generation & Evaluation<br/>Eliminate external verifier & enable test-time scaling"]

Key Designs

1. Tool Calibration: Penalizing Avoidable Tool Calls via Weak Competence Priors

Standard reinforcement learning approaches apply uniform per-step action penalties to discourage latency, but this proves counterproductive when finetuning pretrained foundation models: it fails to distinguish between questions solvable from internal knowledge and challenging questions where retrieval is indispensable, suppressing needed exploration on difficult queries. SVRL resolves this by constructing a search-aware factor grounded in a weak competence prior. Prior to RL training, the base policy \(\pi_\theta\) performs a single greedy forward pass without tools on each training instance to generate \(\hat{y}\); if \(\hat{y} = y^*\), the sample is labeled as search_free. During online RL rollouts, an attenuation penalty is applied exclusively when the agent invokes tools on a search_free instance: $\(r_{\mathrm{aware}} = \begin{cases} \lambda_{\mathrm{sa}}, & \text{if instance is } search\_free \text{ and a search tool is called} \\ 1, & \text{otherwise} \end{cases}\)$ where \(\lambda_{\mathrm{sa}} = 0.1\). This condition-specific penalty curbs redundant calls while preserving the policy's incentive to search on genuinely demanding problems.

2. Candidate Query Sampling & Diversity: Elevating Retrieval Query Formulation

Multimodal agents often produce underspecified or noisy search queries, causing early retrieval failures that permanently derail reasoning trajectories. To break the informational bottleneck of single query formulation, SVRL prompts the agent to propose \(k\) candidate queries simultaneously, executing one valid and unique candidate at random. To incentivize structured exploration, a query-count reward factor is introduced: $\(r_{\mathrm{count}} = \begin{cases} 1, & \text{if trajectory makes no text-search call} \\ \max(\epsilon, \hat{k} / k), & \text{if search is called, where } \hat{k} \text{ is the number of valid unique queries} \end{cases}\)$ with \(\epsilon > 0\) serving as a numerical lower bound. This objective explicitly forces the model to articulate multiple complementary query angles, substantially improving the retrieval recall for complex visual entities.

3. Trace-Internal Structured Self-Verification: Internalizing Evidence Denoising

Retrieved web content is inherently contaminated with irrelevant or contradictory distractors. Rather than passing raw text into the context or adding latency-heavy external re-ranking models, SVRL instructs the agent to emit a structured binary usefulness vector (e.g., <verify> 1, 0, 1, 0 </verify>) alongside concise natural-language justification before formulating its answer. At training time, an oracle verifier equipped with ground-truth answers evaluates both candidate queries and returned snippets to produce target labels \(v^*\), yielding the query alignment factor \(r_{\mathrm{qalign}}\) and the snippet alignment factor \(r_{\mathrm{salign}}\): $\(r_{\mathrm{svrl}} = r_{\mathrm{aware}} \cdot r_{\mathrm{count}} \cdot r_{\mathrm{qalign}} \cdot r_{\mathrm{salign}}\)$ To prevent reward hacking where models generate flattering verification scores while producing wrong answers, process rewards are multiplicatively gated by answer correctness: $\(r(\tau) = (1 - \alpha) \cdot r_{\mathrm{acc}}(\tau) \cdot r_{\mathrm{svrl}}(\tau) + \alpha \cdot r_{\mathrm{fmt}}(\tau)\)$ At test time, the external oracle is entirely removed; the agent relies entirely on its internalized verification capability to filter noisy evidence.

4. Dr. GRPO Policy Optimization: Mitigating Gradient Spikes Under Trajectory-Dependent Rewards

Vanilla GRPO calculates group-relative advantages \(\hat{A}_i\) by normalizing rewards with intra-group empirical mean and standard deviation. However, under multifaceted trajectory-dependent rewards, within-group reward variance can fluctuate sharply or collapse to zero. Standard-deviation and sequence-length normalizations excessively inflate policy gradient norms, triggering response-length drift and training instability. SVRL adopts Dr. GRPO, eliminating heuristic normalizers that distort updates. This stabilization enables efficient and stable convergence using a compact group size of \(G=4\) (halved from \(G=8\) in MMSearch-R1) on a single 8ร—A100 GPU node.

Loss & Training

The overall optimization objective follows Dr. GRPO with a Kullback-Leibler divergence penalty against the reference model: $\(\max_{\theta} \; \mathbb{E}_{(I, q, y^*) \sim \mathcal{D}} \left[ \mathbb{E}_{\tau \sim \pi_\theta(\cdot \mid I, q)} \left[ r(\tau) \right] - \beta \, D_{\mathrm{KL}}(\pi_\theta(\cdot \mid I, q) \,\|\, \pi_{\theta_{\mathrm{ref}}}(\cdot \mid I, q)) \right]\)$ where the reference model \(\pi_{\theta_{\mathrm{ref}}}\) is the base Qwen-2.5-VL-7B-Instruct. The model is trained for 400 steps with a batch size of 32, group size of 4, and format weight \(\alpha\). Web search queries interface with Google Search APIs through a DynamoDB persistent caching layer, avoiding network non-determinism and duplicate API expenses.

Key Experimental Results

Main Results

The model was evaluated across five VQA benchmarks, encompassing in-distribution datasets (FVQA-test, InfoSeek) and out-of-distribution transfer benchmarks (MMSearch, LiveVQA, SimpleVQA). Reported metrics include LLM-as-judge accuracy Acc (%) and search ratio SR (%).

Model FVQA-test (Acc / SR) InfoSeek (Acc / SR) MMSearch (Acc / SR) LiveVQA (Acc / SR) SimpleVQA (Acc / SR)
Direct Answer
Qwen-2.5-VL-7B 26.7 / 0.0 20.1 / 0.0 12.8 / 0.0 17.8 / 0.0 38.4 / 0.0
Qwen-2.5-VL-32B 24.7 / 0.0 25.8 / 0.0 15.7 / 0.0 18.7 / 0.0 40.1 / 0.0
Qwen-2.5-VL-72B 27.1 / 0.0 28.0 / 0.0 15.7 / 0.0 20.1 / 0.0 42.2 / 0.0
GPT-4o 41.7 / 0.0 42.7 / 0.0 22.2 / 0.0 26.9 / 0.0 46.6 / 0.0
RAG Workflow
Qwen-2.5-VL-7B 51.6 / 100 53.7 / 100 52.2 / 100 48.0 / 100 51.6 / 100
Qwen-2.5-VL-32B 57.0 / 100 56.8 / 100 57.9 / 100 49.6 / 100 54.5 / 100
Qwen-2.5-VL-72B 62.2 / 100 59.4 / 100 59.6 / 100 56.0 / 100 61.0 / 100
GPT-4o 66.0 / 100 59.1 / 100 62.5 / 100 59.6 / 100 63.4 / 100
Adaptive Search
MMSearch-R1++ 56.4 / 80.3 55.1 / 59.7 53.8 / 88.5 48.4 / 76.2 57.4 / 42.5
SVRL-full-7B (Ours) 65.3 / 61.4 64.7 / 52.3 60.3 / 82.0 57.8 / 48.6 58.3 / 27.0

Ablation Study

Ablations on FVQA-test and InfoSeek isolate the impact of individual reward components and verification setups:

Config FVQA-test Acc (%) FVQA-test SR (%) InfoSeek Acc (%) InfoSeek SR (%) Note
MMSearch-R1 51.1 64.2 49.8 63.5 Baseline GRPO method
MMSearch-R1++ 56.4 80.3 55.1 59.7 ReAct prompt with standard GRPO
SVRL-query-align 57.7 97.6 58.2 98.0 Query alignment only; exhibits search hacking
SVRL-llm-judge 58.1 36.7 58.0 50.6 Train-time LLM judge causes verbose hedging
SVRL-search-aware 59.9 80.3 57.0 85.8 Calibrates search timing via competence prior
SVRL-ttv 60.6 80.3 58.3 85.2 External GPT-5 test-time verification
SVRL-self-verification 64.2 61.2 64.1 59.2 Trace-internal self-verification (standalone)

Key Findings

  • Internalized Self-Verification Drives the Largest Performance Gain: Progressing from MMSearch-R1++ to full self-verification yields an 8 to 9 percentage point accuracy gain across both in-distribution benchmarks. Strikingly, internalized self-verification (64.2%) surpasses test-time verification with an external GPT-5 judge (60.6%), showing that the model learns not just to filter evidence, but to dynamically integrate verified snippets into coherent CoT reasoning.
  • Search-Aware Gating Prevents Severe Reward Hacking: Introducing query-alignment rewards without search-aware gating causes the policy to issue tool calls on ~98% of queries to exploit query rewards. The weak competence prior effectively eliminates redundant calls on simple instances, balancing accuracy and inference efficiency.
  • Enabling Strong Test-Time Search Scaling: In parallel search scaling experiments (budget expanding from 1 to 15 searches), SVRL continues to demonstrate steady accuracy gains. In contrast, the baseline MMSearch-R1++ plateaus after 5 searches due to evidence confusion. This demonstrates that robust self-verification is prerequisite to unlocking the benefits of expanded test-time search compute.

Highlights & Insights

  • Internalizing Verifiers into Generative Reasoning Tokens: By converting external verification into model-generated <verify> tokens and rationales, the framework provides dense credit assignment during RL training while maintaining zero extra model latency during inference.
  • Zero-Tool Competence Priors for Reward Shaping: Utilizing a single greedy forward pass from the un-finetuned model as an offline prior elegantly prevents uniform search penalties from discouraging necessary exploration.
  • Denoising as the Bottleneck to Agentic Test-Time Compute: The findings demonstrate that scaling test-time search compute in multimodal agents fails not from tool invocation limits, but from the inability to filter retrieved noise. Self-verification resolves this fundamental constraint.

Limitations & Future Work

  • Reliance on Closed-Source Oracle Judges During Training: Training-time supervision of query and snippet alignment depends on GPT-5 paired with ground truth, which could impart judge-specific inductive biases into the compact model.
  • Format and Prompt Sensitivity: The policy relies heavily on strict ReAct formatting and XML-style parsing; out-of-distribution robustness across diverse prompting environments requires further investigation.
  • Shallow Multi-Turn Interaction Horizon: Evaluated tasks focus primarily on one- or two-hop search interactions rather than open-ended interactive web browsing involving dynamic clicks, page scrolling, and form inputs.
  • vs MMSearch-R1: While MMSearch-R1 pioneered GRPO finetuning for multimodal search, its uniform search penalties and sparse terminal rewards lead to high noise acceptance (45% irrelevant snippets). SVRL introduces search calibration and structured self-verification, achieving ~9% higher accuracy with a ~20% reduction in search calls.
  • vs RAG Workflow: Traditional RAG queries external knowledge unconditionally (SR=100%), which degrades efficiency and introduces noise on visually grounded queries. SVRL performs adaptive on-demand retrieval, outperforming heavy RAG baselines at lower tool costs.
  • vs External Test-Time Verifiers: Prior systems attach external rerankers or judge LLMs to filter search noise at inference time. SVRL proves that a compact 7B model can internalize verification entirely within its native reasoning trace through RL, eliminating multi-model deployment overhead.

Rating

  • Novelty: โญโญโญโญโ˜† Pioneers internalizing evidence verification and search calibration into multimodal agent RL finetuning.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluation across 5 benchmarks, in-depth component ablations, tool behavior statistics, and test-time compute scaling.
  • Writing Quality: โญโญโญโญโญ Exceptionally clear framing, rigorous failure analysis, and excellent illustration of reward hacking dynamics.
  • Value: โญโญโญโญโญ Delivers a highly practical blueprint for creating self-contained, tool-augmented multimodal agents without external verifier dependencies.