Learning Consistency in Reward Modeling for Multi-Modal Reasoning¶
Conference: ECCV 2026
Paper: ECCV Official
Full-text Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5859.txt
Area: Multimodal VLM / LLM Reasoning
Keywords: Consistency Reward Model, Multimodal Reasoning, Reinforcement Learning, GRPO, Soft Reward
TL;DR¶
Consistency Reward Models (CRM) reframe multimodal reasoning reward evaluation into chain-of-thought (CoT) guided semantic and mathematical equivalence checking between candidate outputs and ground-truth references in text space, deriving continuous soft rewards coupled with advantage-sign reject sampling to unlock 4× training data and achieve state-of-the-art RL performance.
Background & Motivation¶
Post-training via reinforcement learning (RL) has emerged as the principal catalyst for eliciting multi-step reasoning capabilities in large language models and vision-language models. However, the performance ceiling and empirical stability of policy optimization remain fundamentally constrained by the fidelity of the reward signal. Conventional preference-based reward models trained under the Bradley-Terry objective rely on scalar discriminative heads; in intricate multimodal reasoning contexts, this scalar structure strips the model of the capacity to engage in deliberative chain-of-thought verification and forces small verifiers to estimate the absolute correctness of elaborate derivations, inevitably precipitating reward hacking.
To circumvent the hallucination risks of preference models, recent research has increasingly pivoted toward rule-based reward verification using symbolic engines such as SymPy and math parser tools. Although symbolic systems provide deterministic precision on closed domains, they fail dramatically when encountering semantic and mathematical equivalences—such as alternate units ("7 days" versus "one week"), non-standard LaTeX expressions, or permissible numerical approximations—and mandate brittle formatting requirements. This rigidity results in pervasive false negatives that mistakenly penalize valid reasoning chains and forces data curation pipelines to discard vast volumes of unformatted multimodal data, severely bottlenecking the scale of usable RL supervision.
Addressing the dual failure modes of discriminative scalar models and rigid symbolic engines, this work redefines the foundational premise of reward modeling: a reward verifier does not need to solve the underlying multimodal problem from scratch, but merely needs to verify semantic and mathematical equivalence between a generated response and an authoritative reference. The core idea is to build generative Consistency Reward Models (CRM) operating purely in text space to compare candidate responses against reference answers via chain-of-thought, derive continuous soft rewards from special-token logits, and enforce advantage-sign reject sampling to eliminate GRPO policy oscillation while unlocking 4× uncurated training data.
Method¶
Overall Architecture¶
The core philosophy of CRM shifts absolute correctness evaluation into bidirectional equivalence checking: given multimodal question \(q\), candidate model response \(s\), and ground-truth answer \(g^*\), the reward model learns the mapping \(R(q, s, g^*) = f(q, s, g^*; \theta)\). The operational pipeline encompasses large-scale consensus dataset curation, text-space representation decoupling, generative chain-of-thought consistency judgment, and continuous soft reward computation with sign-consistency sample filtering for downstream policy RL.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multimodal and Math Question Corpus<br/>Numina-Math and MathV360K extraction"] --> B["Strict 3-Judge Consensus 860K Consistency Dataset<br/>13 generation models with GPT-4o and dual Doubao CoT"]
B -->|Supervised Fine-Tuning 3 Epochs| C["Consistency Reward Model CRM-7B"]
D["Multimodal Query Inputs<br/>Question q, candidate response s, ground truth g*"] --> E["Text-Space Verification Architecture<br/>Strip visual tokens, retain context and candidate derivations"]
E --> C
C --> F["Generative CoT Consistency Reasoning Mechanism<br/>Execute semantic analysis, self-reflection, and token emission"]
F --> G["Dual Hard-Soft Rewards and Sign-Consistent Reject Sampling<br/>Derive logit margin reward and filter conflicting advantage signs"]
G --> H["Downstream Multimodal RL Policy Optimization<br/>Drive stable scaling across 7B to 72B VLM policies"]
Key Designs¶
1. Strict 3-Judge Consensus 860K Consistency Dataset: Eliminating Large-Scale Unbiased Supervision Scarcity
Existing judge datasets lack granular supervision tailored for complex multi-step reasoning and diverse vision-language response formats. To establish a robust foundation for consistency modeling, the authors constructed a unified 860K verified training pipeline. Question sourcing aggregates text-based mathematical problems from Numina-Math (filtering out synthetic math and metamath) and 340k multimodal problems from MathV360K. Response collection gathers diverse outputs across 13 representative models—encompassing the Qwen family, Mistral-Large, GPT-4o, DeepSeek-V3, the DeepSeek-R1-Distill series, and QvQ-70B-preview—capturing wide distributions of reasoning depth and error modes. For annotation, GPT-4o performs initial consistency classification, followed by two independent runs of Doubao-1.5-vision-pro with explicit chain-of-thought reasoning; only samples where all three evaluations achieve unanimous agreement are preserved, creating a balanced 1:1 consistent/inconsistent corpus of 860K high-fidelity instances.
2. Text-Space Verification Architecture: Decoupling Multimodal Complexity and Eliminating Visual Noise
While multimodal tasks inherently include visual inputs, empirical results show that training a vision-language reward verifier (Qwen-2.5-VL-7B) achieves 96.7% validation accuracy, which underperforms a pure text-based language model verifier at 98.0%. The underlying reason is that consistency verification centers on symbolic reduction, semantic equivalence, and mathematical validation between textual statements and reference values rather than raw pixel grounding; introducing visual tokens injects cross-modal attention noise and distracts the backbone from precise formal checking. CRM therefore adopts a decoupled text-space architecture: visual image tokens are discarded during reward modeling, preserving only the textual problem formulation, context constraints, candidate answers, and ground-truth references. This design dramatically minimizes inference latency while maximizing the symbolic deduction capacity of the language model backbone.
3. Generative CoT Consistency Reasoning Mechanism: Bypassing Problem-Solving via Analysis and Reflection
Traditional scalar reward verifiers utilize a linear projection head over the final hidden state, precluding multi-step symbolic reduction. CRM implements a fully generative formulation that structures the verification process into a three-stage lightweight chain-of-thought:
- Analysis: Conducts detailed semantic and algebraic comparison between candidate \(s\) and reference \(g^*\) within the problem context, applying mathematical theorems, formula simplification, and unit conversion;
- Reflection: Systematically audits the judgment logic, verifies whether all sub-questions have been addressed, checks for semantic drift in lengthy outputs, and confirms that the verification does not depend on re-solving the question independently;
- Conclusion: Emits a specialized judgment token, <|+|> denoting "Consistent" or <|-|> denoting "Inconsistent".
By structuring verification as an equivalence test rather than an independent solution, a compact 7B model reliably supervises long reasoning outputs from much larger models while delivering fully interpretable audit rationales.
4. Dual Hard-Soft Rewards and Sign-Consistent Reject Sampling: Smoothing Optimization Signals and Correcting Advantage Inversion
Using the predicted probabilities of the specialized conclusion tokens derived from vocabulary logits, CRM calculates both binary hard rewards and continuous soft rewards: $\(r_{\text{hard}} = \begin{cases} 1, & P(\texttt{<|+|>}) > P(\texttt{<|-|>}) \\ 0, & \text{otherwise} \end{cases}\)$ $\(r_{\text{soft}} = P(\texttt{<|+|>}) - P(\texttt{<|-|>})\)$ The continuous metric \(r_{\text{soft}} \in [-1, 1]\) captures granular degrees of semantic alignment and confidence, providing richer learning signals than binary labels. However, when deploying continuous rewards in Group Relative Policy Optimization (GRPO), group-level reward normalization can introduce advantage sign inversion: if all candidate rollouts in a group are fundamentally correct (or all incorrect), small differences in soft reward logits cause below-average yet correct responses to receive negative advantage values (\(A(q, s) < 0\)), mistakenly penalizing valid reasoning. To prevent destructive policy updates, CRM enforces advantage-sign reject sampling: $\(\mathcal{D}_{\text{train}} = \left\{ (q, s) \;\middle|\; \operatorname{sign}(A(q, s)) = \operatorname{sign}(r_{\text{hard}}(q, s)) \right\}\)$ Any rollout sample whose computed advantage sign contradicts its binary hard reward decision is discarded from policy gradient updates, aligning continuous policy optimization with verifier correctness.
Loss & Training¶
The consistency reward model is initialized from Qwen2.5-Instruct (evaluated across 1.5B and 7B variants), configured with a maximum sequence length of 4096 tokens, and trained for 3 epochs using a cosine learning rate schedule with a peak learning rate of 2e-5. Distributed training and inference serving are accelerated via PyTorch FSDP and vLLM. In downstream GRPO training of Qwen-2.5-VL policies, the rollout batch size is set to 512 with 10 candidate responses sampled per query at temperature 1.0, a global training batch size of 1280, and a constant learning rate of 1e-6. By overlapping CRM token generation with reference model log-probability and policy forward evaluations, CRM introduces only a 6% wall-clock latency overhead compared to rule-based execution.
Key Experimental Results¶
Main Results¶
On four multimodal mathematical reasoning benchmarks—MathVista, MathVerse, MathVision, and WeMath—reinforcement learning policies guided by CRM were compared against proprietary frontiers, open-source baselines, and baseline reward frameworks under identical training setups.
| Model | Size | MathVista | MathVerse | MathVision | WeMath | Avg. |
|---|---|---|---|---|---|---|
| Closed-Source Frontiers | ||||||
| GPT-5 | - | 81.9% | 81.2% | 72.0% | 71.1% | 76.6% |
| Gemini 2.5 Pro | - | 80.9% | 76.9% | 69.1% | 78.0% | 76.2% |
| GPT-4o | - | 63.8% | 50.2% | 30.4% | 68.8% | 53.3% |
| Open-Source 7B/8B Baselines | ||||||
| Qwen-3-VL-8B | 8B | 77.2% | 62.1% | 53.9% | 64.2% | 64.4% |
| Qwen-2.5-VL-7B | 7B | 68.2% | 47.9% | 25.4% | 62.1% | 50.9% |
| MM-Eureka-7B | 7B | 73.0% | 50.3% | 26.9% | 66.1% | 54.1% |
| Identical RL Comparisons (Qwen-2.5-VL-7B Base) | ||||||
| Rule-RL-7B-Baseline (MathRuler) | 7B | 72.4% | 52.3% | 28.6% | 67.5% | 55.2% |
| xVerify-RL-7B | 7B | 74.2% | 53.5% | 29.3% | 67.8% | 56.2% |
| CompassVerifier-RL-7B | 7B | 74.4% | 54.8% | 29.2% | 68.6% | 56.8% |
| CRM-RL-7B (Ours, hard reward) | 7B | 74.3% | 56.0% | 29.8% | 69.6% | 57.4% |
| CRM-RL-7B (Ours, soft reward) | 7B | 75.2% | 57.1% | 30.6% | 70.5% | 58.4% |
| Model Scaling (Qwen-2.5-VL-72B Base) | ||||||
| Qwen-2.5-VL-72B (Base without RL) | 72B | 74.8% | 57.6% | 38.1% | 72.4% | 60.7% |
| CRM-RL-72B (Ours, soft reward) | 72B | 76.5% | 61.0% | 41.7% | 74.3% | 63.4% |
On the public verification benchmark VerifyBench-Hard, CRM-7B also demonstrates superior standalone judging accuracy over existing verifier architectures:
| Model / Verifier Method | Numeric (Num) | Expression (Exp) | Multiple Choice (MC) | String (Str) | Overall Accuracy (AVG) |
|---|---|---|---|---|---|
| Qwen3-8B | 68.65% | 78.41% | 73.02% | 66.52% | 70.90% |
| xVerify-8B-I | 69.05% | 76.14% | 93.49% | 81.74% | 83.10% |
| CompassVerifier-7B | 78.97% | 85.23% | 95.35% | 76.09% | 85.90% |
| CRM-7B (Ours) | 79.37% | 84.09% | 96.74% | 84.35% | 88.40% |
Ablation Study & Training Stability¶
Ablation experiments evaluate the impact of Chain-of-Thought (CoT) reasoning on both discrete hard rewards and continuous soft rewards, along with long-horizon training stability across optimization steps.
| Configuration | MathVista | MathVision | Average | Mechanism Analysis |
|---|---|---|---|---|
| Rule-RL-7B-Baseline | 72.4% | 28.6% | 50.5% | Rigid symbolic matching baseline |
| CRM-RL-7B (Hard reward, with CoT) | 74.3% | 29.8% | 52.1% | Discrete decision with step-by-step audit |
| CRM-RL-7B (Hard reward, w/o CoT) | 73.4% (-0.9%) | 28.9% (-0.9%) | 51.2% | Direct classification drops on boundary cases |
| CRM-RL-7B (Soft reward, with CoT) | 75.2% | 30.6% | 52.9% | Best setup; calibrated logit margin guides policy |
| CRM-RL-7B (Soft reward, w/o CoT) | 73.1% (-2.1%) | 28.8% (-1.8%) | 51.0% | Severe overconfidence and variance harm optimization |
Evaluating training stability across optimization steps on MathVerse under identical 220K training data:
| Reward Method | Step-200 | Step-600 | Step-900 | Step-1000 | Step-1200 | Training Dynamics |
|---|---|---|---|---|---|---|
| Rule (MathRuler) | 52.1% | - | - | - | - | Rapid training collapse after 200 steps |
| xVerify | 52.3% | 53.0% | 53.3% | - | - | Performance plateau and termination by 900 steps |
| CompassVerifier-7B | 52.1% | 53.6% | 54.1% | 54.5% | - | Stagnation and instability after 1000 steps |
| CRM (Ours) | 52.3% | 53.7% | 54.8% | 55.4% | 55.7% | Sustained linear improvement through 1200 steps |
Key Findings¶
- CoT reasoning is indispensable for continuous soft reward calibration: While omitting CoT yields a moderate 0.9% drop under hard rewards, it induces a severe 2.1% performance collapse under soft rewards (75.2% down to 73.1%). Analyzing reward distributions across 50 hard misjudged samples reveals that non-CoT models output extreme scores on errors (False Positives average 0.75, False Negatives average -0.63). Integrating CoT sharply reduces variance on erroneous judgments (FP drops to 0.24, FN contracts to -0.39), providing well-calibrated confidence signals.
- Unlocking 4× uncurated data without rule-induced collapse: Rule-based systems collapse after 200 optimization steps because they cannot parse diverse unstandardized multimodal responses from the 220K Llava-OneVision set. In contrast, CRM achieves a 77.8% win rate over VLMEvalKit and maintains continuous performance gains across 1200 steps.
- Text-space judgment outperforms image-conditioned verification: Decoupling visual tokens achieves higher validation accuracy (98.0% vs. 96.7%) and cleaner generalization, proving that multimodal reasoning verification fundamentally operates as a text-space semantic equivalence comparison.
Highlights & Insights¶
- Task Reframing from Problem-Solving to Equivalence Checking: Avoids the architectural paradox of training an expensive outcome verifier that must be stronger than the policy itself, enabling a 7B verifier to reliably supervise a 72B policy model.
- Advantage-Sign Reject Sampling: Identifies the subtle flaw of group reward normalization in GRPO where continuous soft rewards inadvertently penalize uniformly correct rollout batches, resolving it with an elegant sign-filtering criterion.
- Broad Zero-Shot Generalization Across Tasks: Transfers seamlessly to non-mathematical domains without task-specific tuning, attaining 86.9% accuracy on OCRBench (+0.7% RL gain) and expanding to open creative writing via checklist verification (CharXivDQ +3.4%, InfoVQA +2.6%).
Limitations & Future Work¶
- Dependency on Verified Ground-Truth References: The consistency framework strictly requires an authoritative reference answer, precluding direct deployment in fully unsupervised self-play or open-ended exploratory reinforcement learning.
- Tail Truncation Vulnerability on Long-Chain Proofs: To bound serving latency, excessive responses undergo tail truncation, which may misjudge multi-page proofs where the decisive resolution appears at the very end of the output.
- Future Directions: Extending consistency modeling to step-level process reward models (PRMs) under weak references and adapting equivalence checking to 3D spatial reasoning and embodied action trajectories.
Related Work & Insights¶
- vs. Rule-Based Reward Systems (MathRuler / SymPy): Symbolic systems rely on hand-engineered parsing rules, frequently failing on natural language paraphrases and multi-part queries; CRM accommodates diverse syntactic expressions, directly ingesting 4× more data while preventing training collapse.
- vs. Preference Reward Models (Bradley-Terry Formulation): Preference models compress comparisons into scalar values without explicit reasoning steps; CRM employs generative CoT and special-token logit margins to deliver interpretable, smooth reward signals.
- vs. Outcome Reward Models (Gen-ORM / xVerify): Standalone ORMs attempt to solve problems independently from image inputs, leading to high hallucination rates on challenging benchmarks (achieving only a 17.8% win rate against VLMEvalKit); CRM leverages ground-truth answers as anchors to achieve an 89.2% win rate.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates multimodal reward modeling as text-space consistency verification with soft margin rewards and sign-consistent sample filtering.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive 860K dataset curation, standalone judge benchmarks, long-horizon training stability curves, and cross-task generalization.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous methodology, transparent discussion of failure modes, and thorough mathematical formulation.
- Value: ⭐⭐⭐⭐⭐ Removes the data bottleneck imposed by rigid rule systems, serving as vital infrastructure for multimodal reasoning post-training.