Wan-R1: Verifiable-Reinforcement Learning for Generalizable Video Reasoning¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: EventHosts
Area: Multimodal VLM
Keywords: Video Reasoning / Reinforcement Learning / Flow-GRPO / Verifiable Rewards / Embodied Navigation
TL;DR¶
Wan-R1 adapts Flow-GRPO to flow-matching video generation and establishes verifiable trajectory-grounded and embedding-level reference-anchored rewards, overcoming multimodal reward hacking and substantially advancing out-of-distribution spatial reasoning in complex 3D mazes and robotic path planning.
Background & Motivation¶
Recent breakthroughs in generative video models—such as Sora, Veo, and Wan—have demonstrated remarkable fidelity in synthesizing dynamic physical scenes. This progress has catalysed an emerging paradigm: reasoning via visual generation. Instead of generating textual chains of thought with discrete tokens, video models can directly formulate spatial relationships, kinematic dynamics, and temporal causality within continuous visual rollouts, holding immense promise for long-horizon spatial reasoning, maze navigation, and mapless robotic path planning.
However, existing video reasoning systems predominantly depend on supervised fine-tuning (SFT) over expert demonstration videos. This paradigm reveals acute vulnerability under out-of-distribution (OOD) scenarios: models tend to memorize surface-level geometric textures and fixed coordinate shortcuts rather than acquiring generalizable spatial topological understanding. When tested on complex 3D structures, irregular non-grid curves, or unseen visual skins, the exact match accuracy of SFT models drops catastrophically. Reinforcement learning (RL) offers an intuitive avenue to encourage exploratory trial-and-error beyond superficial imitation. Yet, applying RL to video generation faces a severe reward design dilemma: learned multimodal reward models (VLMs) suffer from severe reward hacking—the generator readily synthesizes plausible-looking, smooth camera rollouts that trick the evaluator without actually solving the navigation problem—while naive sparse binary completion rewards fail to provide sufficient gradient signals within high-dimensional continuous latent spaces.
To resolve this impasse, this work discards opaque neural evaluators and anchors policy gradients in strictly verifiable geometric ground truth and continuous reference trajectories. Core idea: adapt Flow-GRPO to flow-matching image-to-video architectures, construct a multi-component verifiable trajectory reward decoupling exact match, consecutive precision, and structural fidelity, and introduce an embedding-level reference-anchored reward for continuous robotic navigation, achieving stable, hacking-free visual reasoning optimization.
Method¶
Overall Architecture¶
Wan-R1 addresses visual trace reasoning (VTR) and embodied navigation tasks. Conditioned on a task prompt \(c\) and an initial observation frame \(I_0\), the model generates a temporally coherent video rollout \(V = \{f_1, f_2, \dots, f_N\}\) depicting an agent traversing from start to goal. The pipeline integrates a stochastic flow-matching solver with group relative advantage estimation, leveraging denoising-step reduction to accelerate online rollouts, tracking the agent's spatial trajectory, and scoring it against ground-truth paths to compute verifiable rewards for policy gradient updates.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Conditions<br/>Initial Frame I0 + Task Prompt c"] --> SDE["Flow-SDE Stochastic Sampling & Step Reduction<br/>ODE to SDE conversion | Strain=30 fast generation of G video candidates"]
SDE --> Rollouts["Candidate Video Group {V1, V2, ..., VG}"]
Rollouts --> EvalBranch{Task Domain Branch}
EvalBranch -->|Grid/3D/Regular Mazes| RewGame["Multi-Component Verifiable Trajectory Reward<br/>Extract continuous trajectory via tracking | Compute REM + RPR + RMF"]
EvalBranch -->|Embodied Quadruped Navigation| RewRobot["Embedding-Level Reference-Anchored Reward<br/>Frozen vision encoder | Compute Rcos + Rtemp + Rend"]
RewGame --> Adv["Group Relative Policy Optimization & KL Constraint<br/>Within-group mean-std advantage normalization | KL penalty to reference policy"]
RewRobot --> Adv
Adv --> Update["LoRA Policy Gradient Update<br/>Clipped surrogate objective optimizing flow velocity network"]
Key Designs¶
1. Flow-SDE Stochastic Sampling & Step Reduction: overcoming deterministic exploration barriers and computation bottlenecks Standard inference in flow-matching models relies on a deterministic ordinary differential equation (ODE): $\(d x_t = v_\theta(x_t, t) dt\)$ Since deterministic ODE trajectories prevent the exploration required for RL trial-and-error, Wan-R1 reformulates the sampling trajectory into an equivalent stochastic differential equation (SDE) as in Flow-GRPO: $\(d x_t = \left( v_\theta(x_t, t) - \frac{\sigma_t^2}{2} \nabla \log p_t(x_t) \right) dt + \sigma_t d w\)$ The discretized rollout integrates score-corrected velocity \(\tilde{v}_\theta\) and Gaussian perturbation \(\sigma_t \sqrt{\Delta t} \epsilon\), enabling the model to sample a group of \(G\) diverse rollouts \(\{V^i\}_{i=1}^G\) per problem instance. Furthermore, to overcome the prohibitive computational overhead of multi-step video denoising, the framework introduces a training-time denoising reduction strategy: utilizing a compressed \(S_{\text{train}} = 30\) denoising steps during online RL trajectory collection while retaining the full \(S_{\text{infer}} = 50\) steps during final evaluation. Empirical results show this acceleration preserves sample quality while providing controlled stochasticity that acts as a beneficial exploration regularizer.
2. Multi-Component Verifiable Trajectory Reward: balancing exploration density, path closure, and structural fidelity In structured game environments (VR-Bench), relying on vision-language reward models triggers catastrophic reward hacking. Wan-R1 initializes an object tracker on the agent's bounding box in the initial frame and tracks its center coordinates across all frames, extracting the discrete predicted path \(\tau_{\text{pred}} = \{p_1, p_2, \dots, p_N\}\). This path is objectively evaluated against the ground-truth optimal solution \(\{v_{ij}\}_{j=1}^{n_i}\) via three disentangled components: - Exact Match Reward (\(R_{\text{EM}}\)): strictly verifies whether every step along the predicted trajectory matches the complete optimal path: $\(R_{\text{EM}} = \prod_{j=1}^{n_i} \mathbb{I}(\hat{v}_{ij} = v_{ij})\)$ - Precision Reward (\(R_{\text{PR}}\)): measures the proportion of consecutively correct steps along the optimal path, providing smooth, dense gradient guidance that prevents the vanishing-gradient failure of purely sparse rewards during early exploration. - Maze Fidelity Reward (\(R_{\text{MF}}\)): penalizes visual cheating where the generative model removes, shifts, or melts maze walls in later frames to bypass blocked passages. It measures pixel-level static consistency across the background regions of frame \(I_0\) and subsequent sampled frames \(I_m\).
The combined game reward is defined as: $\(R = \alpha R_{\text{EM}} + \beta R_{\text{PR}} + \gamma R_{\text{MF}}\)$ Configured with \(\alpha = 0.3, \beta = 0.5, \gamma = 0.2\), prioritizing the dense precision reward while strictly requiring global path correctness and physical consistency.
3. Embedding-Level Reference-Anchored Reward: resolving continuous visual supervision without discrete actions In real-world robotic navigation (Target-Bench), supervision is provided as an expert demonstration rollout \(V^* = \{f_1^*, \dots, f_{T^*}^*\}\) recorded via teleoperation rather than a discrete grid sequence. Wan-R1 introduces an embedding-level reference-anchored reward leveraging a frozen vision encoder \(\phi\) (such as DINOv2 or CLIP) with \(\ell_2\)-normalized representations: \(e_t = \text{norm}(\phi(f_t))\). After linearly interpolating the shorter sequence to a unified length \(T_c\), three geometric metrics are computed: - Mean Frame Cosine Similarity (\(R_{\text{cos}}\)): evaluates perceptual alignment across the entire visual journey: $\(R_{\text{cos}} = \frac{1}{T_c} \sum_{t=1}^{T_c} \langle e_t, e_t^* \rangle\)$ - Temporal Consistency (\(R_{\text{temp}}\)): computes normalized cumulative displacement discrepancies in feature space, compelling the synthesized motion to match the chronological progress of the demonstrator and penalizing standing still or wandering. - Endpoint Fidelity (\(R_{\text{end}}\)): measures boundary alignment at the start and terminal frames to guarantee final arrival at the target.
The unified embedding reward is formulated as: $\(R_{\text{emb}} = \alpha_{\text{emb}} R_{\text{cos}} + \beta_{\text{emb}} R_{\text{temp}} + \gamma_{\text{emb}} R_{\text{end}}\)$ with \((\alpha_{\text{emb}}, \beta_{\text{emb}}, \gamma_{\text{emb}}) = (0.5, 0.2, 0.3)\), providing robust, objective guidance for continuous physical planning.
4. Group Relative Policy Optimization & KL Constraint: eliminating value networks and curbing degenerate policies To avoid maintaining a separate heavy value network (Critic) alongside the diffusion transformer, Wan-R1 normalizes advantages within the candidate group: $\(\hat{A}^i = \frac{R(V^i, c) - \text{mean}\left(\{R(V^j, c)\}_{j=1}^G\right)}{\text{std}\left(\{R(V^j, c)\}_{j=1}^G\right)}\)$ This group normalization cancels out global problem difficulty bias, ensuring that a rollout is rewarded solely for outperforming its peers on the exact same maze. A KL divergence penalty \(\beta_{\text{KL}}\) against the reference policy \(\pi_{\text{ref}}\) is incorporated into the surrogate objective. Ablation experiments demonstrate that setting \(\beta_{\text{KL}} = 0.1\) effectively curbs policy drift and path redundancy, yielding efficient, direct trajectories.
Key Experimental Results¶
Main Results¶
On the VR-Bench benchmark spanning five maze paradigms—Regular Maze (Base), Irregular Maze (Irreg), TrapField (Trap), 3D Maze (3D), and Sokoban (Soko)—Wan-R1 is evaluated against leading proprietary video models, open-source base models, and the supervised fine-tuned baseline (Wan-SFT).
| Model Category | Model | Base EM (↑) | Irreg EM (↑) | Trap EM (↑) | 3D EM (↑) | Soko EM (↑) | Base SD (↓) | 3D SD (↓) |
|---|---|---|---|---|---|---|---|---|
| Closed-Source | Veo-3.1-pro | 0.0 | 4.2 | 1.4 | 0.0 | 0.0 | 140.7 | 40.1 |
| Closed-Source | Sora-2 | 1.4 | 5.6 | 0.0 | 0.0 | 4.2 | 302.9 | 92.4 |
| Closed-Source | Kling-v1 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 162.3 | 84.4 |
| Closed-Source | MiniMax-Hailuo-2.3 | 0.0 | 1.4 | 2.8 | 0.0 | 0.0 | 464.0 | 50.1 |
| Open-Source | Wan2.2-TI2V-5B | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 388.7 | 5.4 |
| SFT Baseline | Wan-SFT | 33.3 | 56.9 | 38.9 | 65.3 | 4.2 | 10.3 | 3.9 |
| Ours | Wan-R1 | 61.1 | 47.9 | 90.3 | 94.4 | 30.6 | 5.2 | 1.9 |
| Gain | vs. Wan-SFT | +27.8 | -9.0 | +51.4 | +29.1 | +26.4 | -5.1 | -2.0 |
On Target-Bench for real-world quadruped navigation, Wan-R1 achieves the highest overall score among all models, reducing Average Displacement Error (ADE) and Final Displacement Error (FDE) by approximately half relative to the base model while drastically cutting Miss Rate (MR).
Ablation Study¶
Ablation experiments analyze the sensitivity of the KL regularization penalty \(\beta_{\text{KL}}\) and training denoising steps \(S_{\text{train}}\) on path optimality and sample efficiency:
| Setting Category | Configuration | Exact Match EM (%) | Success Rate SR (%) | Precision Rate PR (%) | Step Deviation SD (↓) | Note |
|---|---|---|---|---|---|---|
| KL Regularization | \(\beta_{\text{KL}} = 0\) | 62.5 | 75.0 | 79.0 | 6.7 | Omitting KL penalty induces wandering and high path redundancy |
| KL Regularization | \(\beta_{\text{KL}} = 0.04\) | 61.1 | 75.0 | 78.6 | 5.2 | Moderate regularization constrains trajectory deviation |
| KL Regularization | \(\beta_{\text{KL}} = 0.1\) | 66.7 | 75.0 | 81.5 | 3.7 | Optimal trade-off anchoring rollouts close to reference manifold |
| Denoising Reduction | \(S_{\text{train}} = 5\) | 59.7 | 70.8 | 77.5 | 4.0 | Overly aggressive step reduction degrades rollout perceptual fidelity |
| Denoising Reduction | \(S_{\text{train}} = 30\) | 61.1 | 75.0 | 78.6 | 5.2 | Retains full-step accuracy while multiplying rollout throughput |
| Denoising Reduction | \(S_{\text{train}} = 50\) | 62.5 | 73.6 | 79.7 | 2.9 | Full steps yield marginal gain at massive compute cost |
Key Findings¶
- Substantial gains on spatial depth and obstacle avoidance: Wan-R1 raises exact match from 65.3% to 94.4% on stereoscopic 3D mazes, and from 38.9% to 90.3% (+51.4%) on TrapField with near-perfect precision (97.7%), proving that trial-and-error RL successfully eliminates blind wall-penetration habits common in SFT.
- Mandatory SFT cold-start prerequisite: Initiating Flow-GRPO directly from the raw Wan2.2 base model yields zero reward due to the strictness of verifiable metrics. SFT serves as an indispensable bootstrap placing the policy into an informative reward basin.
- Confirmation of VLM reward hacking: Training against Qwen2.5-VL evaluations causes the policy to generate convincing forward motions that blindly enter cul-de-sacs, confirming that visual plausibility frequently fools neural judges in video spaces.
Highlights & Insights¶
- Verifiable grounding insulates RL from visual deception: Extracting explicit agent coordinates via object tracking completely circumvents the vulnerability of learned vision-language evaluators to superficial camera dynamics.
- Seamless coupling of Flow-SDE sampling and step reduction: Combining stochastic SDE exploration with 30-step rollout collection renders online RL fine-tuning of 5B-parameter diffusion video architectures computationally practical.
- Transfer from discrete benchmarks to physical robotic control: Formulating reference rollouts via DINOv2 cosine similarity and endpoint fidelity bridges the gap between synthetic grid games and continuous real-world quadruped navigation.
Limitations & Future Work¶
- Reliance on structured ground truth: The current framework requires clean object tracking or synchronized reference videos, limiting its out-of-the-box applicability to open-ended generative tasks lacking formal verifiers.
- False negative penalties from single reference paths:VR-Bench rewards trajectories strictly against a single BFS shortest path, unjustly penalizing alternative paths of equal length. Expanding to all-shortest-path verification is a clear future improvement.
- Performance attenuation on continuous curves: On irregular curved mazes without coordinate grid alignments, rigid trajectory matching struggles with minor curve discrepancies, slightly trailing SFT on the hardest subset.
Related Work & Insights¶
- vs. Flow-GRPO (Liu et al., 2025): Flow-GRPO pioneered applying GRPO to diffusion flow matching for image aesthetic alignment; Wan-R1 extends this foundation into multi-step spatiotemporal video planning with verifiable physical feedback.
- vs. VR-Bench (Yang et al., 2025): VR-Bench introduced maze-solving benchmarks for video generation under an SFT paradigm; Wan-R1 proves that online RL with objective rewards is essential to overcome SFT distribution-shift brittleness.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering implementation of verifiable RL and anti-hacking reward frameworks for video generative reasoning.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Thorough evaluation across five maze topologies, multiple difficulty levels, unseen texture skins, and real-world Target-Bench robotic setups.
- Writing Quality: ⭐⭐⭐⭐⭐ Cohesive theoretical motivation, insightful analyses of VLM reward vulnerabilities, and balanced discussion of engineering trade-offs.
- Value: ⭐⭐⭐⭐⭐ Establishes a concrete path forward for transitioning video generative models from visual synthesizers into robust world simulators and reasoning agents.