Beyond Prediction: Steering VLM Agents with Retrospective World Modeling¶
Conference: NeurIPS2026
arXiv: 2609.39101
Area: Multimodal VLM
Keywords: retrospective world modeling, inverse action attribution, self-consistency reward, multi-turn agents, credit assignment
TL;DR¶
RWM makes VLM agents infer the previous action from textual belief states derived from consecutive observations and converts action-transition consistency into training feedback; the full configuration achieves the paper's reported 0.81 Overall success rate across four task families, but its gains jointly involve retrospective reasoning, Bi-Level GAE, external judging, and SCR rather than SCR alone.
Background & Motivation¶
A vision-language model (VLM) agent must do more than understand an image: it must continue making decisions after its actions change the environment. ReAct alternates reasoning and actions, while VAGEN further asks the model to describe the current state, rehearse a candidate action's consequences, and then produce an executable action. This prospective world modeling supports planning but can also turn the model's own prediction into an assumed fact. In Sokoban, for example, the agent may imagine that moving upward will push a box without noticing that a wall makes the action ineffective. A linguistically coherent reasoning trace can therefore sustain an incorrect environmental assumption.
The issue is not merely inaccurate prediction: realized transitions are underused. Conventional sparse rewards evaluate the trajectory when the task ends and rarely identify which intermediate action disagreed with the environment. Once two consecutive observations are available, the agent can ask a different question: rather than โwhat will this action cause?โ, ask โwhich action best explains this change?โ. Failed actions, wall collisions, and explanations inconsistent with observations can then supply feedback on the next turn instead of waiting for eventual task failure.
Inverse inference is not automatically a reliable judge, however. The previous action is explicitly recorded in the interaction history, so the model can copy it without understanding the transition; different actions may also produce identical observations. The paper therefore combines retrospective reasoning, a history-isolated scoring interface, and training rewards. Its aim is to encourage behaviors better supported by observed changes, not to establish a rigorous proof of physical causality. Core Idea: infer actions inversely from the two belief states derived from consecutive real observations, then supervise cross-turn self-consistency through the executed action's scoring margin over its strongest alternative.
Method¶
Overall Architecture¶
The inputs are the current visual observation, task goal, and interaction context; the output is an executable environmental action. The same VLM policy describes belief states, explains past transitions, performs prospective planning, and produces actions. No separate inverse-dynamics network is trained. Here, the โworld modelโ is the policy's internal language-based reasoning capability, not an oracle for true state transitions.
The system has two related but distinct data flows. The execution flow constructs the current belief when a new observation arrives, retrospectively explains the preceding transition, and then plans and executes the current action. The training flow sends consecutive beliefs to the history-isolated scoring interface, computes SCR for the previous action, and assigns this delayed signal back to the previous turn's reward. It does not verify an action before execution using a predicted state, nor does it compute the reward simply by reading one sampled retrospective answer.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
O["New observation and goal"] --> B["Dual-belief retrospection<br/>and current planning"]
B --> A["Execute current action"]
A --> N["Environment returns next observation"]
N -->|Update belief next turn| B
B -.->|Consecutive beliefs; training scoring| H["History-agnostic scoring"]
H -.-> S["Self-consistency reward"]
S -.-> T["Cross-turn credit assignment"]
T -.->|Training only; update policy| B
Solid arrows denote execution and observation; dashed arrows denote reward computation and training feedback. The four key designs are dual-belief retrospection, history-agnostic scoring, self-consistency reward, and cross-turn credit assignment, in that order. RWM-Base retains retrospective reasoning with sparse-reward PPO; the complete dense-reward training flow belongs to RWM-Full.
Key Designs¶
1. Dual-belief retrospection: explain a realized change before planning the next action
The agent cannot observe the POMDP's true hidden state and instead generates textual belief states from screenshots and context. At the beginning of turn \(t\), it has both the saved previous belief \(\hat{s}_{t-1}\) and a current belief \(\hat{s}_t\) generated from the new observation. Inverse inference uses these two observation-supported descriptions rather than the imagined next state from the previous turn. The paper explicitly regenerates the next belief from the new observation, avoiding the mistake of treating a successful prediction as evidence that reality actually followed it.
The attributed action \(\hat{a}_{t-1}\) is the model's internal explanation of the previous transition, not an action label supplied by the environment or necessarily a unique true cause. It explains the preceding turn. The model then combines the current belief, task goal, and retrospective context to propose a current candidate action, predict its consequences, and produce the current executable action. The first observation has no preceding transition, so its retrospective field contains a null marker and contributes no corresponding retrospective reward.
The reasoning trace contains observation description, retrospection, current-action reasoning, next-state prediction, and the final action in sequence. The paper alternates between <Res> and <Rea> for reasoning in its equations, while the full appendix prompts use <reasoning>. These presentation-level tag differences should not be interpreted as additional modules. The important distinction is between explaining the past and simulating the future: retrospection does not replace prospection, and prospection cannot substitute for the next real observation.
This design lets changes following a failure inform current planning. If the player did not move, for example, the model can revise its assumption that the space ahead is traversable. It does not possess future observations before execution, however, and cannot guarantee prevention of the first incorrect action. It is better understood as cross-turn closed-loop correction than as an a priori safety filter.
2. History-agnostic scoring: expose the transition rather than the previous action
If the retrospective model can read the full interaction history, it only needs to locate the preceding final answer to score the executed action highly. The training reward may improve without learning how actions affect the environment. RWM therefore exposes only the two belief states and the candidate action set to its discriminative scoring interface, not the full action history. The model scores how well every candidate in the finite action space explains the transition rather than generating only one winning action.
Formally, the score vector is \(v_t=F_\theta(\hat{s}_t,\hat{s}_{t+1};\mathcal{A})\), with one component per candidate action. The same policy implements this interface, but its conditioning differs from ordinary long-history decision-making. The <Retro> explanation generated in the normal reasoning trace must be distinguished conceptually from the complete action-score vector used for the reward: a sampled explanation is not the numerical definition of margin-SCR.
An implementation boundary must be retained. Section 4.2 says the scores are derived from candidate action-token logits, whereas Appendix A.2 and the actual prompts require an integer score in \([-5,5]\) for each action. The appendix explicitly sets \(S_{\max}=5\) and \(S_{\min}=-5\), so the normalization constants are known. However, the full text does not clearly specify how raw token logits become bounded integer scores or how multi-token action names are aggregated. It does not justify adding an implementation based on whole-action sequence log-probabilities or token-by-token enumeration of continuous coordinates.
Isolating the full history removes the direct route of copying the previous answer, but does not mathematically eliminate all information leakage. Belief states are themselves generated from observations and context; if their wording includes something like โjust moved right,โ the scorer can still receive indirect action information. The ablation supports the history-agnostic interface's effectiveness but does not separately audit belief-content leakage. This note therefore does not adopt the paper's absolute claim that memory exploitation is completely eliminated.
3. Self-consistency reward: prefer the executed action over its strongest alternative explanation
For a finite discrete action space, SCR compares the executed action's score with the highest score among all other actions. This is more informative than asking whether one sampled retrospective action equals the executed action. Weak support and strong disagreement can yield different rewards even when a sampled explanation is incorrect, and the comparison explicitly incorporates the hardest-to-distinguish alternative.
SCR is positive when the executed action strictly outranks every alternative, negative when another action outranks it, and zero when it ties with the strongest alternative. Under the appendix's bounded integer scoring, the denominator is \(10\) and SCR lies in \([-1,1]\). This boundedness requires scores to satisfy the declared bounds and cannot be transferred directly to unprocessed raw logits. Extrinsic rewards still evaluate task success, format compliance, and step costs. SCR is a regularizer: an easily explained action that fails to complete the task should not dominate one that genuinely advances the goal.
Another comparison applies softmax to the scores and subtracts the uniform policy's log-probability from the executed action's log-probability. It shares discriminative scoring with the margin method but transforms the scores differently: a low-probability executed action can incur a negative penalty much larger than the benefit of a correct action. Direct token matching instead requires no scoring and supplies a binary signal. The appendix describes an abstract log-likelihood range extending toward negative infinity; with a finite action set and scores strictly bounded in \([-5,5]\), the implemented negative bound is finite. Its abstract argument should not be restated as inevitable unboundedness in that specific implementation.
The discrete margin formula is not transferred unchanged to coordinate-valued actions. ManiSkill actions combine types such as pick, place, and push with continuous spatial parameters. Appendix C.2 instead retrospectively infers the type and coordinates: the types must match, after which the Euclidean distance between attributed and executed coordinates determines spatial consistency.
Here, \(D_{\max}\) denotes the maximum distance. The authors do not provide its numerical value, coordinate-normalization details, or the construction of a parameter vector for push's two coordinate triples. The supported description is therefore type gating plus a distance penaltyโnot exhaustive finite-candidate token-margin ranking of continuous actions. No clipping rule should be added. This is a task-adapted SCR variant, not a proven equivalent implementation of the discrete scoring interface.
4. Cross-turn credit assignment: evidence arrives next turn, but the reward belongs to the previous action
The consequence of the action at turn \(t\) can only be scored after the observation at turn \(t+1\) arrives. SCR computed at that point belongs to \(a_t\), not the action \(a_{t+1}\) currently being prepared. Assigning it to the latter would update an unrelated action for support or disagreement concerning the preceding transition. RWM explicitly assigns the signal back to the turn that produced that transition.
Turn-level rewards also cannot directly update every generated token. RWM-Full adopts VAGEN's Bi-Level GAE: the outer level uses turn rewards and the critic to estimate turn advantages, then injects this semantic feedback into the final generated token of the corresponding turn. The inner level propagates it backward through that turn's token sequence. Generated observation descriptions, retrospection, planning, and actions can thus receive feedback aligned with that turn's outcome rather than relying only on sparse success rewards at the trajectory endpoint.
This mechanism aligns rewards and advantages in time; it is not a causal decomposition proving the true contribution of every reasoning sentence. Retrospection and prospection share a model, so beliefs, explanations, and actions can jointly adapt to the reward. High SCR does not automatically establish agreement with external physical laws. Training gains should be judged through task success and action effectiveness together, not merely increasing internal rewards.
A Worked Example¶
In the Sokoban case from Appendix E, the box initially lies above and to the right of the player, with the target above the box. The model plans to move right so that the player stands directly below the box, and executes Right. The next observation indeed describes the box as directly above the player, so retrospection uses the previous belief and this new belief rather than reusing the earlier prediction as input.
The case gives scores [-1,0,0,5] in action order [Up, Down, Left, Right]. Right scores \(5\), while the strongest alternative scores \(0\). Equation (5) therefore yields a raw SCR of \((5-0)/10=0.5\) for this transition. With Sokoban's configured \(\alpha=0.1\), the SCR contribution added to the preceding turn's reward is \(0.05\), not the total turn reward.
The current plan then changes to Up to push the box. On the following turn, retrospective scores [5,0,-5,-5] support the Up action just executed; another Up subsequently pushes the box onto the target. The case illustrates that retrospective rewards arrive one turn after the action and that a displayed score vector is not a single sampled Retro action. The SCR arithmetic above is calculated from the paper's formula and case, not an additional experimental measurement.
Loss & Training¶
RWM-Base uses the reasoning structure with a retrospective field but only sparse task and format rewards, with conventional token-level GAE; it does not include dense SCR. RWM-Full jointly adds Bi-Level GAE, external LLM-as-a-Judge evaluation of visual state estimations and predictions, and intrinsic SCR. The external judge is GPT-4.1 nano, so this component cannot be described as having no external supervision. โIntrinsicโ accurately characterizes SCR's source only.
The policy is updated through PPO clipping, and the critic is a scalar value head on the VLM's final hidden state. A binary mask excludes environmental prompts and computes policy gradients only for model-generated reasoning and action tokens. The appendix also states that inner token TD-errors typically include a KL penalty relative to a reference model, but does not fully enumerate all optimization hyperparameters.
The backbone is Qwen2.5-VL-3B, with a global batch size of \(128\), actor/critic learning rates of \(10^{-6}\) and \(10^{-5}\), and \(4\) H800 80GB GPUs. Evaluation uses temperature and top-p of \(1.0\), at most \(1024\) new tokens, and a maximum of \(4\) interaction turns. These short-turn experiments do not directly establish performance in interactions lasting hundreds of steps.
Key Experimental Results¶
Main Results¶
The table below selects success rates from Table 1, retaining task-family averages and its reported Overall column. All trained methods use Qwen2.5-VL-3B. Sokoban averages Standard/Hard, Navigation averages Base/Common Sense, and ManiSkill averages Place/Stack/Drawer/Align.
| Config | Sokoban average | Navigation average | ManiSkill average | FrozenLake | Overall (reported) |
|---|---|---|---|---|---|
| Qwen2.5-VL-3B | 0.06 | 0.25 | 0.00 | 0.09 | 0.08 |
| Vanilla-PPO | 0.16 | 0.29 | 0.00 | 0.21 | 0.12 |
| ReAct-RL | 0.24 | 0.38 | 0.00 | 0.30 | 0.17 |
| VAGEN-Base | 0.37 | 0.49 | 0.72 | 0.43 | 0.56 |
| RWM-Base | 0.52 | 0.66 | 0.94 | 0.75 | 0.76 |
| VAGEN-Full | 0.48 | 0.55 | 0.94 | 0.54 | 0.71 |
| RWM-Full | 0.66 | 0.68 | 0.97 | 0.77 | 0.81 |
Reported Overall values are not the simple averages of the four task-family columns shown here, and the text does not clearly specify aggregation weights. For example, RWM-Full's four-column arithmetic mean is \(0.77\), whereas its reported Overall is \(0.81\). This note preserves the published figures without correcting them or reinterpreting \(0.81\) as an equal-weight mean across the four families.
Within matched configurations, RWM-Base improves the reported Overall over VAGEN-Base by \(20\) percentage points, and RWM-Full improves over VAGEN-Full by \(10\) percentage points. On Sokoban-Hard, the Base comparison is \(0.46\) versus \(0.33\), and the Full comparison is \(0.63\) versus \(0.46\). The Base results support the retrospective structure's usefulness, but because Base has no dense SCR, its \(20\)-point gain is not SCR's isolated contribution.
Ablation Study¶
The following table comes from Table 2 and evaluates history-leakage regulation. Its title does not explicitly identify Base or Full, so the configuration names are retained without assigning an additional training setting.
| Config | Sokoban | FrozenLake | Note |
|---|---|---|---|
| Qwen2.5-VL-3B | 0.06 | 0.09 | Untrained backbone |
| VAGEN | 0.41 | 0.43 | Prospective control |
| RWM, regulation-free history | 0.39 | 0.40 | Previous action accessible directly from history |
| RWM, action dropout, 0.10 | 0.44 | 0.52 | Mild masking of action-related tokens |
| RWM, action dropout, 0.20 | 0.50 | 0.59 | Moderate masking |
| RWM, action dropout, 0.50 | 0.33 | 0.38 | Heavy masking disrupts contextual coherence |
| RWM, history-agnostic scoring | 0.58 | 0.75 | Scoring conditioned only on the transition |
Key Findings¶
- History-agnostic scoring improves over regulation-free history by \(19\) and \(35\) percentage points. Giving a model access to the answer does not teach inverse inference; input isolation is more reliable than merely prompting it to ignore the previous action.
- The Sokoban weight analysis reports that \(\alpha=0.50\) lowers success from \(0.58\) to \(0.29\) while the intrinsic reward saturates earlier. Excessive SCR lets easily explained behavior dominate task objectives.
- The text reports Sokoban action effectiveness above \(0.95\), compared with VAGEN's peak near \(0.8\). This supports reducing ineffective moves, but is task-metric evidence rather than independent validation of physical causality.
- Averages over three evaluations and \(128\) test cases are not accompanied by confidence intervals or significance tests. Approximate curve values should not be treated as exact endpoints. Proprietary models are evaluated zero-shot, so the comparison concerns task-specialized training versus general zero-shot use, not equal training budgets.
Highlights & Insights¶
- Turning โcan the outcome explain the action?โ into training feedback is closer to interaction quality than checking action-string validity alone. The strongest-alternative margin also expresses competition among explanations rather than whether one sample happened to match.
- Assigning rewards to the turn that actually produced the transition is a reusable engineering principle. Agent rewards based on subsequent observations must resolve temporal ownership before token-level training is considered.
- The ablations show that the information interface is part of reward design. If the scorer directly sees the answer, a sophisticated reward formula cannot prevent shortcuts.
Limitations & Future Work¶
- The authors identify inverse-mapping ambiguity in stochastic environments and multi-action turns: UpโLeft and LeftโUp can yield the same terminal state. Supporting only one explanation can incorrectly penalize another valid path, motivating set-based retrospection.
- The authors acknowledge memory and OOM bottlenecks for long-turn RL. The current maximum of \(4\) turns, single-box Sokoban, and non-slippery FrozenLake provide limited evidence for genuinely long-horizon or stochastic dynamics.
- Beliefs compress observations but can also carry hallucinations and leaked action information. With the same policy generating beliefs and scores, coordinated reward accommodation remains possible, motivating independent observation verification or belief-content audits.
- The correspondence between logits and integer scores, multi-token action scoring, and ManiSkill coordinate handling are underspecified. Reproduction requires clarifying these interfaces rather than inferring a unified full-action-space implementation from the formulas alone.
Related Work & Insights¶
- vs VAGEN: VAGEN emphasizes observation and future prediction; RWM adds inverse action explanations of realized transitions. Full inherits VAGEN's Bi-Level GAE and external judging, so the entire training infrastructure is not a new contribution.
- vs ReAct: ReAct alternates free-form reasoning and actions; RWM gives past transitions and future simulations explicit roles and turns past consistency into training feedback.
- vs inverse-dynamics curiosity methods: Classical methods learn action-relevant representations with inverse models and encourage exploration through forward prediction error. RWM regularizes policy behavior with inverse-attribution consistency rather than rewarding novelty.
Rating¶
- Novelty: 4/5 โ Embedding inverse action attribution in multi-turn VLM reasoning and rewards is distinctive, while building on inverse dynamics and world modeling.
- Experimental Thoroughness: 3/5 โ Four task families and leakage ablations are useful, but short-turn budgets, aggregation conventions, and reproduction interfaces limit the evidence.
- Writing Quality: 3/5 โ The cross-turn mechanism is clear; scoring implementation and some absolute causal claims lack precision.
- Value: 4/5 โ Useful for delayed-observation rewards and input isolation, without treating internal consistency as a guarantee of true causality.