PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided VLM¶
Conference: ECCV 2026
Paper: ECCV 2026 Official
Code: https://mengsiwei.github.io/MAGIC/
Area: Multimodal VLM / 3D Vision
Keywords: Dynamic 3D Generation, Physical Simulation, Vision-Language Models, Material Point Method, Active Motion Probing
TL;DR¶
Addressing the severely under-constrained nature of inferring physical properties from a single static image, PhyMAGIC introduces a training-free framework that actively probes physical properties by generating motion videos with an I2V model, iteratively refines parameters via a confidence-guided VLM loop, and simulates physically consistent 3D dynamics using a differentiable Material Point Method (MPM) solver.
Background & Motivation¶
Significant progress has been made in generating static 3D assets with high geometric and textural fidelity from a single image, yet extending these assets to physically consistent dynamics remains an open and formidable challenge. Modern video diffusion models can synthesize visually compelling motion trajectories, but they lack explicit physical grounding over critical attributes such as mass, elasticity, Poisson's ratio, and friction. Consequently, their synthesized video sequences frequently suffer from non-physical artifacts, including momentum violations, severe object interpenetration, and distorted material deformation responses. Existing physics-aware generative paradigms either embed hard-coded physical priors directly into network architectures or rely on supervised learning to estimate properties from datasets. However, the former restricts generalization to novel materials and scenes, while the latter heavily depends on scarce physical ground-truth annotations and remains highly vulnerable to perceptual ambiguity.
The root bottleneck lies in the fact that physical property inference from a single static image is fundamentally under-constrained. Across diverse real-world objects, identical static appearances can conceal vastly different densities, Young's moduli, or yield stresses. These motion-critical governing parameters are completely invisible in a static snapshot and are only manifested when an object undergoes dynamic deformation and external force interactions. While foundation vision-language models (VLMs) possess intuitive physical commonsense, they typically operate in a passive, single-forward-pass prediction manner. When visual cues in a single image are ambiguous or sparse, VLMs inevitably produce ungrounded hallucinations without any mechanism for self-verification, and their outputs cannot directly drive numerical physics simulators.
To resolve the fundamental tension between static perceptual ambiguity and dynamic physical determinism, this paper reformulates the paradigm: rather than passively predicting properties from limited static pixels, the system actively initiates "motion probe experiments." Different dynamic motions reveal complementary physical cues—for instance, free fall and impact expose elasticity and mass, whereas bending or spinning reveals stiffness and rigidity limits. Core idea: reformulate single-image physical inference as an active video evidence acquisition process by synthesizing targeted motion probes via an image-to-video diffusion model, iteratively refining physical parameters through a confidence-guided VLM feedback loop, and compiling them into hybrid physical specifications that drive a differentiable Material Point Method (MPM) simulator.
Method¶
Overall Architecture¶
PhyMAGIC consists of three tightly coupled components: Motion Probe Generation, VLM-Based Physical Reasoning, and Physics-Grounded 3D Dynamics. Starting from a single input image and a text instruction, PhyMAGIC uses a pretrained image-to-video (I2V) diffusion model to actively synthesize dynamic probe videos, transforming latent material properties into explicit motion trajectories. Next, a VLM analyzes the probed video sequences over multiple self-consistency passes to evaluate property estimates alongside confidence scores; any attribute falling below a calibrated confidence threshold triggers automated prompt refinement to gather complementary dynamic evidence. Finally, once all attributes converge, the VLM infers simulator-native descriptors and combines them with material parameters into a unified Hybrid Physical Parameters (HPP) specification. This specification initializes 3D Gaussian particles in a differentiable Material Point Method (MPM) solver, simulating high-fidelity, physically consistent 3D dynamics.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Single Static Image + Instruction"] --> B["Motion Probe Generation<br/>Synthesize dynamic video with CogVideoX-I2V"]
B --> C["Confidence-Driven Prompt Refinement<br/>GPT-4o self-consistency scoring & low-conf gating"]
C -->|Attributes below confidence threshold| B
C -->|All parameters reach confidence threshold| D["Hybrid Physical Specification<br/>Unify material attributes & simulator descriptors into HPP"]
D --> E["MPM-Based 3D Dynamic Simulation<br/>3DGS particles transfer state across grid-particle steps"]
E --> F["Rendered Physically Consistent Dynamic 3D Video"]
Key Designs¶
1. Motion Probe Generation: Externalizing Latent Properties into Video Cues Static images cannot directly convey whether an object is hollow, stiff, or rubbery, but these physical properties become visually discriminable under active mechanical perturbations such as bouncing, colliding, or stretching. Instead of treating the generated video as an infallible ground-truth simulation, PhyMAGIC repurposes an open-source, resource-efficient image-to-video diffusion model (CogVideoX-I2V) as an "active physical probing instrument." Given an input image \(I_0\) and a prompt, CogVideoX generates a video sequence \(V_0 = \{I_0, I_1, \dots, I_T\}\). To remove non-informative static background frames while preserving high-information physical transitions, PhyMAGIC applies a motion-aware subsampling strategy driven by optical flow magnitude, extracting key dynamic frames that highlight maximum acceleration and deformation as temporal evidence for downstream VLM reasoning.
2. Confidence-Driven Prompt Refinement: Closed-Loop Uncertainty Reduction Direct single-pass VLM predictions often suffer from prompt-motion misalignment or insufficient visual evidence, leading to catastrophic errors in parameter estimation. PhyMAGIC introduces a structured, coarse-to-fine self-consistency evaluation mechanism. A foundation VLM (GPT-4o) evaluates the input image and probe video across up to three independent stochastic samples to estimate physical attributes \(P_i\) (such as density \(\rho\), Young's modulus \(E\), Poisson's ratio \(\nu\), and yield stress \(\sigma_y\)). The agreement ratio across samples serves as the attribute confidence score \(c_i \in [0, 1]\). Using a calibrated threshold \(\gamma = 0.6\), uncertain attributes are selected via a binary gating mask:
For low-confidence attributes where \(m_i = 1\), the VLM automatically refines the descriptive prompt: \(p^{(t+1)} = \text{VLM}(p^{(t)}, \{P_i \mid m_i = 1\})\). The refined prompt explicitly specifies dynamic actions designed to expose the ambiguous parameters (e.g., requesting a drop impact if elasticity is ambiguous, or a lateral collision if friction and yield stress are uncertain). The newly generated probe video provides complementary evidence, and this iterative loop continues until all parameters exceed \(\gamma\) or the iteration budget is exhausted.
3. Hybrid Physical Specification & Differentiable MPM Simulation: Bridging Semantics and Numerical Dynamics A major conceptual barrier between multimodal reasoning and physics simulation is the severe abstraction gap. A VLM produces high-level descriptions (e.g., "a rigid body of density \(1200\text{ kg/m}^3\)"), but a numerical physics engine cannot execute without an initial-boundary value formulation including boundary conditions, force vectors, and initial velocities. Conversely, specifying boundary conditions alone without constitutive parameters yields invalid simulations. PhyMAGIC bridges this gap with a unified Hybrid Physical Parameters (HPP) specification:
where \(\mathcal{P}_{\text{mat}} = \{\rho, E, \nu, \sigma_y\}\) represents material properties normalized into standard SI units, and \(\mathcal{D}_{\text{sim}} = \{v_0, B, f_{\text{ext}}, \dots\}\) denotes simulator-native descriptors (such as initial velocity vectors, bounding-box collision planes, and contact modes). The VLM selects structured motion templates (e.g., free-fall, compression, spinning, rebound, side-push) based on the instruction and probe context. Concurrently, the single image is reconstructed into a 3D Gaussian Splatting (3DGS) field via Trellis, where each Gaussian kernel is converted into a physical particle \(\mathcal{G}_p^t = (x_p^t, \Sigma_p^t, \alpha_p, c_p, \theta_p^t)\) carrying the HPP state. These particles are simulated in a differentiable Material Point Method (MPM) solver implemented in NVIDIA Warp, running particle-to-grid (P2G) and grid-to-particle (G2P) transfers to enforce continuous momentum and mass conservation.
Loss & Training¶
PhyMAGIC is an entirely training-free pipeline. It requires no task-specific fine-tuning, no LoRA adaptors, and no dataset-specific backpropagation. All foundational models remain frozen: 3D geometry is reconstructed using pretrained Trellis; motion probes are generated via frozen CogVideoX-5B; physical reasoning is executed via GPT-4o API calls; and dynamic simulation runs directly in a numerical MPM solver over 128 to 200 simulation steps on a single NVIDIA RTX 4090 GPU.
Key Experimental Results¶
Main Results¶
The authors evaluate PhyMAGIC across benchmark scenarios from PhysGaussian, PhysGen, and complex real-world single-image scenes spanning elastic, granular/sand, and rigid objects. PhyMAGIC is benchmarked against leading open-source video diffusion models (OpenSora 2.0, CogVideoX-5B) and state-of-the-art physics-aware 3D generative frameworks (OMNIPHYSGS, PhysDreamer, Physics3D).
| Method | Type | Average CLIP similarity ↑ | Average Aesthetic score ↑ | Image-Motion-FID ↓ | Human: Physical Plausibility ↑ | Human: Text Consistency ↑ |
|---|---|---|---|---|---|---|
| OpenSora 2.0 | I2V Video Diffusion | 0.233 | 16.98 | - | 2.02 | 2.18 |
| CogVideoX-5B | I2V Video Diffusion | 0.239 | 30.66 | - | 2.58 | 2.49 |
| CogVideoX* (w/ our prompt refinement) | I2V Video Diffusion | 0.240 | 31.86 | - | 2.97 | 3.00 |
| OMNIPHYSGS | Physics-Aware 3D | 0.197 | - | 106.60 | - | - |
| PhysDreamer | Physics-Aware 3D | 0.213 | - | 107.47 | - | - |
| Physics3D | Physics-Aware 3D | 0.217 | - | 98.91 | - | - |
| PhyMAGIC (Ours) | Training-Free Active Probing | 0.251 | 30.69 | 94.69 | 3.00 | 3.07 |
Ablation Study¶
The paper validates the progressive convergence of physical parameters across refinement iterations on representative scenarios (swing ficus, rolling basketball, driving car) and ablates the confidence-gated probing mechanism on a multi-scene subset.
| Config / Iteration Stage | Swing ficus Accuracy (%) ↑ | Rolling basketball Accuracy (%) ↑ | Driving car Accuracy (%) ↑ | Mat. acc. (%) ↑ | BC acc. (%) ↑ | Probes/scene ↓ |
|---|---|---|---|---|---|---|
| One-shot VLM + fixed boundary | - | - | - | 60.0 | 60.0 | 0.00 |
| One-shot VLM + best template | - | - | - | 76.0 | 80.0 | 0.00 |
| Iterative probing, no confidence gating | - | - | - | 83.0 | 80.0 | 1.00 |
| Iteration 1 (Initial static prediction) | 62.92 | 50.27 | 29.17 | - | - | - |
| Iteration 2 (First probe video feedback) | 82.29 | 73.75 | 44.50 | - | - | - |
| Iteration 3 / Full (Closed-loop convergence) | 90.63 | 90.00 | 93.75 | 90.0 | 90.0 | 0.80 |
Key Findings¶
- Iterative probing eliminates severe initial misjudgments: In Iteration 1 (static image only), VLM reasoning often misclassifies deceptively styled objects, such as classifying a driving car as elastic (achieving only 29.17% overall accuracy). By Iteration 3, after two rounds of targeted dynamic probe evidence, accuracy jumps to 93.75% with exact convergence to rigid-body material parameters.
- Confidence gating delivers higher accuracy with lower probe cost: Ungated iterative probing executes probes uniformly across all scenes (1.00 probe/scene), whereas confidence gating (\(\gamma = 0.6\)) selectively skips confident attributes, reducing probe consumption by 20% (0.80 probes/scene) while boosting both material and boundary-condition accuracy to 90.0%.
- Spatiotemporal slices demonstrate strict motion localization: Comparative spatiotemporal slices and optical flow visualizations show that pure video diffusion models cause unintended pseudo-motion in static regions (e.g., wobbling flowerpots or morphing ground planes). In contrast, PhyMAGIC enforces rigorous physical conservation laws through the MPM solver, strictly confining motion to dynamic target components.
Highlights & Insights¶
- Repurposing generative models as active perceptual probes: Instead of relying on generative video models to directly deliver the final dynamics, PhyMAGIC ingeniously treats video diffusion as an interactive dynamic experiment generator, unlocking pretrained temporal priors to reveal latent physical cues.
- Confidence-driven closed-loop self-correction: By measuring self-consistency entropy across multiple VLM samples, the system detects epistemic uncertainty without ground truth, automatically reformulating text prompts to gather missing physical evidence.
- Bridging high-level reasoning and low-level continuum mechanics: The hybrid physical specification (HPP) decouples semantic reasoning from numerical computation, preserving the broad commonsense generalization of VLMs while leveraging MPM solvers for strict momentum and collision compliance.
Limitations & Future Work¶
- Coupling with 3D reconstruction quality: Simulation fidelity is fundamentally constrained by initial 3D mesh and Gaussian field quality generated by Trellis. Delicate thin-walled geometries or topological voids can cause non-uniform particle discretization, leading to localized MPM solver instability.
- Semi-quantitative precision of VLM estimates: While VLM reasoning successfully identifies material categories and order-of-magnitude moduli, it remains semi-quantitative and struggles with high-precision parameters such as micro-scale friction coefficients or high-frequency dampings.
- Absence of dense articulated multi-body support: The current simulation pipeline primarily targets single-body elastoplastic and rigid dynamics, lacking automated support for complex kinematic joints (e.g., robotic arms) or multi-body contact cascades. Future extensions could leverage differentiable rendering gradients between probe videos and simulation rollouts to fine-tune physical parameters.
Related Work & Insights¶
- vs PhysGaussian (CVPR 2024): PhysGaussian pioneered coupling continuum mechanics with 3DGS, but all constitutive parameters required manual human tuning; PhyMAGIC achieves fully automated, closed-loop physical inference and dynamic synthesis without human intervention.
- vs PhysGen (ECCV 2024): PhysGen relies on a single-pass static VLM inference that is vulnerable to visual ambiguity and restricted to rigid objects; PhyMAGIC uses active motion probes and iterative refinement to handle elastic, granular, and rigid materials robustly.
- vs Physics3D (arXiv 2024) / PhysDreamer (ECCV 2024): Physics3D and PhysDreamer rely on expensive score distillation sampling (SDS) or gradient-based video optimization that takes hours and risks local minima; PhyMAGIC adopts a training-free active probe plus forward MPM framework that runs efficiently and stably.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneering formulation of video generation as active physical probes coupled with closed-loop VLM confidence gating]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across metrics, ablations, optical flow, spatiotemporal slices, and 61-participant user study]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical formulation, structured exposition, and coherent pipeline presentation]
- Value: ⭐⭐⭐⭐⭐ [Provides a practical, highly generalizable, training-free pathway for physically grounded 3D/4D asset generation]