PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos¶
Conference: NeurIPS2026
arXiv: 2609.38377
Area: Video Generation
Keywords: physical consistency, video evaluation, spatiotemporal representations, pairwise ranking, score calibration
TL;DR¶
PhyProbe trains a lightweight scoring head on a frozen video encoder, combining ranking, severity regression, and endpoint anchoring from heterogeneous sources to score individual videos; it achieves the highest pairwise accuracy on three of four physical benchmarks, but remains a proxy for perceived plausibility rather than a test of physical laws.
Background & Motivation¶
Generated videos can preserve sharp textures and coherent motion while showing interpenetrating objects, implausible fluid behavior, or abrupt changes in object form. Attractive appearance does not imply plausible dynamics, and a binary plausible/implausible judgment cannot distinguish minor deviations from severe failures. An evaluator for video generation needs to establish both which video is more severely violated and whether the same numerical value has a comparable meaning across datasets.
Existing vision-language models (VLMs) typically assess physical plausibility through prompts and language descriptions, potentially overlooking subtle motion changes; fine-tuning on human annotations can also encourage dataset-specific semantic shortcuts. Directly fitting absolute human ratings presents another problem: annotators primarily detect salient abnormalities, rating scales shift as generation quality improves, and physical violations frequently co-occur with blur, flicker, and other visual defects. Relative ordering may be reliable without fixing an absolute scale, whereas absolute labels provide magnitude without constituting precise measurements.
Rather than asking an evaluator to explain every physical law, the paper formulates continuous measurement over video representations: the pretrained encoder supplies motion and visual structure, while the scoring head learns ordering, magnitude, and scale endpoints from different sources. Core Idea: integrate pairwise ranking, noisy regression, and anchoring through a lightweight probe on frozen representations, making single-video violation scores comparable in both order and approximate scale.
Method¶
Overall Architecture¶
The input is a real or generated video, and the output is a physical violation score, with lower values indicating greater plausibility. Inference requires neither a paired reference video, a shared prompt, nor a textual explanation: fixed-length sampling and a frozen encoder are followed by statistics pooling, a three-layer MLP produces a raw value, and the value is clipped to \([0,1]\).
The central difference between training and inference is that video pairs and multisource labels constrain this single-video scoring function during training. The system consists of a Frozen Statistical Probe, Heterogeneous Supervision Construction, and Joint Scale Calibration; pairs support training comparisons and evaluation metrics rather than being required inputs to inference.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Video / training video pair"] --> B["Frozen Statistical<br/>Probe"]
T["Seven training sources"] --> C["Heterogeneous<br/>Supervision Construction"]
B -->|Raw training scores| D["Joint Scale<br/>Calibration"]
C -->|Ranking / magnitude / anchors| D
D -.->|Update scoring head only| B
B -->|Clip at inference| O["Single-video violation score"]
Key Designs¶
1. Frozen Statistical Probe: separate representation learning from scale learning
The default backbone is PE-Core-G14-448, with approximately 2B parameters, kept entirely frozen during training. Videos are sampled at 16 fps into 48-frame clips, resized with their aspect ratios preserved, padded to a square with mid-gray bars of pixel value 0.5, and normalized according to the backbone. The scoring head therefore receives a standardized sampling format rather than additional variation induced by different dimensions and cropping policies.
Patch tokens are first spatially averaged within each frame to obtain frame-level features. Temporal mean, standard deviation, and maximum are then concatenated into a descriptor with three times the feature dimension. The mean retains overall content, the standard deviation captures variation across frames, and the maximum preserves salient activations. This maximum is a feature-wise statistic, not the maximum of per-frame physical-error probabilities.
The descriptor feeds a three-layer MLP with GELU activations; the final statistical scoring head has 1.02M trainable parameters. The pretrained backbone carries the burden of representing spatiotemporal information, rather than retraining a large VLM or adding an attention-heavy head. This design facilitates comparisons between representations and makes heterogeneous supervision primarily adjust how violation severity is read out, rather than simultaneously reshaping the entire visual space.
Statistics pooling itself does not explicitly preserve frame order. It can exploit dynamic information already encoded in the backbone features, but mean, standard deviation, and maximum cannot reconstruct a complete causal process; length-insensitive aggregation also does not establish validated evaluation of arbitrary-duration videos. Dilution of brief events and imperfect long-horizon causal encoding are important boundaries of this lightweight choice.
2. Heterogeneous Supervision Construction: prevent correspondence from becoming the only shortcut
The training corpus contains approximately 371K video pairs from PAI-Bench, VideoPhy2, VideoFeedback2, BrokenVideos, ImpossibleVideos, TRAVL, and Kinetics-700. It includes realโgenerated pairs sharing image/text conditioning, generatedโgenerated pairs without semantic correspondence, and realโreal pairs within the same action category. Comparisons restricted to a shared prompt might encourage local content shortcuts; cross-prompt comparisons require the scores to express violation severity across different scenes.
The supervision from every source is not uniformly treated as precise ground truth. PAI-Bench ranks real reference videos above their corresponding generated videos and uses real references as low-violation anchors. VideoPhy2 pairs videos within action categories according to physical correctness ratings; VideoFeedback2 pairs videos across prompts and retains rating differences of at least 1. Two real Kinetics-700 videos constitute a ranking tie and anchor the low-violation endpoint; a tie is not discarded but encourages similar scores.
ImpossibleVideos supplies cross-content comparisons between real videos and counterfactual failure cases. TRAVL instead aggregates each video's physical question-answer labels by majority vote, discards tied votes, and pairs plausible and implausible videos. Both provide ordering and plausible/violated endpoints, but these remain dataset-level operational labels rather than a general rule that every generated video should receive a severe violation score.
For physical correctness ratings from 1 to 5, the authors use violation-flooring rather than a simple linear inversion: ratings 5, 4, 3, 2, and 1 map to 0, 0.5, 0.667, 0.833, and 1, respectively. The intention is to reserve space for weakly perceptible violations while compressing visible violations into the higher range. This is an empirical design choice, not a conversion into physical units; in particular, \([0,0.5)\) lacks direct severity annotations, and its fine-grained interpretation has not been independently validated.
BrokenVideos has no direct pairwise preference labels. The authors construct a heuristic anomaly score from mask area, center proximity, and the fraction of frames containing a mask, weighted by 0.6, 0.3, and 0.1, respectively. Mask presence ratio measures temporal persistence, while center proximity increases the contribution of defects near the image center. Pairs require a raw heuristic gap of at least 0.4, and regression magnitudes are subsequently mapped to \([0.5,0.9]\). These labels also cover visual artifacts and are therefore not pure measurements of dynamical errors.
A larger source is not automatically more reliable. Dataset-weighted sampling prevents the 140K Kinetics-700 pairs from overwhelming smaller human-rating sources, and pair pools are resampled every 2.5K steps. An appendix study validates 85 BrokenVideos pairs with four annotators: the heuristic agrees with the human majority on 80.6% of non-tied pairs (58/72). This supports auxiliary supervision without eliminating label noise.
3. Joint Scale Calibration: constrain ordering, magnitude, and endpoints separately
Pairwise ranking uses the BradleyโTerry formulation: the sigmoid of the difference between two raw scores represents the probability that the first video is more severely violated. Labels 1, 0, and 0.5 indicate a more violated first video, a more violated second video, and a realโreal tie, respectively; binary cross-entropy teaches relative ordering. Because ranking depends only on differences, it cannot independently establish a score offset and absolute magnitude, so high ranking accuracy does not guarantee a usable scale.
The magnitude branch fits normalized noisy severity labels with Huber loss at \(\delta=0.5\). Large errors receive a linear penalty, reducing sensitivity to a small number of noisy labels. The anchoring branch uses mean squared error to fit 0 for high-confidence plausible examples and 1 for clear failures. The former supplies intermediate magnitude information, while the latter fixes the endpoints near interpretable references; their roles cannot be collapsed into a single generic regression label.
Writing the unclipped scoring-head output as \(\tilde{s}\), the training objective is:
Here \(y_v\) is the severity regression target, \(I(v)\) is an anchor label of 0 or 1, and the ranking term is cross-entropy over the score-difference probability described above. The equation preserves the paper's three supervision components while explicitly distinguishing raw outputs; some main-text training equations use \(s\) throughout, whereas the appendix specifies that the head is unconstrained during training.
Only at inference is the reported score obtained:
Thus bounded reported scores result from clipping, not a terminal sigmoid or a loss that strictly guarantees in-range outputs. Clipping can also turn extreme raw values into ties, so the raw ranges in Table 4 must be distinguished from the final reporting interval. The paper explicitly notes that Spearman and Pearson correlations assess ordering and linear association, respectively; neither establishes absolute calibration.
Loss & Training¶
Only the scoring head is trained, using AdamW with learning rate \(3\times10^{-4}\) and weight decay 0.01. Training warms up for 500 steps, then uses cosine decay to zero over 10K total steps, with batch size 8. The authors report training within 24 hours on four NVIDIA H100 GPUs and evaluate EMA weights with decay 0.9999; this does not imply that the system avoids the computation of a large backbone.
Ranking labels also receive smoothing according to annotation gaps within the current batch slice for the same dataset: the largest-gap pair uses 0.05, near-tied pairs receive up to 0.095, and pairs without scalar annotations use 0.05. Ambiguous comparisons therefore receive less sharp targets than clear differences. Source sampling probabilities are PAI-Bench 0.059, VideoPhy2 0.176, VideoFeedback2 0.294, BrokenVideos 0.118, ImpossibleVideos 0.059, TRAVL 0.176, and Kinetics-700 0.118.
Key Experimental Results¶
Main Results¶
The following selection comes from Table 2. Accuracy is in %, and correlation columns report Spearman \(\rho\); R denotes real videos and G generated videos. ImplBp contains 150 pairs; VP2p contains 1280 pairs sharing text prompts; PhyDExp and VF2p each contain 2000 pairs without semantic correspondence. Average accuracy is weighted by benchmark pair counts, not a simple mean of the four columns.
| Method | ImplBp (RโG) | VP2p (GโG) | PhyDExp (RโG) | VF2p (GโG) | Weighted average | VP2 \(\rho\) | VF2 \(\rho\) |
|---|---|---|---|---|---|---|---|
| VP2-AutoEval | 28.0 | 28.5 | 30.5 | 36.3 | 32.1 | 0.359 | 0.235 |
| V-JEPA Surprise | 33.3 | 46.3 | 49.0 | 45.5 | 46.6 | 0.065 | 0.121 |
| VideoScore2 | 68.7 | 49.7 | 60.6 | 65.0 | 59.9 | 0.197 | 0.462 |
| Gemini-3.1-Flash-Lite | 91.3 | 61.8 | 89.0 | 74.5 | 77.3 | 0.302 | 0.419 |
| Gemini-3.1-Pro | 95.2 | 70.5 | 86.2 | 73.7 | 78.1 | 0.277 | 0.402 |
| PhyProbe | 98.0 | 66.1 | 98.9 | 75.2 | 82.4 | 0.293 | 0.596 |
PhyProbe's Pearson \(r\) is 0.604 on VF2 and 0.302 on VP2. Human plausibility and violation scores must have their directions aligned, for example using \(1-s(v)\); a high violation score must not be interpreted as high plausibility. PhyProbe does not lead every metric: Gemini-3.1-Pro reaches 70.5% on VP2p, exceeding 66.1%, and VP2-AutoEval also has higher VP2 correlation.
General-preference results come from Table 3. On MonetBench with ties, PhyProbe reaches 72.7%, versus 68.2% for VisionReward and 56.1% for VideoReward; without that tie protocol, the values are 71.0%, 72.2%, and 58.2%, respectively. PhyProbe achieves 62.7% on Rapidata-I2V and 65.3% on RewardBench Overall without ties, below VideoReward's 73.6%. These results support transfer rather than universal superiority on general preference.
Ablation Study¶
The following selection comes from Table 4 and uses the same PE-Core backbone. Score ranges are effective ranges before clipping; negative zero is preserved from the source.
| Config | PhyDExp | VF2p | VP2p | ImplBp | VF2 \(\rho\) | Raw output range |
|---|---|---|---|---|---|---|
| Without ranking loss | 94.7 | 75.6 | 61.9 | 99.2 | 0.614 | \([-0.07,0.88]\) |
| Without magnitude loss | 98.7 | 75.3 | 62.3 | 94.1 | 0.584 | \([-1.89,3.45]\) |
| Without anchor loss | 68.7 | 79.3 | 64.4 | 84.9 | 0.664 | \([-0.40,4.13]\) |
| Full model | 98.9 | 75.2 | 66.1 | 98.0 | 0.596 | \([-0.00,1.15]\) |
The next selection comes from Table 6, comparing aggregation heads on the same backbone; parameter counts include only trainable components.
| Scoring head | Trainable parameters | VP2p accuracy (%) | VP2 \(\rho\) | VF2p accuracy (%) | VF2 \(\rho\) |
|---|---|---|---|---|---|
| Mean-pooling MLP | 0.36M | 61.0 | 0.277 | 73.2 | 0.578 |
| Attention pooling | 6.92M | 63.7 | 0.292 | 75.0 | 0.612 |
| Statistics pooling (mean + std + max) | 1.02M | 66.1 | 0.293 | 75.2 | 0.596 |
Key Findings¶
- Anchoring strongly constrains realโgenerated comparisons: removing it reduces PhyDExp from 98.9% to 68.7%, while VF2p rises to 79.3%. Higher local ranking or correlation cannot substitute for a stable scale and cross-domain performance.
- Statistics pooling achieves higher pairwise accuracy with far fewer parameters than attention pooling, but its VF2 correlation of 0.596 is below attention's 0.612. It is a performanceโcomplexity choice, not dominance on every metric.
- With the identical V-JEPA2 ViT-H/16 backbone in Table 7, PhyProbe improves over Surprise by 48.1, 13.9, 21.7, and 30.8 percentage points across the four benchmarks. A stronger backbone alone does not explain the gains.
- In Table 9, removing VideoFeedback2 reduces VF2p from 75.2% to 58.0%; removing VideoPhy2 reduces VP2p from 66.1% to 58.3%. Heterogeneous supervision provides partial transfer alongside substantial dataset dependence.
- Linear normalization in Table 10 yields 96.7%, 60.0%, 96.5%, and 71.0% across the four benchmarks, below violation-flooring throughout. The auxiliary human study in Table 11 contains only 30 pairs and four non-experts, with Fleiss \(\kappa=0.148\); it does not establish strong validation of the scale.
Highlights & Insights¶
- Separating ordering, magnitude, and endpoints is a transferable evaluator design. Heterogeneous labels need not all supply both reliable preferences and precise absolute values.
- A probe on frozen features allows representation quality and the readout objective to be examined separately. The same-backbone comparison is important evidence beyond comparing a 2B evaluator with a prediction-error baseline.
- Prompt-free, single-video inference lowers deployment friction. Cross-content comparability still rests on supervision and calibration assumptions; a prompt-free interface alone does not demonstrate elimination of semantic shortcuts.
Limitations & Future Work¶
- Long-horizon causal violations, instantaneous state changes, and unusual but physically valid real motions can be misjudged; fixed clips and statistics aggregation do not guarantee coverage. Multiscale sampling and event-order-preserving representations are possible directions.
- Supervision entangles perceived physics with visual quality, and BrokenVideos specifically includes nonphysical artifacts. Demonstrating isolated physical measurement requires controlling scene, prompt, generator, and visual quality while changing only physical factors; the paper lacks such a controlled benchmark.
- Appendix C states that VF2 test data share all prompts and generators with training. Supervision instances are disjoint, but this is not prompt- or generator-level distribution shift. VP2 has no prompt overlap and partial generator overlap, so their generalization implications must not be conflated.
- \([0,0.5)\) lacks direct severity labels, and anchors are approximate endpoints; even the full model's raw outputs reach 1.15. A stable reference scale is an empirically supported design goal, not an absolute measurement in physical units.
- Table 2 reports VideoScore2 VP2p accuracy as 49.7%, whereas Table 8 reports 49.8% at minimum rating gap 1 and labels the column VideoScore. This numerical/protocol difference is preserved rather than silently reconciled. Appendix B.1.1 also uses โraw scoresโ wording; the explicit normalization mapping in B.1 and training description in B.2 govern the mechanism described here.
- Physical plausibility does not establish that depicted events really occurred. The evaluator is not a tool for determining authenticity or factual correctness.
Related Work & Insights¶
- vs VideoPhy2-AutoEval: The latter uses a fine-tuned VLM to output discrete physical-commonsense scores, whereas PhyProbe learns a continuous readout. Appendix Table 19 reports a 58.8% model tie rate on VP2p, 28.5% accuracy over all pairs, and 69.1% on model-non-tied pairs. Overall accuracy comparisons should explain ties induced by discrete scores rather than simply asserting a complete lack of physical understanding.
- vs VideoScore2: The latter evaluates multiple quality dimensions through language reasoning; PhyProbe provides no textual explanation and integrates multisource physical supervision. VideoFeedback2 supplied by the former is also a training source here, so corresponding test results must be interpreted alongside distribution overlap and leave-one-dataset-out experiments.
- vs V-JEPA Surprise: Prediction error is not inherently physical violation severity. A supervised scoring head substantially improves results on identical representations, suggesting that pretrained world representations should be treated as a basis for readout rather than their prediction error being assumed to be the final metric.
Rating¶
- Novelty: 4/5 โ The frozen probe is simple; its division of measurement roles under heterogeneous supervision is more valuable.
- Experimental Thoroughness: 4/5 โ Pair-type, loss, backbone, aggregation, and leave-one-source-out experiments are included, but controlled isolation of physical factors is missing.
- Writing Quality: 4/5 โ The method and calibration boundaries are clear, with minor discrepancies in table values and label wording.
- Value: 4/5 โ Useful for lightweight continuous evaluation of video generation, not a substitute for physical verification or authenticity assessment.