๐ฌ Video Generation¶
๐ง NeurIPS2026 ยท 8 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (169) ยท ๐ท CVPR2026 (182) ยท ๐ฌ ICLR2026 (97) ยท ๐ฌ ACL2026 (4) ยท ๐งช ICML2026 (32) ยท ๐ค AAAI2026 (11)
๐ฅ Top topics: Video Generation ร2
- Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
-
SplitMoE divides the sparse experts of a video diffusion Transformer into semantic and generic groups, guides semantic routing with clean VAE features, and improves video quality over same-source baselines with 27B total and 14B activated parameters, rather than forcing every expert to process an equal number of tokens.
- GLARE: Generating Listening Heads with Appropriate Reactions
-
GLARE models listener feedback such as nodding and smiling as typed temporal events, improving reaction consistency on RealTalk and Seamless through prosody-conditioned flow matching and training-time reaction supervision, while its metrics measure single-reference consistency rather than uniquely correct social responses.
- Motion Forcing: Decoupling Ego and Object Motion via Sparse Inputs for Structured Video Generation
-
Motion Forcing first converts sparse object trajectories and camera motion into dynamic depth, then renders RGB video with a shared diffusion backbone, achieving FVD 157.8, FVMD 205.2, and Physics-IQ 33.2 on Waymo while trading some distributional visual similarity for better motion consistency and physical plausibility.
- PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos
-
PhyProbe trains a lightweight scoring head on a frozen video encoder, combining ranking, severity regression, and endpoint anchoring from heterogeneous sources to score individual videos; it achieves the highest pairwise accuracy on three of four physical benchmarks, but remains a proxy for perceived plausibility rather than a test of physical laws.
- PISCO: Precise Video Instance Insertion with Sparse Control
-
PISCO propagates a few instance keyframes into existing footage through variable-density conditioning, pre-encoding frame completion and post-encoding masking, and depth and appearance augmentation; on PISCO-Bench, first-and-last-frame control with its 14B model reduces whole-video FVD from VACE's 371 to 204, although the compared methods receive different inputs.
- Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models
-
DyMoS subtracts an adjustable bias from non-reference-query-to-reference-key self-attention logits in the text-conditioned branch during early image-to-video denoising, raising Wan 2.2's VBench Dynamic Degree from 51.7 to 64.8, although stronger dynamics do not automatically imply more realistic motion or higher reference fidelity.
- Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
-
ReMind trains video generators to recover states across interruptions using event frame graphs, reliable-anchor training, and KV caches that retain original spatiotemporal addresses, reaching the highest STEVO-Bench Total of 300.9, but only 14.5% State Progress; the evidence primarily supports short-horizon recovery rather than long-term unobserved physical evolution.
- ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
-
ViTeX-Bench uses paired training data from real videos, a frozen evaluation split, and 13 metrics across three axes to distinguish text correctness, motion stability, and background preservation, while providing a motion-aligned glyph-conditioned reference editor, ViTeX-Edit-14B, with the highest mean character accuracy among the evaluated video-native editors rather than the best performance on every metric.