Skip to content

๐ŸŽฌ Video Generation

๐Ÿง  NeurIPS2026 ยท 8 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (169) ยท ๐Ÿ“ท CVPR2026 (182) ยท ๐Ÿ”ฌ ICLR2026 (97) ยท ๐Ÿ’ฌ ACL2026 (4) ยท ๐Ÿงช ICML2026 (32) ยท ๐Ÿค– AAAI2026 (11)

๐Ÿ”ฅ Top topics: Video Generation ร—2

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

SplitMoE divides the sparse experts of a video diffusion Transformer into semantic and generic groups, guides semantic routing with clean VAE features, and improves video quality over same-source baselines with 27B total and 14B activated parameters, rather than forcing every expert to process an equal number of tokens.

GLARE: Generating Listening Heads with Appropriate Reactions

GLARE models listener feedback such as nodding and smiling as typed temporal events, improving reaction consistency on RealTalk and Seamless through prosody-conditioned flow matching and training-time reaction supervision, while its metrics measure single-reference consistency rather than uniquely correct social responses.

Motion Forcing: Decoupling Ego and Object Motion via Sparse Inputs for Structured Video Generation

Motion Forcing first converts sparse object trajectories and camera motion into dynamic depth, then renders RGB video with a shared diffusion backbone, achieving FVD 157.8, FVMD 205.2, and Physics-IQ 33.2 on Waymo while trading some distributional visual similarity for better motion consistency and physical plausibility.

PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos

PhyProbe trains a lightweight scoring head on a frozen video encoder, combining ranking, severity regression, and endpoint anchoring from heterogeneous sources to score individual videos; it achieves the highest pairwise accuracy on three of four physical benchmarks, but remains a proxy for perceived plausibility rather than a test of physical laws.

PISCO: Precise Video Instance Insertion with Sparse Control

PISCO propagates a few instance keyframes into existing footage through variable-density conditioning, pre-encoding frame completion and post-encoding masking, and depth and appearance augmentation; on PISCO-Bench, first-and-last-frame control with its 14B model reduces whole-video FVD from VACE's 371 to 204, although the compared methods receive different inputs.

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models

DyMoS subtracts an adjustable bias from non-reference-query-to-reference-key self-attention logits in the text-conditioned branch during early image-to-video denoising, raising Wan 2.2's VBench Dynamic Degree from 51.7 to 64.8, although stronger dynamics do not automatically imply more realistic motion or higher reference fidelity.

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

ReMind trains video generators to recover states across interruptions using event frame graphs, reliable-anchor training, and KV caches that retain original spatiotemporal addresses, reaching the highest STEVO-Bench Total of 300.9, but only 14.5% State Progress; the evidence primarily supports short-horizon recovery rather than long-term unobserved physical evolution.

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

ViTeX-Bench uses paired training data from real videos, a frozen evaluation split, and 13 metrics across three axes to distinguish text correctness, motion stability, and background preservation, while providing a motion-aligned glyph-conditioned reference editor, ViTeX-Edit-14B, with the highest mean character accuracy among the evaluated video-native editors rather than the best performance on every metric.