Skip to content

๐Ÿ“น Video Understanding

๐Ÿง  NeurIPS2026 ยท 1 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (110) ยท ๐Ÿ“ท CVPR2026 (187) ยท ๐Ÿ”ฌ ICLR2026 (48) ยท ๐Ÿงช ICML2026 (17) ยท ๐Ÿค– AAAI2026 (27) ยท ๐Ÿง  NeurIPS2025 (39)

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

TT-VidT reconstructs later frames using a first-frame anchor that retains full spatial information and compact learned motion tokens, substantially improving motion-sensitive tasks under a shared recipe of roughly 1.7 million videos and 8 epochs while reducing TT3D encoder compute to 456.1 GF, without establishing universal video-understanding superiority or strict appearanceโ€“motion disentanglement.