๐น Video Understanding¶
๐ง NeurIPS2026 ยท 1 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (110) ยท ๐ท CVPR2026 (187) ยท ๐ฌ ICLR2026 (48) ยท ๐งช ICML2026 (17) ยท ๐ค AAAI2026 (27) ยท ๐ง NeurIPS2025 (39)
- TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
-
TT-VidT reconstructs later frames using a first-frame anchor that retains full spatial information and compact learned motion tokens, substantially improving motion-sensitive tasks under a shared recipe of roughly 1.7 million videos and 8 epochs while reducing TT3D encoder compute to 456.1 GF, without establishing universal video-understanding superiority or strict appearanceโmotion disentanglement.