Skip to content

πŸ“Ή Video Understanding

🎞️ ECCV2026 · 20 paper notes

πŸ“Œ Same area in other venues: πŸ“· CVPR2026 (187) Β· πŸ”¬ ICLR2026 (48) Β· πŸ§ͺ ICML2026 (17) Β· πŸ€– AAAI2026 (27) Β· 🧠 NeurIPS2025 (39) Β· πŸ“Ή ICCV2025 (56)

πŸ”₯ Top topics: Anomaly Detection Γ—2 Β· Human Pose Γ—2

A Dual-Transformer Architecture with Cross-Attention for Multi-Camera View Recommendation

The paper separates the encoding of previously shown footage from the evaluation of current camera candidates, letting candidates query temporal memory through cross-attention; with frozen SwinV2-Tiny features and focal loss, it reaches 69.65% [email protected] on TVMCE, a thresholded recommendation metric rather than forced six-way selection accuracy.

Adaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation

The method compresses action-feature sequences into sparse DDIM latent anchors and allocates anchors by reconstruction error, reaching 68.0% ASFormer accuracy on Breakfast versus 62.8% for GNI at the same temporal anchor budget, while remaining below the 70.7% full-data result.

Aligning Anything: Hierarchical Motion Estimation for Video Frame Interpolation

The paper injects SAM object regions into contextual feature extraction, flow distillation, and feature supervision, retaining pixel-level motion flexibility while improving object-level coherence; under matched retraining settings, AMT-G improves from 36.42 to 36.62 dB on Vimeo90K.

Amplify, Aggregate, and Adjust: VideoMAE-based Holistic-Subtle Aggregation for Micro-Action Recognition

A3-MAE extends an RGB-only VideoMAE by locating and amplifying subtle-motion regions, exchanging holistic and local information bidirectionally, and calibrating distillation supervision with instance-wise confidence, improving MA-52 Action Top-1 from 66.47% to 67.92% and iMiGUE from 66.44% to 67.08%.

Bayesian Uncertainty Attribution-Guided Fine-Tuning for Open-Set Action Recognition

The paper turns ensemble uncertainty from a test-time rejection score into a representation-refinement signal, using three-stage fine-tuning and two-phase distillation to reach 88.50 AUROC with a single TPN student on UCF101/HMDB51, compared with 88.62 for the refined ensemble.

Bounding-Box Trajectories Matter for Video Anomaly Detection

TrajVAD promotes bounding-box trajectories already produced by detection and tracking to the primary anomaly signal, learning normal motion with a class-aware normalizing flow; its trajectory-only variant reaches 87.7 AP on ShanghaiTech, while reliability-gated pose fusion reaches 88.6 AUROC and 90.9 AP, although pose does not help on every dataset.

Cambrian-P: Pose-Grounded Video Understanding

Cambrian-P adds per-frame camera pose regression to video MLLM training and reconciles geometry learning with question answering through interleaved training and frame sampling jitter, improving VSI-Bench from 69.2 to 73.7 against a matched no-pose baseline; the gains primarily reflect better learned representations rather than inference-time pose-token use.

ChronusOmni: Improving Time Awareness of Omni-Modal Large Language Models

ChronusOmni interleaves absolute-time text, video frames, and corresponding audio for Ola-7B, then learns six temporal tasks through audiovisual dense-captioning SFT and task-reward GRPO, reaching 79.85 audio-to-time [email protected] on ChronusAV and 34.5 zero-shot temporal-grounding mIoU on LongVALE.

ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement

ClearText-Video (CTVid) connects low-level restoration with text-grounded video QA through high-quality, degraded, and restored versions of the same content, using 4,639 videos to show that sharper-looking inputs need not be read more accurately and that in-dataset fine-tuning can improve accuracy while reducing an uncertainty-aware metric.

CLUE-VAD: Structured Semantic Clues for Understanding Explainable Events in Video Anomaly Detection

CLUE-VAD organizes video evidence into Action, Environment, and Object captions, performs weakly supervised detection through category-aware fusion, and generates explanations using segment attention and a clue-conditioned language model, achieving 89.47% AUC / 88.24% AP on UCF-Crime / XD-Violence without establishing superiority over every multimodal detector.

ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control

ConTrack treats object motion tracking as the priority task constraint and learns robotic hand control with online weight adaptation, a reachable-state reset library, and contact priors, achieving mean progress of 0.899 and contact F1 of 0.784 within 5000 PPO updates per clip, without leading every error metric.

CTEPM: Continuous-Time Event Process Memory for Long-Video Language Models

CTEPM represents long videos as continuous-time event processes with semantic marks and conditional intensities, computes temporal evidence before generating an answer, and improves LongVideoBench accuracy from 46.0% to 49.2% under the same 64-frame budget.

DART: Deformable Adaptive Reasoning with Temporal Queries for Online Skeleton-Based Action Recognition

DART uses a Deformable Multi-scale Temporal Network to locate informative historical skeleton features and learnable queries to distill action evolution cues from compact memory, achieving per-frame recognition without explicit future forecasting and reaching 73.99% AUC on NTU60 X-Sub, 3.04 percentage points above InfoGCN++.

DisentangledTMR: Privacy-Preserving Skeleton Motion Retargeting via Factorized Transformers

DisentangledTMR separately encodes a source person's action and a reference person's skeletal identity, then recombines them through a factorized Transformer; at its recommended partial-retargeting setting on NTU60, pre-trained re-identification falls from 75.4% to 18.1%, with 75.8% pre-trained action recognition and 87.1% action recognition after retraining.

DnA: Denoising Attention for Visual Tasks

DnA uses softmax/softmin queries over shared keys to retrieve complementary interactions, then combines separate value projections with a learned per-head coefficient, raising ViT-B ImageNet-1K Top-1 from 81.1% to 81.9% and improving video classification and visual-probe question answering at additional parameter and computational cost.

Dynamic Inverse Rendering for Enhanced Material-Lighting Decomposition

DIR turns rigid motion during hand-held capture into illumination constraints, using three stages of tracking, geometry refinement, and physical inverse rendering to recover relightable assets; on synthetic diffuse objects, hand-held capture with estimated poses achieves 40.33 dB albedo PSNR versus 32.68 dB for static multiview capture with ground-truth poses.

EatVid-Bench: A Multimodal Fine-Grained Eating Behavior Video Dataset

EatVid-Bench turns 682 real-world eating videos into traceable hierarchical behavioral annotations and 8,184 QA pairs, shows that strong health-related descriptions can coexist with weak objective perception, and uses AGCoT-LoRA to improve Qwen2.5-VL-7B's objective weighted score from 0.1623 to 0.3214 on the 64-frame Hard subset.

EgoEverything: A Benchmark for Human Behavior–Inspired Long-Context Egocentric Video Understanding in AR Environment

EgoEverything samples question targets using real eye-tracking traces and combines delayed questioning, dual-agent evidence gathering, and human filtering to build an AR memory benchmark with over 100 hours of video and over 5,000 multiple-choice questions, where the best evaluated model, Gemini 1.5 Pro, achieves 63.1% accuracy, trailing human accuracy of 83.5% by 20.4 percentage points.

EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage

EgoPolice builds a benchmark of real body-worn camera footage using objective action definitions and per-second human annotation, with supervised linear probing and zero-shot five-way questions exposing fine-grained recognition and domain-transfer weaknesses; even Gemini 2.5 Pro achieves only 76.9% accuracy on one-minute questions, which does not justify autonomous deployment.

Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding

MERIT stores ultra-long videos as short episodic records with complementary textual keys, then retrieves and filters neighboring context only after receiving a question, achieving 71.2% accuracy on EgoLifeQAβ€”5.6 percentage points above WorldMM with the same GPT-5 solverβ€”using a simple memory structure.