๐ง Human Understanding¶
๐ง NeurIPS2026 ยท 7 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (119) ยท ๐ท CVPR2026 (151) ยท ๐ฌ ICLR2026 (45) ยท ๐งช ICML2026 (5) ยท ๐ค AAAI2026 (20) ยท ๐ง NeurIPS2025 (21)
๐ฅ Top topics: Human Pose ร2 ยท Face & Gaze ร2
- BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion
-
BiMoGen combines a unified motion-text vocabulary with a bidirectional masked diffusion model for motion generation and captioning, using correspondence learning before conditional generation and generation-aware self-correction to achieve T2M R@1 of 0.555 and M2T CIDEr of 60.2 on HumanML3D, without outperforming existing unified models on every metric.
- FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery
-
FactorizedHMR first deterministically estimates the torso, shape, and camera-space trajectory, then fixes that structure while masked flow matching completes limbs, the head, and world motion; its clearest benefits concern severe occlusion and trajectory drift rather than uniformly superior conventional metrics.
- GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction
-
GazeFlow generates whole gaze trajectories through flow matching conditioned on local video features and global task tokens, reaching F1 0.491 and average angular error 9.01 on EGTEA Gaze+ while improving displacement statistics, but its default model uses future frames and is not directly an online eye tracker.
- PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
-
Instead of adding an RGB action-recognition branch, PoseBridge preserves visual semantics inside pose estimation before they are compressed into joint coordinates, then transfers them through skeleton-conditioned bridging and semantic prototype adaptation, improving over the strongest compared baseline by 13.3โ17.4 percentage points across eight Kinetics-200/400 splits.
- Re:Cognize: Open-Set Comic Character Re-Identification
-
Re:Cognize separates identity emergence from identity maintenance through four gallery protocols over the same reading-order stream, finds that growth under predicted labels harms recognition with random seeds, and uses Re:Cast's character averages, page evidence, and cross-page binding to recover part of the benefit of correctly labelled updates under explicit conditions.
- STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts
-
STRIDE precompiles textual crowd contexts into behavioral questions, deterministic measurement functions, and expected ranges, then checks generated trajectories against them; its benchmark contains 936 scenarios, Text-Crowd achieves an overall score of 0.645, and humans agree with benchmark answers on 80% of pairs, although the latter result comes from a limited annotation subset.
- Towards Unified Dynamic Face Landmark Detection
-
The paper describes landmarks using a face part and a normalized sequence position, then uses image-conditioned queries and iterative decoding to train one model across annotation formats and predict landmarks on demand; the unified ViT-B achieves full-set NME of 4.05, 2.80, and 1.02 on WFLW, 300W, and AFLW-19, respectively, offering unified training and interfaces rather than superiority on every metric.