Skip to content

๐Ÿง‘ Human Understanding

๐Ÿง  NeurIPS2026 ยท 7 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (119) ยท ๐Ÿ“ท CVPR2026 (151) ยท ๐Ÿ”ฌ ICLR2026 (45) ยท ๐Ÿงช ICML2026 (5) ยท ๐Ÿค– AAAI2026 (20) ยท ๐Ÿง  NeurIPS2025 (21)

๐Ÿ”ฅ Top topics: Human Pose ร—2 ยท Face & Gaze ร—2

BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion

BiMoGen combines a unified motion-text vocabulary with a bidirectional masked diffusion model for motion generation and captioning, using correspondence learning before conditional generation and generation-aware self-correction to achieve T2M R@1 of 0.555 and M2T CIDEr of 60.2 on HumanML3D, without outperforming existing unified models on every metric.

FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery

FactorizedHMR first deterministically estimates the torso, shape, and camera-space trajectory, then fixes that structure while masked flow matching completes limbs, the head, and world motion; its clearest benefits concern severe occlusion and trajectory drift rather than uniformly superior conventional metrics.

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

GazeFlow generates whole gaze trajectories through flow matching conditioned on local video features and global task tokens, reaching F1 0.491 and average angular error 9.01 on EGTEA Gaze+ while improving displacement statistics, but its default model uses future frames and is not directly an online eye tracker.

PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition

Instead of adding an RGB action-recognition branch, PoseBridge preserves visual semantics inside pose estimation before they are compressed into joint coordinates, then transfers them through skeleton-conditioned bridging and semantic prototype adaptation, improving over the strongest compared baseline by 13.3โ€“17.4 percentage points across eight Kinetics-200/400 splits.

Re:Cognize: Open-Set Comic Character Re-Identification

Re:Cognize separates identity emergence from identity maintenance through four gallery protocols over the same reading-order stream, finds that growth under predicted labels harms recognition with random seeds, and uses Re:Cast's character averages, page evidence, and cross-page binding to recover part of the benefit of correctly labelled updates under explicit conditions.

STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts

STRIDE precompiles textual crowd contexts into behavioral questions, deterministic measurement functions, and expected ranges, then checks generated trajectories against them; its benchmark contains 936 scenarios, Text-Crowd achieves an overall score of 0.645, and humans agree with benchmark answers on 80% of pairs, although the latter result comes from a limited annotation subset.

Towards Unified Dynamic Face Landmark Detection

The paper describes landmarks using a face part and a normalized sequence position, then uses image-conditioned queries and iterative decoding to train one model across annotation formats and predict landmarks on demand; the unified ViT-B achieves full-set NME of 4.05, 2.80, and 1.02 on WFLW, 300W, and AFLW-19, respectively, offering unified training and interfaces rather than superiority on every metric.