๐ง Human Understanding¶
๐๏ธ ECCV2026 ยท 19 paper notes
๐ Same area in other venues: ๐ท CVPR2026 (151) ยท ๐ฌ ICLR2026 (45) ยท ๐งช ICML2026 (5) ยท ๐ค AAAI2026 (20) ยท ๐ง NeurIPS2025 (21) ยท ๐น ICCV2025 (41)
๐ฅ Top topics: Translation ร2 ยท Face & Gaze ร2
- ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
-
ActionPlan generates temporally aligned semantic action plans before progressively synthesizing 3D human motion, combining offline generation and streaming output in one model and reducing streaming FID from MotionStreamer's 11.790 to 5.735 on the HumanML3D-272 test set.
- ANFI: Rethinking Neighbor Feature Interaction in Person Re-ID
-
ANFI learns both affinity and discrepancy interactions between neighboring person images, trains their sample-wise fusion with noisy relation supervision, and achieves 88.1% mAP on CUHK03 in the same-backbone comparison, 1.2 percentage points above its strongest comparator.
- BackTranslation2.0 -- A Linguistically Motivated Metric to Assess Sign Language Production
-
Proposes BackTranslation2.0, a linguistically motivated text-to-sign language translation evaluation metric. Utilizing a two-phase agent framework (10 dedicated tools for evidence extraction + 4 LLM cross-comparison modules), it produces deterministic scores across four dimensions: grammatical correctness, phonological accuracy, motion fluency, and generation fidelity. It achieves a Pearson \(r=1.00\) / Spearman \(\rho=1.00\) correlation with human judgment on known corruptions and synthetic data, significantly outperforming existing backtranslation and motion similarity baselines.
- BiCE-HG: A Bi-Conditional Egocentric Hand Gesture Dataset for Intelligent Reality Systems
-
BiCE-HG crosses three mobility states with two lighting configurations and provides spatial illuminance records alongside hand skeleton data; its dynamic baseline reaches 89.24% validation accuracy, but the evaluation is not participant-independent and does not establish cross-lighting generalization.
- Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
-
Dress-ED constructs the first large-scale benchmark dataset (146k verified quadruplets) that unifies VTON (virtual try-on), VTOFF (virtual try-off), and text instruction-guided garment editing. It proposes a unified multimodal diffusion baseline model, Dress-EM, based on MLLM + dual-path connector + DiT, which thoroughly outperforms general-purpose editing models and domain-specific VTON models on instruction-guided garment editing tasks.
- FaceMoE: Mixture of Experts for Low-Resolution Face Recognition
-
FaceMoE replaces the single FFN in Transformers with multiple sparsely activated MoE experts and a Top-k router, allowing different experts to automatically specialize in distinct semantic regions of the face (high-frequency texture, low-frequency smooth, landmarks). This achieves resolution-aware feature extraction. It comprehensively outperforms the state-of-the-art (SOTA) on three low-resolution benchmarks (BRIAR, IJB-S, TinyFace) while suffering almost no degradation in high-resolution (HR) pre-trained performance.
- FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
-
FlowerDance combines MeanFlow (a flow matching variant that replaces instantaneous velocity prediction with interval-average velocity) with Physical Consistency Constraints (PCC), a BiMamba backbone, and channel-level cross-modal fusion. It generates high-quality 3D dance motions in only 5โ20 sampling steps. While achieving state-of-the-art (SOTA) quality on FineDance and AIST++, its inference speed (2008 FPS) far exceeds the previous best method (MatchDance at 345 FPS), leaving ample computational budget for real-time 3D rendering.
- InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing
-
InterEdit proposes a new task, Text-guided 3D Dyadic Motion Editing (TMME), builds the first large-scale dyadic motion editing dataset InterEdit3D (comprising 5,161 source-target-text triplets), and designs the InterEdit method based on conditional diffusion models. By utilizing semantic-aware plan token alignment to capture high-level editing intentions and interaction-aware frequency token alignment to constrain interaction rhythm using DCT frequency band energy, InterEdit significantly outperforms four baseline methods in editing faithfulness and motion realism.
- Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning
-
This work reformulates gaze target estimation from pixel-level regression to a hierarchical reasoning problem. It first establishes candidate attention objects using object-level semantic representations, constructs a field-of-view (FOV) cone geometric prior using gaze direction to constrain the search space, and finally achieves precise localization via multi-scale residual fusion. This method achieves state-of-the-art (SOTA) performance on GazeFollow, VideoAttentionTarget, ChildPlay, and GOO-Real using only 7.1M parameters.
- Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion
-
Odoriko proposes the first unified, multimodal human motion generation framework. By hierarchically injecting gender and SMPL shape parameters as explicit conditioning signals into the diffusion backbone, the generated motion reflects the biological morphological characteristics of the subjects. Meanwhile, it achieves or surpasses the performance of current task-specific methods across three tasks (text-to-motion, music-to-dance, and video-to-motion estimation) with a minimal parameter footprint.
- PIAvatar: Physically Interactive Avatars via Deformation Gradient Decoupling
-
PIAvatar proposes an MPM-based physically interactive 3D human avatar framework. By explicitly decoupling user-defined kinematic velocity from the deformation gradient update, it eliminates unintended internal stresses generated during motion driving. It also embeds a skeletal structure to achieve real-time tracking of deformed poses through closed-form optimization, supporting bidirectional human-human and human-object physical interactions alongside non-rigid surface deformations simultaneously within a unified MPM simulation framework for the first time.
- Self-supervised Garment Dynamics with Persistent Wrinkles
-
This paper proposes the first self-supervised neural network garment simulator. By leveraging dynamic rest bending energy and physically-inspired curriculum learning, the approach explicitly models the elasto-plastic deformation of fabrics. It generates natural persistent wrinkles for the first time within a self-supervised framework, outperforming existing methods across various garments, body shapes, and motions.
- SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset
-
SICAGE formulates culture-aware co-speech gesture generation as a domain generalization problem: treating each speaker as a domain, it extracts speaker-independent cultural representations from audio-visual and textual modalities using Fishr regularization or adversarial learning. These representations then condition a real-time diffusion generator, ALaDiT, significantly improving gesture realism, diversity, beat synchronization, and cultural consistency on the self-built 106-hour, four-culture TED4C-L dataset.
- SIGNER: Temporally Grounded Sign Language Generation via Time-Resolved Conditioning
-
SIGNER proposes a time-resolved conditioning framework. By constructing a time-resolved Gloss conditioning sequence and injecting it via Local Temporal Fusion (LTF) during diffusion denoising, it explicitly preserves the temporal grounding of downstream sign language segments. This solves the core issues of disordered sign order and inaccurate semantics in existing sign language generation methods, significantly outperforming prior SOTA on CSL-Daily and Phoenix-2014T.
- SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation
-
SIGNET proposes treating multiple sign language skeleton backbones pre-trained on large corpora as frozen core motion-visual experts. By utilizing hand-prior-driven attention aggregation and learnable gating fusion, it achieves cross-language motion-level knowledge transfer. With only \(1.5\text{M}\) trainable parameters, it delivers state-of-the-art (SOTA) or comparable results on four translation benchmarks and the WLASL recognition benchmark.
- SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks
-
SignNet-1M leverages 3DGS novel view synthesis, diffusion-based scene/identity editing, and post-rendering augmentation to expand 7 public sign language corpora into approximately 1 million multilingual sign language video clips (ASL/CSL/DGS). It introduces a unified evaluation protocol (Orig/Zero-shot/Trained) to expose and mitigate the robustness vulnerabilities of existing models under view, background, and identity variations.
- SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
-
SyncCache proposes a training-free feature caching acceleration scheme for DiT-based audio-driven portrait animation. By employing Spatial Asymmetric Detection (SAP) to prioritize recalculations in the facial region, Modal Decoupled Caching (MDC) to skip heavy visual modules while refreshing lightweight audio modules step-by-step, and memory-adaptive offline DP-based cache planning, it achieves 4.12x and 3.75x acceleration on HunyuanVideo-Avatar and Wan-S2V respectively, with near-lossless visual quality and lip-sync performance.
- Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation
-
STREAM proposes BEAM, a bimodal decoupled energy-based attention mechanism, which allows text to control the semantic structure of the dance motion ("what to do") while the music merely decorates its temporal rhythm ("when to do"). It resolves the modality collapse problem in joint text-music generation via mathematically guaranteed hierarchical classifier-free guidance, achieving SOTA on the newly annotated Motorica++ dataset.
- UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
-
UniMotion treats human motion as a continuous modality with equal status to RGB within a shared LLM. By replacing discrete tokens with a Continuous Motion Aligned VAE (CMA-VAE) and a dual-path embedder, and incorporating Dual Posterior KL Alignment (DPA) to inject visual semantic priors along with Latent Reconstruction Alignment (LRA) to resolve the motion path cold-start issue, UniMotion unifies and achieves comprehensive state-of-the-art (SOTA) performance across seven understanding, generation, and editing tasks on Motion-Text-RGB tri-modalities.