Skip to content

๐Ÿค– Robotics & Embodied AI

๐Ÿง  NeurIPS2026 ยท 6 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (129) ยท ๐Ÿ“ท CVPR2026 (146) ยท ๐Ÿ”ฌ ICLR2026 (162) ยท ๐Ÿ’ฌ ACL2026 (11) ยท ๐Ÿงช ICML2026 (53) ยท ๐Ÿค– AAAI2026 (30)

๐Ÿ”ฅ Top topics: Navigation ร—2

Action Chunking Proximal Policy Optimization with Feedback Correction

ACPPO-Corr extends PPO with a low-frequency action-chunk planner and a stepwise feedback corrector while retaining a state-only value network, improving final normalized interquartile mean over PPO by 30.4% across 25 simulated robotics tasksโ€”not by 30.4 percentage points of success rate.

Behavioral Foundation Models for Quality Diversity

BFM-QD freezes an offline-pretrained behavioral foundation model, builds a repertoire of high-quality, diverse behaviors by searching latent codes rather than policy weights, and uses closed-form backward inference for small directed mutations, substantially improving sparse navigation and contact-rich manipulation while still requiring environment rollouts and pretraining investment.

GeoWind2Plan: Mission-Time 3D Urban Wind Prediction for Energy-Efficient UAV Planning

GeoWind2Plan converts building geometry and background wind into a mission-corridor 3D mean wind field and jointly optimizes UAV path and speed; on block A at a background speed of 4 m/s, CFD-evaluated energy decreases from \(13.3\pm1.8\) for wind-agnostic planning to \(12.5\pm1.5\) Wh/km, while wind inference for a 20% corridor takes about 3 seconds.

Language-Conditioned World Modeling for Visual Navigation

The paper introduces a language-conditioned visual navigation dataset and compares two alternative approachesโ€”โ€œdiffusion world model + latent-space actorโ€“criticโ€ and โ€œunified autoregressive action/observation predictionโ€: the former has stronger image structural fidelity in seen environments, whereas the latter predicts offline trajectories better in unseen environments, without validating real closed-loop control.

Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

The paper encodes object motion histories as linked linguistic captions, sparse 3D anchors, and visual anchors, improving semantic trajectory retrieval in long egocentric videos while trading expensive one-time construction for sub-second online queries.

ProCompNav: Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries

ProCompNav collects same-category objects in an unknown environment, then recursively identifies distinguishing attributes and asks users yes/no questions, achieving success rates of 23.7%, 28.1%, and 17.0% across the three simulated CoIN-Bench splits while substantially shortening user-simulator responses.