Skip to content

๐ŸงŠ 3D Vision

๐Ÿง  NeurIPS2026 ยท 8 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (512) ยท ๐Ÿ“ท CVPR2026 (751) ยท ๐Ÿ”ฌ ICLR2026 (195) ยท ๐Ÿงช ICML2026 (30) ยท ๐Ÿค– AAAI2026 (79) ยท ๐Ÿง  NeurIPS2025 (116)

๐Ÿ”ฅ Top topics: Face & Gaze ร—3

DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding

DirectUV generates mesh textures directly in the latent UV space of a frozen Flux VAE, embeds 3D surface coordinates into each attention head's rotary positional encoding, and uses the reference image and coarse UV only as conditions, achieving 25.13 PSNR and 0.9527 SSIM on GSO while improving consistency across UV islands and texture detail.

M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals

M-plicits reconstructs oriented point clouds with a full-domain coarse SIREN and progressively band-supervised residual SIRENs, reusing this hierarchy for tracing, mesh extraction, and attribute mapping to deliver real-time rendering and strong synthetic measurement-noise robustness with compact models, without winning every accuracy or speed metric.

NRF-GS: Neural Residual Fields for Expressive and Compact Gaussian Splatting

NRF-GS replaces per-Gaussian spherical-harmonic appearance with shared low-/high-frequency neural residual branches, achieving higher average PSNR with fewer Gaussians under unchanged densification and pruning rules, but not uniformly better perceptual quality or rendering speed.

Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction

AFG uses changes in adjacent frames' internal features to generate a positive frame-level write weight in frozen CUT3R/TTT3R inference, multiplying it with the token gate to reduce redundant overwriting and lower mean KITTI ATE from 68.54 m to 43.96 m without discarding frames, retraining, or growing memory, although not every sequence or error metric improves.

SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding

SceneScaffold allocates a fixed budget of 100 visual tokens to details, entities, spatial references, relations, and a global summary, organizing scene evidence before language reasoning and improving ScanRefer / Multi3DRefer mIoU from 43.3 / 42.7 to 47.0 / 47.9 over 3D-LLaVA.

Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

Seeing Speech learns facial motion along spreading, opening, and protrusion coordinate directions, retrieves local motion features through a speechโ€“articulatory memory, and composes them into 3D mesh animation through topology-aware residuals, reducing facial and lip reconstruction errors on VOCASET and TFHP without recovering the physiological mechanisms of actual speech organs.

Seg3DParts: Segmentation-Grounded Controllable Part-Level 3D Generation

Seg3DParts injects externally supplied part segmentations into a two-stage 3D generator and exchanges structural information across part latents, directly producing separate meshes in a space that preserves training-time relative coordinates; its strength is controllable decomposition with less interpenetration, not automatic discovery of unknown parts.

Towards Generalizable 3D Anomaly Detection via Relational Inconsistency Modeling

GRIM learns a defect criterion from controlled local relational violations in normal point clouds through edge-aware graph refinement and within-sample cluster-deviation modeling, achieving 97.4/94.5 object-/point-level AUROC on Anomaly-ShapeNet and 83.6/89.6 when transferred directly to Real3D-AD.