Skip to content

🧊 3D Vision

🤖 AAAI2026 · 79 paper notes

📌 Same area in other venues: 📷 CVPR2026 (751) · 🔬 ICLR2026 (197) · 🧪 ICML2026 (30) · 🧠 NeurIPS2025 (116) · 📹 ICCV2025 (267) · 🧪 ICML2025 (17)

🔥 Top topics: 3D Gaussian Splatting ×19 · Point Cloud ×13 · Segmentation ×7 · Shape Completion ×5 · 3D Object Detection ×5

3D-ANC: Adaptive Neural Collapse for Robust 3D Point Cloud Recognition

This paper introduces the Neural Collapse (NC) mechanism to 3D point cloud adversarial robustness, constructing a decoupled feature space with a fixed ETF classifier head and an adaptive training framework (RBL+FDL). This improves the adversarial accuracy of DGCNN on ModelNet40 from 27.2% to 80.9%, outperforming the best baseline by 34 percentage points.

3D-Free Meets 3D Priors: Novel View Synthesis from a Single Image with Pretrained Diffusion Guidance

Proposes a framework combining 3D-free methods (HawkI-style test-time optimization) with 3D-based priors (weak guidance maps from Zero123++) that generates camera-controlled views at specified elevation/azimuth angles from a single image without extra 3D data or training. It consistently outperforms Zero123++, HawkI, and Stable Zero123 across metrics like LPIPS and CLIP-Score in complex scenes.

3DTeethSAM: Taming SAM2 for 3D Teeth Segmentation

Adapts the SAM2 foundation model to the 3D teeth segmentation task by rendering 3D meshes into 2D images from multiple views. It designs three lightweight adapters (Prompt Embedding Generator, Mask Refiner, and Mask Classifier) and a Deformable Global Attention Plugin (DGAP) to address automatic prompting, boundary refinement, and semantic classification challenges, achieving a new state-of-the-art with 91.90% T-mIoU on the Teeth3DS dataset.

4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation

The 4DSTR framework is proposed, which significantly enhances the spatial-temporal consistency of 4D Gaussian generation and its adaptability to rapid temporal changes through Mamba-based temporal correlation rectification (correcting the scale and rotation of Gaussian points) and a per-frame adaptive densification and pruning strategy.

Adapt-As-You-Walk Through the Clouds: Training-Free Online Test-Time Adaptation of 3D Vision-Language Foundation Models

Uni-Adapter is proposed—a training-free online test-time adaptation framework for 3D Vision-Language Foundation Models (VLFMs). It registers SOTA performance across multiple 3D corruption benchmarks by tackling distribution shifts with cluster-based dynamic prototype caching and graph-regularized label smoothing.

AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation

This paper reveals a key issue in SDS: the source distribution is dynamically evolving rather than static. To address this, it proposes AnchorDS, which anchors the source distribution by feeding the current rendered image as an image condition into a dual-conditioned diffusion model. This resolves semantic over-smoothing and multi-view inconsistency in SDS, comprehensively outperforming SDS/VSD/SDS-Bridge on T3Bench.

AnchorHOI: Zero-shot Generation of 4D Human-Object Interaction via Anchor-based Prior Distillation

This paper proposes AnchorHOI, which distills interaction and motion priors from image/video diffusion models via two intermediate bridges—anchor NeRF and anchor keypoints. This achieves zero-shot text-driven 4D human-object interaction (HOI) generation, outperforming existing methods in both static 3D and dynamic 4D HOI generation.

Arbitrary-Scale 3D Gaussian Super-Resolution

This paper proposes the Arbi-3DGSR integrated framework. Comprising three core components—scale-aware rendering, generative-prior-guided optimization, and progressive super-resolution—it achieves, for the first time, high-resolution rendering at arbitrary (including non-integer) scales using a single 3DGS model. It improves the PSNR by 6.59 dB compared to vanilla 3DGS at a \(\times 5.7\) scale while maintaining a real-time rendering speed of 85 FPS.

ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance Segmentation

This paper proposes the ASSIST-3D synthetic data pipeline, which generates high-quality annotated data for class-agnostic 3D instance segmentation through three stages: heterogeneous object selection, LLM-guided scene layout generation, and realistic point cloud construction, significantly improving model generalization.

Can Protective Watermarking Safeguard the Copyright of 3D Gaussian Splatting?

This work systematically reveals the vulnerability of 3DGS watermarking frameworks for the first time, proposing the GSPure framework. GSPure accurately separates and removes watermark-related Gaussian primitives through view-aware weight accumulation and geometric feature clustering, reducing watermark PSNR by up to 16.34dB while keeping the original scene loss below 1dB.

Cheating Stereo Matching in Full-Scale: Physical Adversarial Attack against Binocular Depth Estimation

Proposes the first 3D full-surface texture physical adversarial attack targeting stereo matching models. Through a stereo-aligned rendering module and region-aware merging attacks, the adversarial vehicle is seamlessly blended into the background within depth maps, leading to severe failures in autonomous driving perception systems.

Class-Partitioned VQ-VAE and Latent Flow Matching for Point Cloud Scene Generation

Proposes Class-Partitioned VQ-VAE (CPVQ-VAE) and Latent Flow Matching Model (LFMM) to realize the first pure point cloud scene generation method that does not require external database retrieval, reducing Chamfer Distance by 70.4% on complex living room scenes.

CLIPPan: Adapting CLIP as A Supervisor for Unsupervised Pansharpening

This paper proposes CLIPPan, which adapts CLIP via parameter-efficient fine-tuning to understand multispectral/panchromatic/high-resolution multispectral image types and the pansharpening process. It leverages text prompts such as Wald's protocol as semantic supervision signals to achieve unsupervised pansharpening at full resolution without ground truth. It can serve as a plug-and-play module compatible with any pansharpening backbone network.

DANCE: Density-Agnostic and Class-Aware Network for Point Cloud Completion

The proposed DANCE framework achieves density-agnostic point cloud completion through ray-based candidate point sampling and an opacity prediction mechanism, while introducing a classification head to provide semantic priors, achieving state-of-the-art performance on PCN and MVP benchmarks.

DAPointMamba: Domain Adaptive Point Mamba for Point Cloud Completion

This work introduces Mamba (SSM) into Unsupervised Domain Adaptive Point Cloud Completion (UDA PCC) for the first time, proposing the DAPointMamba framework. By utilizing three modules—cross-domain patch-level scanning, cross-domain spatial SSM alignment, and cross-domain channel SSM alignment—it achieves high-quality cross-domain point cloud completion while maintaining linear complexity and a global receptive field.

Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting

This work proposes the TD-Attn framework, which integrates two modules—3D-Aware Attention Guidance (3D-AAG) and Hierarchical Attention Modulation (HAM)—to resolve the multi-view inconsistency (Janus problem) in 3D generation/editing caused by prior viewpoint bias in T2I diffusion models. It can be integrated into existing 3DGS frameworks as a plug-and-play module.

DeepRAHT: Learning Predictive RAHT for Point Cloud Attribute Compression

This paper proposes DeepRAHT, the first end-to-end differentiable Region Adaptive Hierarchical Transform (RAHT) framework for lossy point cloud attribute compression. By leveraging a learnable prediction model and a Laplace-based rate proxy, it achieves compression performance surpassing the G-PCC standard and existing deep learning methods.

Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection

This work proposes the FTKD (Future Temporal Knowledge Distillation) framework. By utilizing two strategies—Future-aware Feature Reconstruction (FFR) and Future-guided Logit Distillation (FLD)—it effectively transfers future frame knowledge from an offline teacher model to an online student model, achieving a 1.3 mAP / 1.3 NDS improvement on nuScenes without adding any inference overhead.

Domain Generalized Stereo Matching with Uncertainty-guided Data Augmentation

UgDA-Stereo is proposed to simulate visual styles of various unseen domains by applying batch-statistics-based Gaussian uncertainty perturbations to the channel-wise means and standard deviations of RGB images. Combined with feature consistency constraints, this plug-and-play approach significantly enhances the cross-domain generalization capability of stereo matching models.

Dynamic Gaussian Scene Reconstruction from Unsynchronized Videos

Proposes a coarse-to-fine temporal alignment module that can be plugged into existing 4D Gaussian Splatting frameworks. It addresses the degradation of dynamic scene reconstruction quality caused by temporal desynchronization in multi-view videos, significantly improving PSNR/SSIM/LPIPS of multiple baseline methods on the DyNeRF dataset.

Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization

This work proposes the WeSTAR framework, which synergizes semantic-aware hierarchical normalized self-training, sparse pairwise ordinal weak supervision, and LoRA weight regularization. This parameter-efficient approach enhances the generalization capability of the depth estimation foundation model (Depth Anything V2) on unseen domains and corrupted data, achieving SOTA performance on multiple OOD benchmarks.

Enhancing Rotation-Invariant 3D Learning with Global Pose Awareness and Attention Mechanisms

This paper proposes the Shadow-informed Pose Feature (SiPF) and the RIAttnConv operator, which enhance the global pose awareness of local rotation-invariant features by introducing global "shadow" reference points learned via Bingham distribution. This addresses the "Wing-tip Feature Collapse" issue where symmetric structures (such as left and right wings of an airplane) cannot be distinguished, achieving SOTA results on ModelNet40 classification and ShapeNetPart segmentation.

EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios

This paper proposes EPSegFZ, a pre-training-free 3D point cloud few-shot/zero-shot semantic segmention framework. By utilizing ProERA to extract high-frequency features, LGPE to fuse textual information for prototype updates, and DRPE to establish precise query-prototype correspondences, EPSegFZ outperforms SOTA methods on S3DIS and ScanNet by 5.68% and 3.82% respectively.

FantasyStyle: Controllable Stylized Distillation for 3D Gaussian Splatting

This paper proposes FantasyStyle, the first 3DGS style transfer framework entirely based on diffusion model distillation. It utilizes a Multi-view Frequency Consistency (MVFC) mechanism to suppress low-frequency components and reduce inconsistencies between perspectives, and designs Controllable Stylized Distillation (CSD) to introduce negative guidance, eliminating content leakage from style images. It outperforms existing VGG and diffusion-based methods in both stylization quality and content preservation.

FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM

By injecting the geometric priors of depth foundation models into an optical flow-based SLAM system, a closed-loop system is formed through three modules: a hybrid optical flow network, a bi-consistent BA layer, and reliability-aware refinement. It achieves SOTA trajectory accuracy and dense reconstruction quality on four major datasets (TUM, EuRoC, 7Scenes, and ETH3D) while running in real-time at 18 FPS.

Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine

FFSE is proposed—an autoregressive 3D-aware image editing framework based on video diffusion models. Combined with a hybrid dataset 3DObjectEditor (real + synthetic), it enables multi-round object translation, scaling, and rotation on real images, similar to a 3D engine. It simultaneously generates realistic background effects such as shadows, reflections, and occlusions, maintaining consistency across editing rounds. It significantly outperforms existing methods in both single-round and multi-round editing.

Gaussian Blending: Rethinking Alpha Blending in 3D Gaussian Splatting

Revisits scalar alpha blending in 3DGS, pointing out that ignoring intra-pixel spatial variation is the root cause of multi-scale rendering artifacts (erosion when zoomed in and dilation when zoomed out). It proposes Gaussian Blending, which models alpha and transmittance as an intra-pixel spatial distribution (a 2D uniform window) to achieve real-time anti-aliasing without retraining, improving the PSNR from 31.59 to 35.80 on Multi-scale Blender.

GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting

GaussianImage++ is proposed to achieve high-quality image representation and compression under limited 2D Gaussian primitives through a distortion-driven densification mechanism and content-aware Gaussian filters, while maintaining real-time decoding speed.

Generalized Geometry Encoding Volume for Real-time Stereo Matching

This paper proposes GGEV, which integrates depth priors from a monocular depth foundation model (Depth Anything V2) into the cost aggregation process in a lightweight manner. It adaptively enhances matching relationships for different disparity hypotheses through Depth-Aware Dynamic Cost Aggregation (DDCA), achieving strong generalization capabilities at real-time speeds.

Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo under Limited Multi-Illumination Cues

This paper proposes GeoUniPS, which introducesgeometric priors from a large-scale 3D reconstruction model (VGGT) into a universal photometric stereo network for the first time. Through a dual-branch illumination-geometry encoder, geometric priors are leveraged to compensate for insufficient multi-illumination cues. Additionally, a perspective projection training dataset, PS-Perp, is introduced to bridge the gap between orthographic projection assumptions and real-world scenes.

Graph Smoothing for Enhanced Local Geometry Learning in Point Cloud Analysis

Analyzes the issues of conventional graph construction methods (such as ball query) generating sparse connections at boundary points and noisy connections at intersection areas. Proposes a graph smoothing module (symmetric adjacency optimization + von Neumann kernel) and a local geometry learning module (adaptive shape features + cylindrical coordinate transformation), achieving competitive performance on classification and segmentation tasks.

Griffin: Aerial-Ground Cooperative Detection and Tracking Dataset and Benchmark

Introduces Griffin, an aerial-ground cooperative (AGC) 3D perception dataset and benchmark framework. It comprises 250+ dynamic scenes (37K+ frames), achieving realistic UAV dynamics, variable cruising altitudes (20–60m), and occlusion-aware annotations through CARLA-AirSim co-simulation, alongside a systematic robustness evaluation protocol.

GSAP-ERE: Fine-Grained Scholarly Entity and Relation Extraction Focused on Machine Learning

Proposes GSAP-ERE, a fine-grained scholarly entity and relation extraction dataset for the machine learning domain. It contains 10 entity types and 18 relation types, with 63K entities and 35K relations annotated across 100 full-text papers. Experimental results demonstrate that fine-tuned models (NER: 80.6%, RE: 54.0%) significantly outperform LLM prompting methods (NER: 44.4%, RE: 10.1%).

GT2-GS: Geometry-aware Texture Transfer for Gaussian Splatting

This paper proposes the GT2-GS framework, which achieves high-quality and view-consistent 3DGS texture transfer through a geometry-aware texture transfer loss, an adaptive fine-grained control module, and a geometry-preserving branch. It outperforms existing 3D style transfer methods in both texture fidelity and scene content preservation.

Hierarchical Direction Perception via Atomic Dot-Product Operators for Rotation-Invariant Point Clouds Learning

This paper proposes DiPVNet, which builds a local L2DP operator and a global DASFT module based on the dual properties of the atomic dot-product operator (directional selectivity + rotation invariance) to achieve hierarchical direction-aware rotation-invariant point cloud learning.

IE-SRGS: An Internal-External Knowledge Fusion Framework for High-Fidelity 3D Gaussian Splatting Super-Resolution

This paper proposes the IE-SRGS framework, which reconstructs high-fidelity super-resolution 3DGS from low-resolution inputs. It fuses high-frequency texture priors from external 2D super-resolution models (external knowledge) with cross-view consistent depth and texture features from multi-scale 3DGS models (internal knowledge) through a mask-guided fusion strategy, achieving state-of-the-art performance on both synthetic and real-world scenes.

Learning Conjugate Direction Fields for Planar Quadrilateral Mesh Generation

This paper proposes an efficient, data-driven method based on DGCNN to generate conjugate direction fields (CDFs), bypassing the high computational overhead of traditional non-linear optimization. It supports user-stroke-guided controllable CDF generation, speeding up CDF computation by 1 to 2 orders of magnitude. Along with the method, a large-scale dataset containing over 50,000 free-form surfaces is released.

MeshA*: Efficient Path Planning With Motion Primitives

This paper proposes the MeshA algorithm, which shifts lattice-based path planning from "searching at the motion primitive level" to "searching at the grid cell level while simultaneously fitting primitive sequences." By defining a new search space called the "extended cell," MeshA achieves a 1.5x-2x runtime speedup compared to standard LBA* while ensuring completeness and optimality.

MeshSplat: Generalizable Sparse-View Surface Reconstruction via Gaussian Splatting

MeshSplat is proposed as the first generalizable sparse-view surface reconstruction framework based on 2D GS. By introducing a weighted Chamfer Distance loss to regularize depth predictions and an uncertainty-based normal prediction network to align 2D GS orientations, it learns geometric priors from novel view synthesis in a self-supervised manner, achieving state-of-the-art performance in sparse-view mesh reconstruction and cross-dataset generalization.

MoBGS: Motion Deblurring Dynamic 3D Gaussian Splatting for Blurry Monocular Video

MoBGS proposes an end-to-end dynamic deblurring 3D Gaussian Splatting framework. Through two core modules, Blur-adaptive Latent Camera Estimation (BLCE) and Latent Camera-induced Exposure Estimation (LCEE), it reconstructs sharp spatiotemporal novel views from blurry monocular videos, significantly outperforming existing SOTA methods on the Stereo Blur dataset.

MonoCLUE: Object-Aware Clustering Enhances Monocular 3D Object Detection

This paper proposes MonoCLUE, which extracts object-level visual patterns (such as hoods, roofs, and other parts) through local clustering and aggregates consistent appearance features across images using generalized scene memory. This enhances the detection capability for occluded and truncated objects in monocular 3D detection, achieving SOTA performance on the KITTI benchmark without relying on extra depth or LiDAR information.

MR-CoSMo: Visual-Text Memory Recall and Direct Cross-Modal Alignment Method for Query-Driven 3D Segmentation

MR-CoSMo is proposed, a coarse-to-fine query-driven 3D segmentation model. It establishes explicit alignment between 3D point clouds and text/2D images via a Direct Cross-Modal Alignment (DCMA) module, and integrates a visual-text memory module (Memory Module) to store high-confidence feature pairs to enhance cross-scene segmentation consistency. It achieves state-of-the-art (SOTA) performance across three tasks: 3D instruction segmentation, referring segmentation, and semantic segmentation.

Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection

This paper proposes MMAssist, which leverages image and text features as a "bridge" to align 3D features between the source and target domains, while combining 2D detection results to enhance the quality of pseudo-labels, significantly improving the performance of LiDAR-based 3D unsupervised domain adaptation for object detection.

NURBGen: High-Fidelity Text-to-CAD Generation through LLM-Driven NURBS Modeling

This paper proposes NURBGen, the first text-to-CAD generation framework based on NURBS surface representation. By fine-tuning an LLM to convert natural language descriptions into structured NURBS parameter JSONs, and introducing a hybrid representation (untrimmed NURBS + analytical primitives) alongside the large-scale partABC dataset, it significantly outperforms existing methods in geometric fidelity and dimensional accuracy.

OceanSplat: Object-aware Gaussian Splatting with Trinocular View Consistency for Underwater Scene Reconstruction

OceanSplat is proposed, which achieves high-fidelity underwater 3D Gaussian Splatting scene reconstruction under scattering media by incorporating trinocular view consistency constraints, synthetic epipolar depth priors, and depth-aware alpha adjustments, significantly reducing floating artifacts and outperforming existing methods.

Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning

Proposes a unified framework, OSU-3DSG, which integrates vision-language models for open-world 3D scene graph generation and supports four interactive tasks (scene question answering, visual grounding, instance retrieval, and task planning) via retrieval-augmented reasoning, achieving performance comparable to supervised methods under unsupervised settings.

OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding

This paper proposes the Generalized Open-Vocabulary 3D Scene Understanding (GOV-3D) task and the corresponding OpenScan benchmark, extending 3D scene understanding from object categories to eight linguistic property dimensions, revealing the severe limitations of existing OV-3D methods in understanding abstract object properties.

Opt3DGS: Optimizing 3D Gaussian Splatting with Adaptive Exploration and Curvature-Aware Exploitation

This paper proposes the Opt3DGS framework, which divides 3DGS training into exploration and exploitation stages. The exploration stage utilizes adaptive weighted SGLD to escape local optima, while the exploitation stage adopts a local quasi-Newton Adam optimizer for precise convergence, achieving SOTA rendering quality without modifying the Gaussian representation.

Parameter-Free Fine-tuning via Redundancy Elimination for Vision Foundation Models

This work identifies a significant number of redundant channels in vision foundation models (such as SAM, SAM2, and DINOv2) and proposes a parameter-free fine-tuning method. By employing an output-difference-based channel selection algorithm to locate optimal replacement pairs, redundant channels are replaced with active ones to enhance feature representations for downstream tasks, achieving an average mIoU improvement of 5 to 11 points.

Parameterized Approximation Algorithms for TSP on Non-Metric Graphs

This paper proposes improved FPT approximation algorithms for the Traveling Salesman Problem (TSP) on non-metric graphs, parameterized by \(p\) (the number of vertices violating the triangle inequality) and \(q\) (the size of a minimum violator set). It improves the approximation ratio from 2.5 to 1.5 under parameter \(p\), and from 11 to 3 under parameter \(q\).

Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph Network

Pb4U-GNet is proposed, which decouples message propagation from feature updating (Propagation-before-Update). By combining resolution-aware propagation depth control and an update scaling mechanism, it achieves garment simulation that generalizes to high-resolution meshes after training only on low-resolution meshes.

PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos

Proposes PFAvatar, a two-stage approach (pose-aware diffusion model fine-tuning + NeRF distillation) to reconstruct high-quality 3D human avatars from real-world "Outfit-of-the-Day" (OOTD) photos, achieving personalization in just 5 minutes, representing a 48x speedup compared to prior methods.

Physics-Informed Deformable Gaussian Splatting: Towards Unified Constitutive Laws for Time-Evolving Material Field

By treating each 3D Gaussian as a Lagrangian material point and introducing a time-evolving material field to predict particle velocity and the constitutive stress tensor, this work incorporates the Cauchy momentum residual as a physical constraint alongside Lagrangian particle flow matching as a data-fitting term. This approach achieves physical consistency and cross-scene generalization in monocular dynamic novel view synthesis, achieving state-of-the-art (SOTA) performance on both a self-constructed physics-driven dataset and the HyperNeRF dataset.

Point-SRA: Self-Representation Alignment for 3D Representation Learning

Point-SRA is proposed to enhance 3D point cloud representation learning by leveraging the complementarity of representations under different mask ratios through Dual Self-Representation Alignment (MAE layer + MFT layer) and MeanFlow probabilistic modeling, outperforming Point-MAE by 5.59% on ScanObjectNN.

Point Cloud Quantization through Multimodal Prompting for 3D Understanding

Proposes PCQ (Point Cloud Quantization), which leverages text embeddings from pre-trained vision-language models as semantic prototypes. It discretizes continuous point cloud features into the text prototype space using Gumbel-Softmax differentiable quantization, achieving significant improvements in 3D understanding when combined with cross-modal feature fusion.

PressTrack-HMR: Pressure-Based Top-Down Multi-Person Global Human Mesh Recovery

This paper proposes PressTrack-HMR, the first top-down pipeline for multi-person global human mesh recovery based solely on pressure signals. It achieves pressure footprint tracking via an innovative UoE similarity metric (93.6% MOTA) and builds MIP, the first multi-person interactive pressure dataset.

Real-Time 3D Object Detection with Inference-Aligned Learning

This paper proposes the SR3D framework, which bridges the gap between training and inference behaviors in indoor dense 3D object detection through two training-phase components: Spatial-Prioritized Optimal Transport Assignment (SPOTA) and Rank-Aware Adaptive Self-Distillation (RAS). It refreshes the state-of-the-art (SOTA) for dense detectors on ScanNet V2 and SUN RGB-D while maintaining a real-time inference speed of 42ms.

Redundant Queries in DETR-Based 3D Detection: Unnecessary and Prunable

Proposes GPQ (Gradually Pruning Queries) to progressively prune a large number of redundant object queries in DETR-based 3D detectors based on classification scores. Removing queries requires no extra learnable parameters and can be directly accomplished by fine-tuning on pre-trained checkpoints, achieving up to a 67.86% FLOPs reduction and a 65.16% inference time reduction on edge devices.

Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective

A new paradigm, Completion-by-Correction, is proposed. It utilizes a pretrained image-to-3D model to generate a topologically complete shape prior, which is then corrected in the feature space to align with partial observations, replacing the traditional Completion-by-Inpainting method. It achieves a 23.5% reduction in average CD and a 7.1% improvement in F-score on ShapeNetViPC.

Rethinking Rainy 3D Scene Reconstruction via Perspective Transforming and Brightness Tuning

Proposes the OmniRain3D dataset (the first rainy 3D scene dataset that simultaneously models perspective heterogeneity and brightness dynamicity) alongside the REVR-GSNet end-to-end framework (joint recursive brightness enhancement + Gaussian primitives optimization + GS-guided deraining) to reconstruct high-fidelity clean 3D scenes from rain-degraded images.

Retrieving Objects from 3D Scenes with Box-Guided Open-Vocabulary Instance Segmentation

A Box-Guided approach is proposed, which leverages detection boxes from the 2D open-vocabulary detector YOLO-World to guide the construction of 3D instance masks from superpoints. Eliminating the need for SAM and CLIP, it achieves high efficiency (< 1 minute per scene) while significantly improving the retrieval capability for low-frequency target categories.

RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image

RTGaze is proposed, a real-time 3D-aware gaze redirection method. By utilizing a hybrid-frequency feature encoder, a gaze injection module, and 3D facial geometric prior distillation, it achieves high-quality gaze redirection from a single image at 61ms/frame, which is over 800 times faster than the previous SOTA 3D method (GazeNeRF).

Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion

This paper proposes the Simba framework, which reformulates point cloud completion as "diffusion on geometric transformation fields" rather than "diffusion on point coordinates." By utilizing Sym-Diffuser to learn the conditional distribution of point-wise affine transformations, it generates a coarse completion. Subsequently, a cascaded Mamba architecture (MBA-Refiner) is employed to progressively refine it to high-fidelity outputs, achieving SOTA performance across multiple benchmarks including PCN, ShapeNet, and KITTI.

SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images

This paper proposes SmartSplat, a feature-aware 2D Gaussian Splatting image compression framework. By adopting three coordinate strategies—gradient-color-guided variational sampling, repulsive uniform sampling, and scale-adaptive color initialization—it achieves high-quality reconstruction of 8K/16K ultra-high-resolution images under extreme compression ratios (up to 5000×) for the first time.

Sparse4DGS: 4D Gaussian Splatting for Sparse-Frame Dynamic Scene Reconstruction

This paper proposes Sparse4DGS, the first 4D dynamic scene reconstruction method designed for sparse-frame inputs. Through two core modules, Texture-Aware Deformation Regularization (TADR) and Texture-Aware Canonical Optimization (TACO), it guides the Gaussian distribution to focus on texture-rich areas, achieving high-quality dynamic novel view synthesis with only 5–30 sparse input frames.

SparseSurf: Sparse-View 3D Gaussian Splatting for Surface Reconstruction

Proposes SparseSurf, which enhances geometric consistency under sparse views through Stereo Geometry-Texture Alignment and Pseudo-Feature Enhanced Geometry Consistency, simultaneously achieving high-precision surface reconstruction and high-quality novel view synthesis, achieving SOTA on DTU, BlendedMVS, and Mip-NeRF360 datasets.

Splat-SAP: Feed-Forward Gaussian Splatting for Human-Centered Scene with Scale-Aware Point Map Reconstruction

Proposes Splat-SAP, a feed-forward method that reconstructs scale-aware point maps from highly sparse binocular camera inputs, enabling high-quality free-viewpoint rendering of human-centered scenes via Gaussian Planes without any 3D supervision.

Splats in Splats: Robust and Effective 3D Steganography towards Gaussian Splatting

Introduces Splats in Splats, the first steganography framework that embeds 3D content into 3DGS assets without modifying any vanilla 3DGS attributes. Through importance-graded spherical harmonics (SH) coefficient encryption and autoencoder-assisted opacity mapping, it achieves 5.31% higher scene fidelity and 3x faster rendering speeds.

SplatSSC: Decoupled Depth-Guided Gaussian Splatting for Semantic Scene Completion

This paper proposes SplatSSC, which addresses the issue of inefficient random initialization and floating artifacts from outlier primitives in the object-centric paradigm through a depth-guided Gaussian primitive initialization strategy and a Decoupled Gaussian Aggregator (DGA). It achieves an IoU gain of 6.3% and a mIoU gain of 4.1% on Occ-ScanNet, while reducing latency and memory costs by over 9.3%.

Split-Layer: Enhancing Implicit Neural Representation by Maximizing the Dimensionality of Feature Space

Proposed Split-Layer, which splits the MLP fully connected layer into multiple parallel branches and integrates their outputs using the Hadamard product. Without increasing parameters or computation, this exponentially expands the feature space dimensionality from \(C\) to \(\binom{C/\sqrt{N}+N-1}{N}\), significantly enhancing the representation capability of Implicit Neural Representations (INR).

STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification

STMI proposes a three-component multi-modal object re-identification framework. It suppresses background noise through SAM segmentation-guided feature modulation (SFM), extracts compact representations via semantic token reallocation (STR), and captures high-order semantic relationships with cross-modal hypergraph interaction (CHI), achieving significant improvements on benchmarks such as RGBNT201.

StreamSTGS: Streaming Spatial and Temporal Gaussian Grids for Real-Time Free-Viewpoint Video

The authors propose StreamSTGS, a streamable spatial-temporal Gaussian grid representation. It encodes canonical 3D Gaussian attributes into 2D images and temporal features into videos, enabling real-time free-viewpoint video streaming (frame size of only 170KB) while ensuring reconstruction quality (PSNR 32.30dB) via a Transformer-guided auxiliary training and a sliding window mechanism.

Surface-Based Visibility-Guided Uncertainty for Continuous Active 3D Neural Reconstruction

Proposes Surface-Based Visibility (SBV), which utilizes SDF-derived surface confidence and a voxel grid update mechanism to accurately estimate visibility of uncertainty during continuous active learning. Guided by SBV, Next-Best View selection achieves an image rendering quality improvement of up to 11.6% across four benchmarks: DTU, Blender, TanksAndTemples, and BlendedMVS.

TG-Field: Geometry-Aware Radiative Gaussian Fields for Tomographic Reconstruction

This paper proposes TG-Field, a geometry-aware Gaussian deformation framework for extremely sparse-view CT reconstruction. By incorporating a multi-resolution hash encoder to model spatial geometric priors, alongside a spatiotemporal attention module and a motion flow network to handle dynamic CT, it achieves state-of-the-art (SOTA) performance in both static and dynamic CT reconstruction.

TOSC: Task-Oriented Shape Completion for Open-World Dexterous Grasp Generation from Partial Point Clouds

This work proposes a new task named Task-Oriented Shape Completion (TOSC), which only reconstructs task-relevant contact areas rather than the entire object. By generating candidates with pre-trained foundation models, filtering the optimal shape with a 3D Discriminative Autoencoder (DAE), and synthesizing dexterous grasps via a FlowGrasp flow-matching model, the proposed method yields performance gains of 16.17% in grasp displacement and 55.26% in Chamfer distance.

Towards Temporal Fusion Beyond the Field of View for Camera-based Semantic Scene Completion

The C3DFusion module is proposed to explicitly align point features of historical and current frames in 3D space, systematically addressing the temporal completion problem of out-of-frame regions in camera-based SSC for the first time, achieving SOTA on SemanticKITTI and SSCBench-KITTI-360.

UniC-Lift: Unified 3D Instance Segmentation via Contrastive Learning

Ours proposes UniC-Lift, a unified single-stage 3D instance segmentation framework. By learning optimizable vector embeddings in 3DGS primitives and training them with contrastive and triplet losses, it directly decodes consistent 3D segmentation labels through a simple "Embedding-to-Label" process. This eliminates post-processing clustering steps like HDBSCAN, reducing training time from over 15 hours to less than 40 minutes.

VGGT-DP: Generalizable Robot Control via Vision Foundation Models

The paper proposes VGGT-DP, a bio-inspired visuomotor policy framework that combines the pre-trained 3D-aware foundation model VGGT as a vision encoder with Diffusion Policy. Through three key designs—frame-wise token reuse, random token pruning, and proprioception-guided vision learning—it significantly outperforms DP and DP3 baselines on high-precision manipulation tasks in MetaWorld.

VPN: Visual Prompt Navigation

Proposes a new paradigm of Visual Prompt Navigation (VPN): users annotate visual trajectories (arrows connecting key waypoints) on a 2D top-down view to guide agent navigation, replacing natural language instructions and image goal instructions. Constucts two datasets, R2R-VP and R2R-CE-VP, along with the VPNet baseline model. Combining view-level and trajectory-level data augmentation leads to outstanding performance in both discrete and continuous environments.