ECCV2026 Paper Notes TODO¶
总计: 771 篇 | 已完成: 771 | 待更新: 0
- 2D Features Are All You Need for 3D Shape Understanding | AC: 5088
- 2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction | arXiv: 2603.19964 | AC: 4194
- 340 FPS Reflection-free Video from Spikes Modulated by a Rapidly Rotating Polarizer | AC: 4598
- 360Anything: Geometry-Free Lifting of Images and Videos to 360° | arXiv: 2601.16192 | AC: 3670
- 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents | AC: 5213
- 360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method | AC: 4124
- 3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism | AC: 3820
- 3D Field of Junctions: A Noise-Robust, Training-Free Structural Prior for Volumetric Inverse Problems | arXiv: 2603.02149 | AC: 5956
- 3D Gaussian Splatting Compression with Object Scalability | AC: 4180
- 3D Gaussian Texture for Real-time Mesoscale Appearance Synthesis and Rendering | AC: 5185
- 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement | AC: 4208
- 3D-Aware VLMs with Implicit and Explicit Geometries | AC: 3224
- 3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing | AC: 4525
- 3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis method for Single-View Aerial-Ground Re-Identification | AC: 5389
- 3D-ReGen: A Unified 3D Geometry Regeneration Framework | AC: 5538
- 3DGS3: Joint Super Sampling and Frame Interpolation for Real-Time Large-Scale 3DGS Rendering | AC: 3945
- 3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints | AC: 4788
- 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering | AC: 5396
- 4D-VGGT: A SpatioTemporal Foundation Model for Dynamic Scene Geometry Estimation | AC: 3366
- 4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video | AC: 4866
- A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation | AC: 3921
- A Benchmark for Heterogeneous Stereo Deblurring with Physically- and Epipolar-constrained Cross Attention | AC: 3687
- A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP | arXiv: 2606.30342 | AC: 4853
- A Comprehensive Analysis about Unsupervised Outlier Detection for Images | AC: 5320
- A Dual-space Patch-driven Complementary Learning Framework for Semi-supervised Multi-organ Segmentation | AC: 3317
- A Dual-Transformer Architecture with Cross-Attention for Multi-Camera View Recommendation | AC: 5242
- A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR | arXiv: 2606.31225 | AC: 3216
- A Mechanism-Driven Theory of Phase Transitions in Active Learning | arXiv: 2607.00144 | AC: 4856
- A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models | AC: 3665
- A Scalable Vector Graphics Latent Space | AC: 5250
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world | arXiv: 2606.21216 | AC: 3432
- A second-order theory of texture for depth from focus | AC: 3452
- A Simple Baseline with Placement Prior for Point-Supervised Oriented Object Detection | AC: 5334
- A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning | AC: 3311
- Abstract the Layout, Focus the Detail: A Dual-Granularity Representation Framework for Zero-Shot 3D Visual Grounding | AC: 5726
- AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation | arXiv: 2606.31204 | AC: 5183
- AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation | AC: 4669
- Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation | arXiv: 2606.31323 | AC: 4695
- Accelerating Diffusion Models via Equal-Risk Caching | AC: 5300
- Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache | AC: 5921
- Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction | arXiv: 2606.24156 | AC: 3636
- Accelerating Text-to-Video Generation with Calibrated Sparse Attention | AC: 3436
- Accurate Zero-shot Quantization via Hierarchical Teacher-Assistant Distillation | AC: 4699
- Achieving Subcategorical Erasure in Text-to-Image Models | AC: 4524
- ActionParty: Multi-Subject Action Binding in Generative Video Games | AC: 4221
- ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning | AC: 5269
- Activation Quantization of Vision Encoders Needs Prefixing Registers | AC: 5152
- Active View Selection with Perturbed Gaussian Ensemble for Tomographic Reconstruction | arXiv: 2603.06852 | AC: 3847
- ActiveStructure: Plane Scene Graph-Guided Active 3D Gaussian Splatting | AC: 5184
- Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO | AC: 4730
- Ada-VNNs: Adaptive Equivariance for Vector Neural Networks | AC: 4500
- AdaBoosting Text Prompts for Vision-Language Models | arXiv: 2607.00684 | AC: 4284
- AdaBridge-SR: Adaptive Bridge Matching for Real-World Image Super-Resolution | AC: 5676
- AdaDexGrasp: Adaptive Dexterous Grasping via 3D Visuo-Tactile Representation Fusion | AC: 5885
- ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs | arXiv: 2606.31054 | AC: 5033
- Adapting MLLMs for Nuanced Video Retrieval | AC: 4071
- Adaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation | AC: 4115
- Adaptive Neural Dynamics for Robust Geometric LiDAR-Inertial State Estimation on UAVs | AC: 3654
- Adaptive Noise Covariance Scheduling under Riemannian Metrics for Diffusion Models | AC: 3974
- Adaptive Spectrum-Aware Feature Disentangled Network for Small Object Detection | arXiv: 2606.29029 | AC: 3168
- AdaptiveSplat: Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction | AC: 4240
- AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking | AC: 4501
- Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods | arXiv: 2606.24484 | AC: 3683
- Adversarial Attack and Disturbance Detection by Hadamard-Coded Output Representations for Object Detection and Semantic Segmentation | AC: 5869
- Adversarial Score Distillation for Stable One-Step Diffusion in Real-World Image Super-Resolution | AC: 4022
- AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World | arXiv: 2606.29716 | AC: 4015
- AeroVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control | AC: 5844
- AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware | arXiv: 2602.16249 | AC: 5712
- AffoGato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale | arXiv: 2506.12009 | AC: 4336
- Affordance-Guided Diffusion Prior for 3D Hand Reconstruction | AC: 4223
- AGE: Agentic Gaussian Editing in 3D Scenarios | AC: 5221
- Agent-OBJ: Prompt-Driven 3D Adversaries for Multi-Modal Perception | AC: 5027
- Agentic Collaborative Cognition for Zero-Shot 3D Understanding | arXiv: 2606.24649 | AC: 4045
- AgentVLN: Towards Agentic Vision-and-Language Navigation | AC: 3693
- Aggregating Cross-Domain Knowledge via Learnable Tokens for Multi-Teacher Distillation | AC: 5104
- AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors | arXiv: 2603.17975 | AC: 3219
- AIMold: An Autonomous AI-based Pipeline for Complex Mold Design | AC: 5227
- AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting | AC: 3540
- AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision | arXiv: 2604.26567 | AC: 5421
- AiSCREAM: Absolute Target Localization with Language-Conditioned Cross-View Alignment for Autonomous Vehicles | AC: 4382
- Align and Segment: Unsupervised Learning for Building Segmentation From Misaligned Labels | AC: 5445
- Aligning Anything: Hierarchical Motion Estimation for Video Frame Interpolation | AC: 5280
- Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation | AC: 3296
- AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport | AC: 4288
- Allo{SR}2: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows | AC: 3463
- AlphaRad: Grounded Zero-Shot Classification in Chest Radiology via α-Corrected Binary Cross Entropy and Factorized Latent Supervision | AC: 3412
- AMCI: Unlock the Potential of Large Multimodal Models for Fine-grained Open-world Classification via Adaptive Memory Context Injection | AC: 3437
- AMCoNav: Asynchronous Multi-module Collaborative Framework for Embodied Visual Navigation | AC: 5222
- Amplify, Aggregate, and Adjust: VideoMAE-based Holistic-Subtle Aggregation for Micro-Action Recognition | AC: 4219
- An Inverse-Adversarial and Difficulty-Adaptive Robust Vision-Language Model | AC: 3380
- Analytic Bayesian Uncertainty for LiDAR Segmentation: A Single-pass Generative Approach | AC: 5831
- Analyzing and Improving Training-Free Fast Sampling of Text-to-Image Diffusion Models | AC: 3912
- AnaPFL: When Closed-Form Solutions Meet Generalizationand Personalization in Personalized Federated Learning | AC: 4595
- Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models | AC: 4828
- Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion | AC: 3196
- Anchored, Not Graded: How Vision-Language Models Fail at Slant-from-Texture Perception | arXiv: 2606.06714 | AC: 4561
- AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation | AC: 4292
- Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time | AC: 5400
- Anchoring on Reality: Breaking the Pseudo-Target Ceiling in Makeup Transfer | arXiv: 2606.31089 | AC: 3284
- AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning | AC: 4172
- AnchorSplat: Fast and Structure Consistent Detail Synthesis for Gaussian Splatting | AC: 4233
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories | AC: 5000
- AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution | AC: 4173
- ANFI: Rethinking Neighbor Feature Interaction in Person Re-ID | AC: 5111
- Anomaly Factory 3D: A Modular Framework for Diverse Pseudo-Anomaly Synthesis in Unsupervised 3D Anomaly Detection | AC: 4995
- Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation | AC: 3308
- Any to Full: Prompting Depth Anything for Depth Completion in One Stage | AC: 3803
- AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation | AC: 5144
- AnyGround3D: Towards Grounding Any 3D Object in the Wild via 2D-to-3D Lifting | AC: 3160
- AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images | arXiv: 2606.31077 | AC: 5344
- AnyStyle: A Single LoRA is Sufficient for Image-Guided Style Transfer | AC: 4851
- AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization | arXiv: 2603.17461 | AC: 3579
- AracNet: Revealing Debiasing Signals across Layers with Shallow Monitors | AC: 4940
- ARC-Loc: Leveraging Azimuthal Ray Convergence as a Geometric Cue for Direct Cross-View Localization | AC: 3676
- ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection | AC: 5146
- Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection | AC: 3983
- Are Video Reasoning Models Ready to Go Outside? | arXiv: 2603.10652 | AC: 4365
- ARGENT: Adaptive Hierarchical Image-Text Representations | AC: 5595
- ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search | AC: 4039
- ARMS: Anchor–Relational Motion Streaming for Seamless Solo-Social Motion Transitions | AC: 5555
- Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations | AC: 5435
- ART-VSR: Adaptive Rectified Trajectories for One-Step Video Super-Resolution | AC: 4653
- Articulat3D: Reconstructing Articulated Digital Twins From Monocular Videos with Geometric and Motion Constraints | arXiv: 2603.11606 | AC: 3496
- ARVAR: Accelerating Visual Autoregressive Model via Attention Retrospect | AC: 3564
- ASSCG: Just-Right Gating over Chattering for Fast–Slow LLM Planning in Autonomous Driving | AC: 3354
- ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving | arXiv: 2606.29286 | AC: 4652
- Asymmetric Anchoring: Opening the Black Box of MLLMs for Forgery Detection | AC: 5294
- Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation | AC: 3573
- ATOMIC: A Domain-Specific Vision-Language Model for Transmission Electron Microscopy | AC: 5128
- ATP-Bench: Towards the Agentic Tool Planning for MLLM Interleaved Generation | AC: 3485
- Attention is Case-Sensitive | AC: 3409
- Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment | AC: 4363
- Attention-based Vision-Language Memory for Spatial Reasoning | AC: 4166
- Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning | AC: 4203
- Attention-Logit Steering to Compositional Generalization for Continual VQA | AC: 5723
- Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models | AC: 3951
- Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video | arXiv: 2512.12165 | AC: 4578
- Audio-Visual Continual Test-Time Adaptation without Forgetting | AC: 4051
- Auto-Prompting: Layer-Specific Prompt Fusion Discovery via Differentiable Search | arXiv: 2606.26379 | AC: 4112
- Automatic Method Illustration Generation for AI Scientific Papers via Drawing Middleware Creation, Evolution, and Orchestration | AC: 4257
- Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens | AC: 4010
- AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation | arXiv: 2607.01051 | AC: 3578
- AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs | AC: 3550
- AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing | AC: 4102
- AV2T-Gen: Aerial Visible to Thermal Generation with Environment and Vehicle State Guidance | AC: 3919
- AViTS:Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation | AC: 3330
- AVQ-Attention: Adaptive Vector-Quantized Attention | AC: 4988
- AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution | arXiv: 2607.00987 | AC: 5371
- AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation | arXiv: 2606.30811 | AC: 5346
- A²-Edit: Precise Reference-Guided Image Editing of Arbitrary Objects and Ambiguous Masks | AC: 4028
- BAAF: Universal Transformation of One-Class Classifiers for Unsupervised Image Anomaly Detection | AC: 5489
- Back-Tracking from Clarity: Self-Learning to See Text from Afar | AC: 3551
- Background Blurring Matters: Improving Visual Grounding by Merging Text-Irrelevant Tokens | AC: 5081
- BackTranslation2.0 - A Linguistically Motivated Metric to Assess Sign Language Production | arXiv: 2606.28673 | AC: 4935
- BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension | AC: 5793
- BATQuant: Outlier-Resilient MXFP4 Quantization via Learnable Block-wise Optimization | AC: 4636
- Bayesian Self-Attention with Local Pixel Correlations for Lightweight Denoising Transformers | AC: 4718
- Bayesian Uncertainty Attribution-Guided Fine-Tuning for Open-Set Action Recognition | AC: 5989
- BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models | AC: 5972
- Be Tangential to Manifold: Discovering Riemannian Metric for Diffusion Models | AC: 4638
- Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning | AC: 4647
- Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset | AC: 4633
- Benchmarking Federated Learning & Knowledge Distillation for Point Cloud Classification | AC: 3844
- Benchmarking MLLMs on Mistake Recognition and Explanation in Single-Step Components of Cooking | AC: 5187
- Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench | AC: 5536
- Benchmarking Vision-Language Models for Microscopic Plant Image Understanding | AC: 4108
- Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation | AC: 3661
- BeTTER: Diagnose the Illusion of Embodied Reasoning in Vision-Language-Action Models | AC: 5364
- BEV-GS: Feed-forward Gaussian Splatting in Bird’s-Eye-View for Road Reconstruction | AC: 4436
- BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations | AC: 4104
- BEVOpen3D: Towards Open-World 3D Object Detection in Bird's-Eye-View | AC: 3368
- Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization | AC: 3796
- Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment | AC: 4698
- Beyond Aesthetics: Quantifying Information Loss in Turbid Scenes | AC: 4433
- Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models | AC: 4079
- Beyond Attention: Convolutional Global Context for Remote Sensing Change Detection | AC: 3835
- Beyond Categorical Matching: Intra-Class Graded Relevance Estimation for Cross-Modal 3D Retrieval | AC: 5096
- Beyond Common Sense: Grounding Logical Anomaly Detection in Inspection Criteria | AC: 4468
- Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation | AC: 3187
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents | AC: 4331
- Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation | AC: 4913
- Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives | AC: 3394
- Beyond Inpainting: Unleash 3D Understanding for Stable Camera-Controlled Video Re-rendering | AC: 3874
- Beyond Isolated Scans: Cross-Phase Alignment of Structure and Topology for 3D Medical Pretraining | AC: 5629
- Beyond Pixel Mimicry: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation | arXiv: 2606.23950 | AC: 5173
- Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation | AC: 5497
- Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR | AC: 5594
- BeyondSight: Object Permanence for End-to-End Autonomous Driving | AC: 5981
- BiCE-HG: A Bi-Conditional Egocentric Hand Gesture Dataset for Intelligent Reality Systems | AC: 5982
- BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming | arXiv: 2606.24740 | AC: 5202
- BioMTBee: Biologically Constrained Multi-View Template-Based 3D Reconstruction of Bumblebee | AC: 5834
- BitRIC: Efficient Neural Compression of LiDAR Range Images via Hierarchical Bitplanes | AC: 5317
- Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding | AC: 3141
- BLOB-Q: Boosting Low Bit ViT Quantization via Global Optimization on Model Distortion | AC: 5591
- Boba: Batched Simulation for Physics-Based Gaussian Digital Twins | AC: 4947
- Boosting 3D Foundation Models with Featureless Pose Optimization | arXiv: 2607.00579 | AC: 4837
- Boosting 6D Object Pose Estimation via Monocular Depth Cues | AC: 5757
- Boosting Correspondence Learning with Structure-Aware Estimator | AC: 4879
- Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation | AC: 4230
- Bootstrapping Articulated 3D Reconstruction from 2D Image Collections | AC: 5953
- Bottom-up modeling of repeated elements via single image analysis-by-synthesis | AC: 3696
- Bounding-Box Trajectories Matter for Video Anomaly Detection | AC: 5841
- BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure | arXiv: 2607.00573 | AC: 4762
- BrainRiem: Riemannian Prototype Learning for Source-Free Cross-Site Brain Network Diagnosis | arXiv: 2606.29200 | AC: 5246
- Breaking the Model Forgetting Cycle in Long-Incremental 3D Object Detection | AC: 5691
- BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning | AC: 4696
- BrepLLM: Enabling Large Language Models to Understand Boundary Representations | arXiv: 2512.16413 | AC: 5778
- Bridging the Geometry Mismatch: Frequency-Aware Anisotropic Serialization for Thin-Structure SSMs | AC: 5356
- Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction | arXiv: 2606.29445 | AC: 5028
- BWAFDA: Block-wise Weighted Attention Fusion with Detail-aware for No-Reference Image Quality Assessment | AC: 4196
- C2E: Boosting Ego-Only 3D Object Detection via Multi-Teacher Contrastive Knowledge Distillation | AC: 4362
- C3-Bench: A Context-Aware Change Captioning Benchmark | arXiv: 2606.25445 | AC: 3688
- CabinSI: Omni-Cabin Spatial Reasoning through Explicit Visual Cognitive Maps | AC: 3315
- Cambrian-P: Pose-Grounded Video Understanding | AC: 4080
- Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting | arXiv: 2606.26754 | AC: 4306
- CaPCL: Caption-Preserved Continual Learning for Text-to-Image Retrieval | AC: 5212
- Caption Bottleneck Models | AC: 5814
- Capturing Spectral and Spatial Patterns for Federated Remote Sensing Segmentation | AC: 5810
- CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models | AC: 4809
- CaRe: Critical Parameter Rectification for Efficient Visual Modeling | AC: 5148
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion | AC: 5264
- Cast and Attached Shadow Detection via Iterative Light and Geometry Reasoning | AC: 5072
- Causal Yet Future-Aware: Dual-Path Temporal Modeling for Online Action Segmentation | AC: 5567
- CausalDrive: Real-time Causal World Models for Autonomous Driving | AC: 3753
- CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation | AC: 5969
- CerDETR: Cell-Prior Empowered DETR for Cervical Lesion Detection | AC: 3977
- CGCE: Classifier-Guided Concept Erasure in Generative Models | AC: 4668
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens | AC: 4447
- ChronoFlow Policy: Unifying Past-Future Interaction Flow in Visuomotor Policy Learning | AC: 3245
- ChronusOmni: Improving Time Awareness of Omni-Modal Large Language Models | AC: 4007
- CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains | AC: 5941
- Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding | AC: 4463
- CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection | AC: 3905
- CL4D: Contrastive Language–4D Pretraining for Vision-Language Reasoning in Dynamic Scenes | AC: 4532
- ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement | AC: 3992
- CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition | AC: 5438
- CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery | AC: 3386
- Clue Matters: Empower Video Reasoning with Brain-Inspired Latent Clue Learning | AC: 4320
- CLUE-VAD: Structured Semantic Clues for Understanding Explainable Events in Video Anomaly Detection | AC: 4744
- ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation | AC: 3740
- CMDS-AD: Cross-Modal Dual-Stream Decoupling for Few-Shot Anomaly Detection | arXiv: 2606.20300 | AC: 3629
- CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization | AC: 5960
- CoDePose: Multi-View 3D Human Pose Estimation via Coupled 2D-3D Denoising Diffusion | AC: 5550
- CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion | arXiv: 2606.30030 | AC: 3247
- CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts | AC: 4776
- CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action Models | AC: 3708
- Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning | AC: 3558
- COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models | AC: 5394
- ComplexMimic: Human–Scene Interaction Imitation in Complex 3D Environments | AC: 3457
- Comprehensive language–image pre-training for 3D medical image understanding | AC: 5833
- Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs | AC: 4499
- Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation | AC: 4207
- Condensing Large-Scale Datasets Directly with Minimal Information Loss | arXiv: 2607.00916 | AC: 5756
- ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation | AC: 3892
- ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning | AC: 4462
- Consistent Video-to-Video Translation via Explicit Correspondences | AC: 3572
- Constrained Rotation Optimization: Revisiting Crop-Based Gaze Estimation | AC: 4912
- Context-Aware Joint Alignment for Cross-Scene Hyperspectral Image Classification | AC: 4764
- Continuous Speculative Decoding for Autoregressive Image Generation | arXiv: 2411.11925 | AC: 5776
- ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control | AC: 3535
- Control-DINO: Feature Space Conditioning for Controllable Video Diffusion | AC: 4902
- ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering | AC: 5370
- Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints | arXiv: 2603.11755 | AC: 3544
- Controlling Motion Transfer in Diffusion Transformers via Attention Heads | AC: 3621
- Conversational Human Audio-visual Talking Dialogue Generation | AC: 4941
- CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization | arXiv: 2606.31219 | AC: 4574
- CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection | arXiv: 2510.14792 | AC: 3265
- CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning | AC: 5196
- CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport | AC: 5259
- CRD-Net: Frequency-Adaptive Feature Injection and Change Decoupling for Building Damage Assessment | AC: 5684
- CTEPM: Continuous-Time Event Process Memory for Long-Video Language Models | AC: 5848
- CulinaryCut: A Physics-aware Vision-Language-Action Benchmark for Food Cutting via Material Point Method | AC: 5439
- CURE: Cumulative Knowledge Reuse for Efficient Device-Server Hybrid Inference in Vision-Language Models | AC: 5036
- Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning | arXiv: 2606.22394 | AC: 5545
- Curvature-Guided Mixing for MLLM Adaptation | arXiv: 2606.24963 | AC: 5041
- CUST : Clustered Unit-level Similarity Transformer for Lightweight Image Super-Resolution | AC: 4220
- CustomX: Unified Character, Action, and Scene Customization in Video World Models | arXiv: 2512.17796 | AC: 4020
- Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency | AC: 5343
- DARL: Efficient Document-to-Markup Generation via Look-Ahead Diffusion Trajectory Sampling | AC: 5755
- DART: Deformable Adaptive Reasoning with Temporal Queries for Online Skeleton-Based Action Recognition | AC: 3850
- DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression | AC: 4843
- Data Circuit Breaker: Identifying Training, Test, and Generated Data in Image Generative Models | arXiv: 2606.23872 | AC: 3593
- DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space | AC: 4727
- DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation | AC: 4282
- DE2TR: Dual Evidence Detection Transformer for Video Temporal Grounding | AC: 5412
- Decoding Multimodal Causality: End-to-End Multimodal Mediation Pathways Inference | AC: 3479
- DeCoFlow: Structural Decomposition of Normalizing Flows for Continual Anomaly Detection | AC: 4729
- Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces | arXiv: 2603.14354 | AC: 4059
- DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering | arXiv: 2602.19323 | AC: 4267
- Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models | AC: 4930
- DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking | AC: 3443
- Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation | arXiv: 2512.20117 | AC: 4151
- Denoised Variance-Based Pruning with Optimal Brain Bias Compensation | AC: 3650
- Denoising the Deep Sky: Physics-Based CCD Noise Formation for Astronomical Imaging | AC: 3316
- Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation | arXiv: 2606.21956 | AC: 5434
- Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views | arXiv: 2606.23557 | AC: 5231
- Dense Video Understanding with Inter-tokenization Acceleration | AC: 3614
- Depth-guided Multi-view Exposure Bracketing for HDR Robot Vision | AC: 5101
- DepWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors | arXiv: 2607.00889 | AC: 3937
- DeRA: Decoupled Representation Alignment for Video Tokenization | AC: 3946
- DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models | AC: 3456
- Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams | AC: 3381
- DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues | arXiv: 2606.26602 | AC: 4297
- DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models | AC: 3667
- Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution | arXiv: 2606.22314 | AC: 5647
- Diffusion-Based Material Regularization for Physics-Based Inverse Rendering | arXiv: 2606.31065 | AC: 3441
- DiNBV-Grasp: Real-Time Distance-Aware Two-Stage Next-Best-View for Robotic Grasping | AC: 4360
- DisentangledTMR: Privacy-Preserving Skeleton Motion Retargeting via Factorized Transformers | AC: 5993
- Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning | arXiv: 2606.25488 | AC: 4400
- Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation | arXiv: 2606.20196 | AC: 3773
- DIVA: Instruction-Aware Vision Token Pruning via Dual-Probe Attention Discrepancy | AC: 4758
- DiverseAD: A Large-Scale Driving Dataset with Diverse Atmospheric Conditions | AC: 4645
- Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models | arXiv: 2605.01896 | AC: 5089
- DLGStream: Dynamic Language-embedded Guassian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming | arXiv: 2606.28840 | AC: 4664
- DnA: Denoising Attention for Visual Tasks | AC: 3828
- DOGE: Differentiable Bézier Graph Optimization for Road Network Extraction | AC: 4405
- Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data | AC: 5440
- Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts | arXiv: 2607.00666 | AC: 4444
- Don’t Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance | arXiv: 2606.27371 | AC: 3156
- Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement | AC: 4878
- DP-BOA: Dirichlet-Process Birth-or-Assign for On-the-Fly Category Discovery | AC: 4517
- Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off | arXiv: 2603.22607 | AC: 4939
- DriftScope: Measuring The Hidden Effects of Diffusion Model Fine-Tuning | arXiv: 2607.00183 | AC: 3510
- Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout | arXiv: 2605.05092 | AC: 3775
- DriveVA: Video Action Models are Zero-Shot Drivers | arXiv: 2604.04198 | AC: 3252
- DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation | arXiv: 2606.31918 | AC: 3336
- DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution | arXiv: 2606.29198 | AC: 5665
- Dual-End Consistency Model | arXiv: 2602.10764 | AC: 5288
- Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution | arXiv: 2606.26716 | AC: 3760
- Dynamic Inverse Rendering for Enhanced Material-Lighting Decomposition | AC: 5513
- E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation | arXiv: 2606.27268 | AC: 4867
- E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes | arXiv: 2604.04834 | AC: 4840
- EAGS: Error-Aware Gaussian Splatting with Dual-Confidence-Guided Modeling for Uncalibrated Driving Scenes | AC: 4526
- EatVid-Bench: A Multimodal Fine-Grained Eating Behavior Video Dataset | AC: 5037
- ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation | AC: 3709
- EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics | AC: 3747
- Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing | arXiv: 2603.03143 | AC: 3702
- Efficient Document Tampering Localization with Multi-Level Discrepancy Features and Unified DCT–Quantization Embedding | arXiv: 2606.22285 | AC: 5506
- Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion | arXiv: 2606.30215 | AC: 4023
- Egocentric Procedure Parsing | AC: 4540
- EgoEverything: A Benchmark for Human Behavior–Inspired Long-Context Egocentric Video Understanding in AR Environment | AC: 4529
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding | arXiv: 2510.26113 | AC: 3203
- EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage | AC: 3471
- EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding | arXiv: 2606.24422 | AC: 4403
- EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning | arXiv: 2511.18242 | AC: 4523
- ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval | arXiv: 2606.20280 | AC: 5615
- Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding | AC: 3478
- Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings | AC: 3466
- EMOTE: Expressive Motion and Shape Disentanglement for Human Animation | arXiv: 2606.28026 | AC: 3741
- EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models | AC: 4214
- Enhancing Embodied Reasoning and Grounding by Novel View Synthesis | AC: 4036
- Entropy-Controlled Flow Matching | arXiv: 2602.22265 | AC: 3810
- EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal | arXiv: 2512.21545 | AC: 5267
- ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis | arXiv: 2603.25168 | AC: 5878
- Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs | arXiv: 2606.20177 | AC: 3870
- Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations | arXiv: 2606.24716 | AC: 5939
- Event-Driven Video Generation | arXiv: 2603.13402 | AC: 5108
- Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks | arXiv: 2603.11689 | AC: 5715
- Exploiting Local Flatness for Efficient Out-of-Distribution Detection | arXiv: 2606.29952 | AC: 5630
- Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation | AC: 5803
- ExPLoRe: Expert Patch-Level Loss Routing for Multi-Objective Masked Image Modeling | arXiv: 2606.31201 | AC: 5992
- ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving | arXiv: 2604.02714 | AC: 5991
- Exploring Efficient Reasoning Segmentation with Small Language Models | AC: 4052
- Fabric Image Demoiréing Benchmark from Synthesis to Restoration | arXiv: 2606.24072 | AC: 3616
- Face Anything: 4D Face Reconstruction from Any Image Sequence | arXiv: 2604.19702 | AC: 4482
- FaceMoE: Mixture of Experts for Low-Resolution Face Recognition | arXiv: 2606.32040 | AC: 3339
- FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis | AC: 5679
- FD²: A Dedicated Framework for Fine-Grained Dataset Distillation | arXiv: 2603.25144 | AC: 3742
- FedLAS: Feature-Modulated Bidirectional Label Smoothing for Neural Network Calibration | arXiv: 2606.28654 | AC: 3206
- FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs | arXiv: 2606.22875 | AC: 3888
- FeVOS: Foresight Expression Video Object Segmentation | arXiv: 2606.25585 | AC: 5598
- Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples | arXiv: 2509.25682 | AC: 5926
- Fidelity- and Perception-Aware Local Implicit Attention for Arbitrary-Scale Image Super-Resolution | arXiv: 2606.21910 | AC: 3972
- Filterless Snapshot Hyperspectral Imaging using Guided Patch Diffusion | arXiv: 2412.02798 | AC: 3834
- Fisher-Routed Mixture of Experts for Federated Class-Incremental Learning | AC: 4485
- FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control | AC: 4328
- FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction | arXiv: 2606.21373 | AC: 4753
- FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation | arXiv: 2606.22424 | AC: 4005
- FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation | arXiv: 2511.21029 | AC: 4239
- FMA-Net++: Motion- and Exposure-Aware Joint Video Super-Resolution and Deblurring | arXiv: 2512.04390 | AC: 5366
- Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE | arXiv: 2606.26938 | AC: 3647
- Following the Flow: Advection-Consistent Modeling for Event-based Small Object Detection | arXiv: 2606.22378 | AC: 5799
- FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution | arXiv: 2606.28745 | AC: 4909
- From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting | AC: 4354
- From Gaze to Meaning: An AI Agent for Unified Zero-Shot Grounding and Explanation | AC: 5517
- From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP | arXiv: 2606.26535 | AC: 3720
- From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs | arXiv: 2606.29461 | AC: 3547
- From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation | arXiv: 2606.23486 | AC: 5450
- FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model | arXiv: 2606.20110 | AC: 3978
- FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal | arXiv: 2603.19036 | AC: 4886
- G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation | arXiv: 2601.03510 | AC: 5646
- GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models | arXiv: 2601.18197 | AC: 5058
- Gaussian Belief Propagation Network for Depth Completion | arXiv: 2601.21291 | AC: 3367
- GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation | arXiv: 2603.26661 | AC: 4503
- GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine | AC: 5462
- GEM: Generative Supervision Helps Embodied Intelligence | AC: 3349
- Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition | arXiv: 2606.22416 | AC: 3293
- GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence | arXiv: 2511.21945 | AC: 3508
- Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval | arXiv: 2602.00813 | AC: 4117
- Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding | AC: 3262
- Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior | arXiv: 2606.31814 | AC: 4869
- GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models | arXiv: 2607.00492 | AC: 3856
- Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics | arXiv: 2603.13859 | AC: 3513
- GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces | AC: 3259
- GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising | arXiv: 2606.30003 | AC: 4817
- Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning | arXiv: 2606.26973 | AC: 3557
- Geometry-Anchored Transport Framework for Exemplar-Free Class-Incremental Learning | arXiv: 2606.25347 | AC: 4569
- Geometry-Aware Style Transfer in 3D Gaussian Splatting | arXiv: 2606.24144 | AC: 4174
- Geometry-Preserving in 3D Gaussian Splatting for LiDAR-Camera Extrinsic Calibration | arXiv: 2606.20103 | AC: 4656
- GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis | arXiv: 2603.14965 | AC: 3876
- GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation | AC: 3217
- GKDT: General Keypoint Detection Transformer | arXiv: 2607.00752 | AC: 3679
- Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing | AC: 3375
- GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction | arXiv: 2604.19624 | AC: 3270
- Graph Coloring for Multi-Task Learning | arXiv: 2509.16959 | AC: 3425
- Grasp-Oriented Non-Prehensile Manipulation via Learning a Graspability Field | AC: 3490
- GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition | arXiv: 2602.03370 | AC: 3955
- GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation | arXiv: 2603.26266 | AC: 3244
- Guiding the Blind: Generalizing GUI Agents to Unseen Websites via Multimodal Tutorials | AC: 4245
- H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks | arXiv: 2606.25578 | AC: 5686
- HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving | AC: 3249
- HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration | arXiv: 2606.28215 | AC: 5188
- HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking | arXiv: 2607.00494 | AC: 4786
- Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation | arXiv: 2606.31071 | AC: 5648
- Hierarchical Spatial and Channel Aggregation for Cross-domain Few-shot Segmentation | arXiv: 2606.24296 | AC: 5612
- HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training | arXiv: 2606.20189 | AC: 3430
- HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning | AC: 5381
- Histogram-constrained Image Generation | arXiv: 2606.31683 | AC: 4232
- Histopathology Multi-modal Embedding for Pathology Composed Retrieval | arXiv: 2502.07221 | AC: 5623
- Horizon3D: Sparse Radar-Camera Fusion for Long-Range 3D Perception in Autonomous Driving | arXiv: 2606.31096 | AC: 5928
- HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding | arXiv: 2602.12957 | AC: 5282
- HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection | arXiv: 2606.31172 | AC: 5119
- HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human–Scene Interactions | AC: 3313
- HuCollisionField: Resolving Self-Collisions via Neural Fields for Human Prediction | arXiv: 2606.29686 | AC: 3864
- Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum | AC: 3751
- Hybrid Event–Frame Sensors: Modeling, Calibration, and Simulation | arXiv: 2511.18037 | AC: 3755
- HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding | arXiv: 2607.00428 | AC: 5643
- HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization | arXiv: 2604.20328 | AC: 5084
- Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting | arXiv: 2607.00159 | AC: 5899
- Image Warping for Image-to-Image Translation | arXiv: 2606.31018 | AC: 5130
- Improving Adversarial Robustness via Activation Amplification and Attenuation | arXiv: 2606.27784 | AC: 3522
- Improving Sparse-View 3DGS Generalization via Flat Minima Optimization | arXiv: 2607.00885 | AC: 5368
- In-context Region-based Drag: Drag Any Region to Any Shape | arXiv: 2606.25907 | AC: 3552
- Information-Regularized Attention for Visual-Centric Reasoning | arXiv: 2607.00434 | AC: 5530
- Interact3D: Compositional 3D Generation of Interactive Objects | arXiv: 2603.16085 | AC: 3861
- InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars | AC: 3402
- InterEdit: Navigating Text-Guided Multi-Human 3D Motion Editing | arXiv: 2603.13082 | AC: 3603
- Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment | arXiv: 2606.30262 | AC: 5537
- Intrinsically Stable Spiking Neural Networks: Overcoming the Performance Barrier in the Absence of Batch Normalization | arXiv: 2606.31695 | AC: 4371
- IREU: Identity-Related Encoder-Only Unlearning for Customized Portrait Generation | arXiv: 2606.29880 | AC: 4120
- ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation | arXiv: 2505.20935 | AC: 5373
- JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising | arXiv: 2606.20563 | AC: 4414
- KATANA: Knowledge-Aligned Topology-Aware Neural Agents for RL-Driven Vision-Language Model Compression | AC: 5897
- Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding | AC: 5468
- Kirin: Animal Motion Generation from In-the-Wild Video | AC: 3829
- Knowledge-Centric Agents for Workflow Generation in ComfyUI | AC: 4258
- LaGen: Towards Autoregressive LiDAR Scene Generation | arXiv: 2511.21256 | AC: 4888
- Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures | arXiv: 2605.04035 | AC: 3643
- Latent Visual Diffusion Reasoning with Monte Carlo Tree Search | arXiv: 2606.27988 | AC: 3822
- LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion | arXiv: 2603.14526 | AC: 3581
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking | arXiv: 2509.12046 | AC: 3209
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection | arXiv: 2512.05663 | AC: 3882
- Learn Once, Edit Anywhere: Visual Direction Transfer for Diffusion Models | arXiv: 2403.19645 | AC: 3264
- Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision | AC: 3825
- Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents | arXiv: 2606.31270 | AC: 5478
- Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding | AC: 4771
- Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval | arXiv: 2607.00374 | AC: 4892
- Learning to Deny: Action Denial in Multimodal Large Language Models | arXiv: 2606.31187 | AC: 3915
- Learning Transferable Dynamics Priors from Action to World Modeling | arXiv: 2606.29501 | AC: 3656
- Learning Video Dynamics with Predictive Differentiable Rendering | arXiv: 2606.31050 | AC: 4642
- LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models | arXiv: 2606.23686 | AC: 4895
- LibraGen: Playing a Balance Game in Subject-Driven Video Generation | AC: 3221
- LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement | arXiv: 2606.23539 | AC: 4839
- LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing | arXiv: 2606.26740 | AC: 4792
- LogicIR: Logic Gate Networks for Image Restoration | arXiv: 2606.26609 | AC: 3151
- LogiCo: A Unified Framework for Logical and Structural Anomaly Detection | arXiv: 2606.28688 | AC: 3604
- Long-term Traffic Simulation via Structured Autoregressive Modeling | arXiv: 2606.31209 | AC: 5085
- Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition | arXiv: 2607.00090 | AC: 3660
- LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation | arXiv: 2509.17773 | AC: 3999
- LUNA: Learning Universal 3D Human Animation Beyond Skinning | arXiv: 2606.31981 | AC: 4925
- M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for optical-SAR Object Detection | arXiv: 2505.10931 | AC: 3158
- Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs | arXiv: 2602.05275 | AC: 5057
- MambaRaw: Selective State Space Modeling for Efficient 4K RAW Image Reconstruction | arXiv: 2606.24479 | AC: 3408
- MASS: Motion-Aligned Selective Scan for Flow-Based Video Frame Interpolation | arXiv: 2606.27718 | AC: 3303
- Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras | arXiv: 2604.18744 | AC: 4933
- MATCH: Flow Matching for Multi-View Anomaly Detection | arXiv: 2606.24375 | AC: 4872
- Matryoshka Gaussian Splatting | AC: 3234
- MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction | arXiv: 2604.01958 | AC: 4385
- MedCAGD: Context-Aware Gated Decoder for Robust Medical Image Segmentation | arXiv: 2607.00409 | AC: 4399
- MeGAS: Thermomechanical Dynamic Gaussian Splatting for Thermophysical Scene Editing | arXiv: 2606.23455 | AC: 3980
- MemLearner: Learning to Query Context Memory for Video World Models | arXiv: 2606.31734 | AC: 3900
- MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization | arXiv: 2607.00902 | AC: 4624
- MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation | arXiv: 2606.26016 | AC: 3631
- MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations | arXiv: 2606.27779 | AC: 3789
- MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein | arXiv: 2606.29462 | AC: 4670
- MirrorPPR: Exemplar-Based Portrait Photo Retouching | arXiv: 2606.29308 | AC: 5694
- MixTTA: Low-Rank Cross-Channel Mixing for Reliable Test-Time Adaptation | arXiv: 2606.28142 | AC: 4612
- MLVC: A Multi-platform Learned Video Codec for Real-World Deployment | arXiv: 2606.28027 | AC: 4404
- MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation | arXiv: 2604.19679 | AC: 3470
- MobileManiBench: Simplifying Model Verification for Mobile Manipulation | arXiv: 2602.05233 | AC: 4029
- ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation | arXiv: 2604.07021 | AC: 4138
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance | AC: 4484
- Moiré Video Authentication: A Physical Signature Against AI Video Generation | arXiv: 2604.01654 | AC: 3761
- MolmoWeb: Open Visual Web Agent and Open Data for the Open Web | AC: 4684
- Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading | arXiv: 2606.28144 | AC: 5324
- MonoSR: Open-Vocabulary Spatial Reasoning on Monocular Images | arXiv: 2511.19119 | AC: 4275
- Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting | arXiv: 2606.30017 | AC: 3145
- MotionAtlas: A High-Quality Dataset and Benchmark for Dense Motion Captioning | arXiv: 2606.29531 | AC: 4897
- MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment | arXiv: 2607.00858 | AC: 5408
- Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction | arXiv: 2606.26812 | AC: 3254
- Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments | arXiv: 2607.00457 | AC: 5948
- Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning | arXiv: 2606.29334 | AC: 3804
- Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts | arXiv: 2607.00371 | AC: 5141
- Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation | arXiv: 2606.22197 | AC: 3218
- MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving | AC: 5547
- NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation | arXiv: 2606.29395 | AC: 4846
- NavWM: A Unified Navigation World Model for Foresight-Driven Planning | arXiv: 2606.24101 | AC: 3973
- NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models | arXiv: 2606.22537 | AC: 3940
- Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating | arXiv: 2603.12598 | AC: 4147
- Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors | arXiv: 2603.15129 | AC: 3536
- Nexus-Vid: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models | arXiv: 2605.31603 | AC: 5094
- NGPS: Structure-Preserving Self-Supervised Denoising via Neighbor-Guided Patch Sampling | arXiv: 2606.23200 | AC: 4200
- No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs | arXiv: 2606.31933 | AC: 5290
- NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment | arXiv: 2606.18066 | AC: 3859
- Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold | arXiv: 2607.00647 | AC: 4934
- NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics | arXiv: 2606.31764 | AC: 4225
- Obliviate: Erasing Concepts from Autoregressive Image Generation Models | arXiv: 2606.28643 | AC: 4923
- OCTOPUS: Multi‑Agentic Universal Compositional Visual Retrieval | AC: 5802
- Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion | arXiv: 2606.21135 | AC: 3176
- OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning | AC: 3376
- OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data | AC: 4238
- OmniNWM: Unifying the State-Action-Reward Triad for Closed-Loop Panoramic Driving Navigation World Models | arXiv: 2510.18313 | AC: 3205
- OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning | AC: 3451
- OmniX: Any-view and Any-time 4D reconstruction via Feed-forward Trajectory Fields | AC: 3323
- On Test-Time Scaling for Vision-Language Models | arXiv: 2606.28864 | AC: 4216
- On the Faithfulness of Post-Hoc Concept Bottleneck Models | arXiv: 2606.30498 | AC: 5447
- On the Vulnerability of Parameter-Level Defenses to Model Merging | arXiv: 2606.30360 | AC: 3680
- One Video, One World: Turning Monocular Video into Physical 4D Scenes | arXiv: 2606.31388 | AC: 3356
- OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization | arXiv: 2607.00289 | AC: 3903
- OP3DSG: Open-vocabulary Part-aware 3D Scene Graph Generation for Real-world Environments | arXiv: 2606.29786 | AC: 4307
- Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints | arXiv: 2606.24353 | AC: 3266
- OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos | arXiv: 2606.25245 | AC: 4440
- OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models | arXiv: 2606.31026 | AC: 5040
- P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling | AC: 5238
- PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates | arXiv: 2512.06845 | AC: 4682
- Pano3D: Unified 3D Reconstruction and Panoptic Segmentation | arXiv: 2606.14307 | AC: 4920
- PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding | arXiv: 2512.20907 | AC: 4323
- Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models | arXiv: 2606.27373 | AC: 4470
- Perceptual Projection Pruning: Diversity-Aware Video Token Pruning for Multimodal Large Language Models | AC: 5817
- Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising | arXiv: 2607.00407 | AC: 5977
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer | arXiv: 2511.19778 | AC: 4089
- PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models | AC: 3450
- PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing | arXiv: 2606.26551 | AC: 5047
- PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation | arXiv: 2512.24551 | AC: 4063
- Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch | arXiv: 2604.09100 | AC: 4773
- Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation | arXiv: 2606.25306 | AC: 5501
- PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation | arXiv: 2606.26916 | AC: 5129
- PIAvatar: Physically Interactive Avatars via Deformation Gradient Decoupling | arXiv: 2606.21162 | AC: 4193
- PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation | AC: 3174
- PLOT: Pseudo-Labeling via Object Tracking for Monocular 3D Object Detection | arXiv: 2507.02393 | AC: 4690
- Pointer-CAD v2: Plan-Then-Construct CAD Generation with Dimension-Aware Parametric Precision | arXiv: 2606.29301 | AC: 3855
- PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models | arXiv: 2606.22540 | AC: 3170
- Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation | arXiv: 2606.29908 | AC: 5211
- Pose Anything Anywhere: Model-free Object Poses from Arbitrary References | arXiv: 2606.23634 | AC: 4278
- Prevention over Correction: Learning Aligned Representations in One-shot Federated Learning | AC: 4293
- PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving | arXiv: 2606.31830 | AC: 4451
- ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models | arXiv: 2603.19466 | AC: 4434
- Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video | arXiv: 2607.00157 | AC: 4740
- Prompt2Effect: Training-Free LoRA Synthesis for Controllable Video Effects | arXiv: 2606.13971 | AC: 3453
- ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering | AC: 5675
- ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs | arXiv: 2605.01971 | AC: 5957
- PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking | arXiv: 2606.30476 | AC: 3664
- QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception | arXiv: 2509.03704 | AC: 3837
- RAGA: Real Time Ray Traced Gaussian Shadow Casting for 3DGS Avatar-Scene Interaction | arXiv: 2606.29329 | AC: 3220
- Rank-Aware Hyperbolic Alignment for Vision–Language Dataset Distillation | arXiv: 2606.29464 | AC: 4352
- RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation | arXiv: 2606.22749 | AC: 5166
- RBE-Flow:Recurrent Bayesian Estimation on Feature Manifolds for Cross-Modal Registration | arXiv: 2606.30492 | AC: 4906
- Real-Time Source-Free Object Detection | arXiv: 2606.31834 | AC: 4975
- RefAlign: Representation Alignment for Reference-to-Video Generation | arXiv: 2603.25743 | AC: 3790
- Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation | arXiv: 2506.23102 | AC: 5301
- Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning | AC: 3298
- Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering | arXiv: 2505.17338 | AC: 4144
- RePer-360: Releasing Perspective Priors for 360° Depth Estimation via Self-Modulation | arXiv: 2603.05999 | AC: 3917
- Repurposing Geometric Foundation Models for Multi-view Diffusion | AC: 3136
- ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision–Language Models | arXiv: 2607.00361 | AC: 5876
- Residual-Guided Expert Specialization for Incomplete Multimodal Learning | arXiv: 2606.30355 | AC: 3678
- ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration | arXiv: 2606.26769 | AC: 4157
- RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception | arXiv: 2606.31895 | AC: 4555
- Rethinking Continual Anomaly Detection on the Edge: Benchmarking Under Realistic Industrial Conditions | arXiv: 2605.24251 | AC: 3516
- Rethinking Garment Conditioning in Diffusion-based Virtual Try-On: Decouple, Don't Denoise | arXiv: 2511.18775 | AC: 4096
- Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection | arXiv: 2606.23069 | AC: 5171
- Rethinking Training and Inference for Trajectory Forecasting: Linking Winner-Take-All back to GMMs | arXiv: 2606.26424 | AC: 4060
- Revisiting Autoregressive Models for Generative Image Classification | arXiv: 2603.19122 | AC: 3984
- Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need | AC: 3431
- Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation | arXiv: 2606.31382 | AC: 4854
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios | arXiv: 2511.18011 | AC: 5365
- Robust Zero-shot Anomaly Detection under Limited Auxiliary Anomaly Priors | arXiv: 2606.29428 | AC: 5411
- RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning | arXiv: 2606.28266 | AC: 3916
- S2-FracMix: Self-Saliency Fractal Mixup | arXiv: 2606.25784 | AC: 3462
- SALT: Self-Consistent Distribution Matching with Cache-Aware Training for Few-Step Video Generation | arXiv: 2604.03118 | AC: 3473
- SAM2Matting: Generalized Image and Video Matting | arXiv: 2606.27339 | AC: 4966
- SARIF: Segment Anything for Robust Image Forensics | arXiv: 2606.21108 | AC: 5314
- ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision–Language Models | arXiv: 2606.29579 | AC: 4961
- SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation | arXiv: 2606.30124 | AC: 5783
- See & Sniff: Learning Visuo-Olfactory Representations | arXiv: 2606.27307 | AC: 4376
- Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation | arXiv: 2606.29941 | AC: 4435
- Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation | AC: 3331
- Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing | arXiv: 2607.00124 | AC: 5741
- Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning | AC: 5852
- Self-supervised Garment Dynamics with Persistent Wrinkles | arXiv: 2606.25065 | AC: 3705
- Semantic Browsing: Controllable Diversity for Image Generation | arXiv: 2606.23679 | AC: 4450
- SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models | arXiv: 2606.27444 | AC: 4801
- SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking | arXiv: 2606.24449 | AC: 4693
- SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions | arXiv: 2607.00369 | AC: 4723
- ShellMaker: Language-Guided Exterior Completion under Structural Constraints | arXiv: 2606.31680 | AC: 3519
- SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset | arXiv: 2606.30001 | AC: 4963
- SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models | arXiv: 2606.27741 | AC: 5178
- SIGNER: Temporally Grounded Sign Language Generation via Time-Resolved Conditioning | arXiv: 2506.07460 | AC: 4313
- SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation | arXiv: 2606.28626 | AC: 4999
- SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks | arXiv: 2606.24361 | AC: 3533
- Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs | arXiv: 2604.20937 | AC: 5652
- SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation | arXiv: 2603.14152 | AC: 3548
- SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery | arXiv: 2511.20157 | AC: 3845
- SkelEM: Explicit Decoupling of Topology and Details for Self-supervised Axial Super-Resolution in Volume Microscopy | arXiv: 2606.30012 | AC: 4327
- Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models | arXiv: 2511.14900 | AC: 4291
- SlowBA: An efficiency backdoor attack towards VLM-based GUI agents | arXiv: 2603.08316 | AC: 5078
- SMART: When is it Actually Worth Expanding a Speculative Tree? | arXiv: 2604.09731 | AC: 4155
- Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective | arXiv: 2512.10244 | AC: 4085
- SONIC: Spectral Optimization of Noise for Inpainting with Consistency | arXiv: 2511.19985 | AC: 5954
- Spanning the Visual Analogy Space with a Weight Basis of LoRAs | arXiv: 2602.15727 | AC: 4237
- Sparsity-Inducing Divergence Losses for Biometric Verification | arXiv: 2606.31664 | AC: 5481
- SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning | AC: 5289
- SPEAR: A Simulator for Photorealistic Embodied AI Research | AC: 3273
- SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation | arXiv: 2607.00525 | AC: 5880
- Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution | arXiv: 2603.06275 | AC: 5583
- Spectral Evolution-Guided Token Pruning in Large Multimodal Models | arXiv: 2606.24165 | AC: 3609
- Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations | arXiv: 2606.23129 | AC: 4373
- SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision | arXiv: 2603.24036 | AC: 3727
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models | arXiv: 2510.12784 | AC: 4190
- StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production–Living Simulations with Stardew Valley | arXiv: 2507.07445 | AC: 5361
- Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging | AC: 4368
- Staying VIGILant: Mitigating Visual Laziness in MLLMs via Information-Theoretic Alignment | arXiv: 2606.26387 | AC: 4110
- Steerable Vision Transformers | arXiv: 2604.02327 | AC: 5945
- Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance | arXiv: 2506.20995 | AC: 4191
- StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors | arXiv: 2606.30545 | AC: 3379
- StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning | arXiv: 2607.00465 | AC: 5503
- StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation | arXiv: 2605.21466 | AC: 3345
- Streaming Dense Voxel Representations for 3D Occupancy Prediction | arXiv: 2503.22087 | AC: 3241
- Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space | arXiv: 2606.21705 | AC: 4066
- Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers | arXiv: 2606.22383 | AC: 5700
- Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding | arXiv: 2604.07753 | AC: 5319
- Syn4D: A Multiview Synthetic 4D Dataset | AC: 3248
- SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation | arXiv: 2606.30849 | AC: 3830
- Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction | arXiv: 2510.03117 | AC: 3459
- TaskTok: Delving into Task Tokens for Task-driven Image Restoration | arXiv: 2606.26615 | AC: 3148
- TaxoMIL: Taxonomy-Constrained Learning for Hierarchical Whole Slide Image Analysis | arXiv: 2606.31100 | AC: 5220
- Tesselating The Earth | arXiv: 2606.27514 | AC: 5452
- Test Time Training for Long Videos via Frame Forgetting Network | arXiv: 2606.26515 | AC: 4271
- Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Generation | arXiv: 2606.22726 | AC: 5082
- Text-Conditioned Background Generation for Editable Multi-Layer Documents | arXiv: 2512.17151 | AC: 3585
- TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts | arXiv: 2606.28077 | AC: 4871
- The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models | arXiv: 2607.00402 | AC: 5377
- The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning | arXiv: 2606.30875 | AC: 3333
- There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion | arXiv: 2606.27905 | AC: 5062
- Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs | arXiv: 2606.31471 | AC: 5278
- TIIF-Bench: How Does Your T2I Model Follow Your Instructions? | AC: 3177
- TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration | AC: 5122
- TORA: Topological Representation Alignment for 3D Shape Assembly | arXiv: 2604.04050 | AC: 3567
- Toward Robust In-Context Segmentation via Concept Guidance | arXiv: 2606.28149 | AC: 5254
- Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning | arXiv: 2511.20196 | AC: 4358
- Towards Consistent and Efficient Dataset Distillation via Diffusion-Driven Selection | AC: 5321
- Towards Generalizable Robotic Manipulation in Dynamic Environments | arXiv: 2603.15620 | AC: 3319
- Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation | arXiv: 2606.30598 | AC: 3943
- Towards Interactive Global Geolocation Assistant | AC: 3263
- Towards Long-Form Spatio-Temporal Video Grounding | arXiv: 2602.23294 | AC: 4321
- Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption | AC: 4508
- Towards Metric-Agnostic Trajectory Forecasting | arXiv: 2607.01133 | AC: 5818
- Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics | arXiv: 2512.13660 | AC: 3404
- Training-free Cross-domain Few-shot Segmentation via Robust Semantic Representation and Matching | AC: 5609
- Training-Free Refinement of Flow Matching with Divergence-based Sampling | AC: 3225
- Transferability Between Understanding and Generation in Unified Multimodal Models | AC: 3137
- Triangular Consistency as a Universal Constraint for Learning Optical Flow | arXiv: 2606.19938 | AC: 3812
- TriMotion: Modality-Agnostic Camera Control for Video Generation | AC: 3398
- Trustworthy Image Authentication using Forensic Knowledge Graphs | arXiv: 2606.23917 | AC: 3521
- UEval: A Benchmark for Unified Multimodal Generation | AC: 3346
- UHD-MFF: Shattering Barriers in Multi-Focus Ultra-High-Definition Image Fusion via Learnable Lookup Tables | arXiv: 2606.31242 | AC: 5384
- Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics | arXiv: 2606.27554 | AC: 3563
- UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving | arXiv: 2601.04453 | AC: 4081
- UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation | arXiv: 2603.22282 | AC: 3405
- UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer | arXiv: 2512.21078 | AC: 4134
- UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation | arXiv: 2606.31451 | AC: 5005
- UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving | arXiv: 2606.25736 | AC: 4186
- UniTranslator: A Unified Multi-modal framework for End-to-end In-Image Machine Translation | arXiv: 2606.24333 | AC: 4127
- UniTriSplat: A Unified 3D Gaussian Splatting Framework with Uniform Spherical Rasterization for Universal Cameras | arXiv: 2606.29794 | AC: 4818
- Universal Image Immunization against Diffusion-based Image Editing via Semantic Injection | arXiv: 2602.14679 | AC: 3278
- Unveiling Transferability in Trajectory Prediction via Latent Scene Embeddings | arXiv: 2606.30777 | AC: 5485
- URoPE: Universal Relative Position Embedding across Geometric Spaces | arXiv: 2604.18747 | AC: 5575
- VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos | arXiv: 2603.12703 | AC: 3399
- Vector Scaffolding: Inter-Scale Orchestration for Differentiable Image Vectorization | arXiv: 2605.11913 | AC: 5518
- Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously | AC: 3229
- VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement | arXiv: 2607.00446 | AC: 5813
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models | AC: 3165
- ViewSplat: View-Adaptive Dynamic Gaussian Splatting for Feed-Forward Synthesis | arXiv: 2603.25265 | AC: 3759
- ViQ: Text-Aligned Visual Quantized Representations at Any Resolution | arXiv: 2606.27313 | AC: 5257
- VisCritic: Visual State Comparison as Process Reward for GUI Agents | arXiv: 2606.24525 | AC: 4958
- Vision Bridge Transformer at Scale | AC: 3232
- VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning | arXiv: 2603.01195 | AC: 4167
- VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context | arXiv: 2606.30288 | AC: 5979
- Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers | arXiv: 2607.00382 | AC: 3543
- VLA Knows Its Limits | arXiv: 2602.21445 | AC: 4137
- VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors | arXiv: 2510.00458 | AC: 3853
- VOCA: Visual Odometry with Codec Awareness | arXiv: 2607.00189 | AC: 3509
- VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction | arXiv: 2509.19297 | AC: 3169
- VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On | arXiv: 2603.11734 | AC: 3764
- WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching | AC: 3393
- Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs | arXiv: 2607.00302 | AC: 5669
- Walking in the Implicit: Interactive World Exploration via Neural Scene Representation | arXiv: 2606.30045 | AC: 3302
- WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation | AC: 5413
- What if? Emulative Simulation with World Models for Situated Reasoning | arXiv: 2603.06445 | AC: 4721
- When Sinks Help or Hurt: Unified Framework for Attention Sink in MLLMs | arXiv: 2604.03316 | AC: 3694
- WorldAgents: Can Foundation Image Models be Agents for 3D World Models? | AC: 5821
- Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation | arXiv: 2606.30248 | AC: 5650
- Zero-shot Depth from Defocus | arXiv: 2603.26658 | AC: 3798
- Zero-Shot Quantization for Object Detectors using Off-the-Shelf Generative Models | arXiv: 2606.31456 | AC: 4509
- ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting | AC: 3373
- LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory | AC: 3286
- MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding | AC: 3142
- VOID: Video Object and Interaction Deletion | AC: 3342
- VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment | AC: 3197
- FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry | AC: 3189
- TETO: Tracking Events with Teacher Observation for Motion Estimation and Frame Interpolation | AC: 3139
- Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? | AC: 3320
- XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity | AC: 3426
- S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight | AC: 3407
- Human Mesh Modeling for Anny Body | AC: 3422
- UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors | AC: 3418
- RAU: Reference-based Anatomical Understanding with Vision-Language Models | AC: 3449
- TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy | AC: 3423
- Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution | AC: 3370
- DiffProxy: Multi-View Human Mesh Recovery via Diffusion-Generated Dense Proxies | AC: 3382
-
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing | AC: 3458
-
CanoVerse: 3D Object Scalable Canonicalization and Dataset for Generation and Pose | AC: 3600
- Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models | AC: 3559
- FaCT-GS: Fast and Scalable CT Reconstruction with Gaussian Splatting | AC: 3541
- Free-Range Gaussians: Non-Grid-Aligned Generative 3D Gaussian Reconstruction | AC: 3583
- IRIS: Intersection-aware Ray-based Implicit Editable Scenes | AC: 3495
- MemoBench: Benchmarking World Modeling in Dynamically Changing Environments | AC: 3594
- Narrative-Driven Paper-to-Slide Generation via ArcDeck | AC: 3590
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing | AC: 3482
- Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility | AC: 3472
- SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving | AC: 3597
- StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views | AC: 3486
- Surprise Forcing: What to Remember, When to Skip in Long Video Generation | AC: 3534
- SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing | AC: 3530
- Tri-Efficient Transfer Learning for Point Cloud Videos | AC: 3624
- VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward | AC: 3500
- ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images | AC: 3574
- VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations | AC: 3595
- Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning | AC: 3599
- VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model | AC: 3580
- World Reconstruction From Inconsistent Views | AC: 3497
- Scaling Laws for Black-box Adversarial Attacks | AC: 3565
- ECHO: Ego-centric Modeling of Human-Object Interactions | AC: 3503
- From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space | AC: 3460
- Shared LoRA Subspaces for almost Strict Continual Learning | AC: 3502
- T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability | AC: 3505
- TACO-Net: Topological Signatures Triumph in 3D Object Classification | AC: 3582
- HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching | AC: 3494
- NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control | AC: 3499
- DiffPro: Joint Timestep and Layer-Wise Precision Optimization for Efficient Diffusion Inference | AC: 3584
- Learning to Recover Task Experts from a Multi-Task Merged Model | AC: 3618