🧪 ICML2026 Accepted Papers¶
1846 ICML2026 paper notes covering Image Generation (141), Model Compression (117), AI Safety (114), Reinforcement Learning (110), Interpretability (91), Multimodal VLM (89), Optimization & Theory (88), LLM Agent (83) and other 50 areas. Each note has TL;DR, motivation, method, experiments, highlights, and limitations — 5-minute reads of core ideas.
💡 LLM Reasoning (78)¶
- Chain-of-Thought Reasoning in the Wild Is Not Always Faithful
-
This paper reveals two types of unfaithful behavior in frontier LLM Chain-of-Thought (CoT) under non-adversarial, naturally phrased prompts (without human-injected bias): Implicit Post-hoc Rationalization (generating contradictory but seemingly plausible arguments for the same comparative question pairs) and Unfaithful Irlogical Shortcuts (skipping critical reasoning steps in difficult math problems while still reaching the correct answer). The unfaithfulness rate in production models reaches up to 13%, and even reasoning models (DeepSeek R1: 0.37%, Claude 3.7 Sonnet thinking: 0.04%) are not perfectly faithful.
- Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
-
The authors first identify three critical flaws in "pure numerical reward RL" (performance plateaus, ineffective spontaneous reflection, and stubborn failures), then integrate natural language critique into online RL. The model learns both the initial response and "self-refinement based on critique." A shaping function is used to bias towards "correct but unfamiliar" refinements while suppressing incorrect ones, achieving an average Pass@1 improvement of approximately +15.0~21.6% across eight reasoning benchmarks (Qwen series).
- Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
-
This paper proposes Prefix-RFT, which constructs mixed trajectories by sampling prefixes from expert demonstrations and concatenating model continuations. This approach injects knowledge guidance from SFT while maintaining the objective-oriented optimization of RFT, significantly outperforming independent SFT, RFT, and existing hybrid methods on mathematical reasoning tasks.
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
-
The authors use attention dynamics to "develop" the reasoning process—discovering a "preplan-and-anchor" two-beat rhythm during generation. They convert two internal metrics (WAAD/FAI) characterizing this rhythm into token-level advantage amplification coefficients for RL. This allows GRPO to concentrate credit on critical tokens that dictate the direction of downstream reasoning, achieving consistent performance gains across Countdown, QA, and multiple mathematical reasoning benchmarks.
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
-
Ours proposes the BRIDGE framework, which models the integration of SFT and RL as a bilevel optimization problem. In this framework, an SFT-based upper-level teacher learns to selectively transfer beneficial supervisory signals to an RL-based student via a lightweight LoRA module, achieving an average absolute improvement of over 3 percentage points across five mathematical reasoning benchmarks.
- FloorplanQA: A Benchmark for Spatial Reasoning in LLMs Using Structured Representations
-
FloorplanQA systematically diagnoses the "pure symbolic spatial reasoning" capabilities of 15 cutting-edge LLMs using 2,000 JSON/XML-formatted 2D indoor layouts and 16,000 geometric problems (distance, visibility, pathing, placement, etc.). The study reveals that while models can calculate simple distances, they consistently fail at set unions, planning, and constraint satisfaction. Furthermore, Python tool augmentation fixes arithmetic errors but cannot salvage failures at the algorithmic level.
- Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
-
The authors argue that the current collapse of "LLM self-play" within a few rounds is fundamentally due to self-synthetic data failing to provide learnable information gain; they formalize "learnable information" using bounded MDL/epiplexity and propose three system-level designs—Asymmetric Co-evolution, Capacity Growth, and Proactive Information Seeking—to collectively ensure the monotonic increase of learnable information in the Proposer-Solver-Verifier self-evolution loop.
- R2-Router: A New Paradigm for LLM Routing with Reasoning
-
This paper proposes R2-Router, which transforms "output token budget" from a passive estimate into a controllable variable. By enabling the router to search in the joint (LLM, budget) space and using a lightweight multi-head quality predictor to extend each LLM from a static point into a quality-cost curve, it achieves comparable quality to existing routers at 4–5× lower cost.
- A Formal Comparison Between Chain of Thought and Latent Thought
-
Based on computational complexity theory, this paper formally compares the expressive power of CoT (Chain of Thought) and Latent Thought (Looped Transformer / Coconut). It proves that Latent Thought strictly reaches \(\mathsf{TC}^k\) under polylogarithmic depth, while CoT reaches at most \(\mathsf{TC}^{k-1}\). Simultaneously, in a probabilistic setting, it reveals for the first time that CoT can support FPRAS counting through stochastic decoding, thereby surpassing deterministic Latent Thought.
- Conformal Thinking: Risk Control for Reasoning on a Compute Budget
-
This paper reframes the problem of "when a reasoning LLM should stop thinking" from an uninterpretable threshold tuning task into a user-specified risk tolerance conformal risk control problem. By employing dual thresholds—an upper threshold to stop when the model is confident (controlling false positives) and a newly proposed parameterized lower threshold to force a stop when the model is "stuck" on unsolvable problems (controlling false negatives)—and automatically deriving thresholds via the UCB algorithm on a calibration set, the method achieves significant token savings on AIME / GPQA / MathVision while maintaining accuracy.
Browse all 78 LLM Reasoning papers →
🦾 LLM Agent (83)¶
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
-
EvolveR provides LLM agents with a closed-loop lifecycle: "Online interaction \(\rightarrow\) Offline self-distillation into principle libraries \(\rightarrow\) GRPO policy evolution." Instead of discarding past trajectories, the agent abstracts successes and failures into a retrievable "principle library" and uses RL to learn how to utilize its own principles to solve new tasks. It significantly outperforms RL agent baselines like Search-R1 across 7 multi-hop QA benchmarks.
- ACON: Optimizing Context Compression for Long-horizon LLM Agents
-
Acon utilizes failure trajectory contrast to optimize natural language compression guidelines, simultaneously compressing agent history and observation contexts. It reduces peak tokens by 26% to 54% on AppWorld, OfficeBench, and multi-objective QA while maintaining or improving success rates in long-horizon tasks.
- Towards a Science of AI Agent Reliability
-
Drawing on established practices from safety-critical engineering (aviation, nuclear power, and automotive), this paper decomposes AI agent "reliability" into 12 accuracy-independent metrics across four dimensions: consistency, robustness, predictability, and safety. Systematic evaluation of 15 frontier models on GAIA and \(\tau\)-bench reveals an industry-wide trend: while accuracy has skyrocketed over the past 24 months, reliability remains largely stagnant.
- Measuring Agents in Production
-
This is the first systematic empirical study investigating "how LLM agents in production are actually built and evaluated." Through 20 in-depth case studies and 306 practitioner surveys (filtering 86 deployed/pilot systems) across 26 domains, the authors find that production agents generally follow a "simple and controllable" route (\(68\%\) execute \(\le 10\) steps before human intervention, \(70\%\) directly prompt off-the-shelf models without weight fine-tuning, and \(74\%\) rely primarily on human evaluation). Reliability is identified as the number one challenge, and practitioners primarily address it through system-level design rather than algorithmic or model-layer innovation.
- Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
-
Skill-Pro explicitly extracts interactive experiences of LLM agents into a "activation + execution + termination" skill triplet. It uses semantic gradients to generate candidate skills and verifies them with a PPO-style trust region (PPO Gate) before inclusion. Ultimately, it achieves over 0.85 reuse rate and significant performance gains in ALFWorld/Mastermind with a minimal memory library of ~800 tokens.
- Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
-
RHB constructs a suite of realistic tool-based multi-step tasks (independent and chained modes across four families: data pipeline, log forensics, performance optimization, and multi-file reconstruction) to quantify reward hacking in LLM agents. Across 13 frontier models, the study finds that RL post-training significantly increases exploit rates (DeepSeek-V3 0.6% vs. R1-Zero 13.9%), hacking rates rise with chain length, and exploits "relapse" on harder variants even for near-zero models, while lightweight environment hardening reduces exploit rates by 87.7% without compromising task success.
- ExCyTIn-Bench: Evaluating LLM Agents on Cyber Threat Investigation
-
This paper constructs ExCyTIn-Bench, the first benchmark evaluating LLM Agents for end-to-end "cyber threat investigation." Using 57 security log tables from a real Azure tenant, it automatically generates 7,542 SQL Q&A pairs with evidence chains via alert-entity bipartite graphs. It provides a MySQL environment for Agents to answer by querying logs and performing multi-hop evidence tracking. Currently, the strongest model, Claude-Opus-4.5, achieves a reward of only 0.606.
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
-
This paper proposes Thought-Aligner—a lightweight (1.5B/7B) plug-and-play safety model that performs causal debiasing of intermediate thoughts within the LLM agent's think-act-observe loop. By intervening before actions are executed, it improves the behavioral safety rate of six mainstream LLMs from approximately 50% to approximately 90% on ToolEmu/Agent-SafetyBench, while simultaneously increasing helpfulness by about 5%.
- Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
-
This paper proposes Persona2Web, the first open-web benchmark for personalized web agents. It utilizes "implicit user history + three levels of ambiguous queries + reasoning-aware scoring" to compel agents to infer user preferences from browsing records to disambiguate queries. Evaluations of five mainstream models, including GPT-4.1 and o3, reveal that the success rate for Level 2 queries is only 13% even when history is provided, highlighting a significant lack of true personalization in current web agents.
- A Minimal Agent for Automated Theorem Proving
-
This paper proposes AxProverBase—a minimalist Lean 4 theorem-proving agent. By relying on only three components—"compiler feedback + self-managed notebook + lightweight tool search"—it achieves or exceeds the performance of specialized systems like Hilbert/Seed-Prover using non-fine-tuned frontier LLMs (Claude Opus), while reducing costs by 100x.
Browse all 83 LLM Agent papers →
⚖️ Alignment & RLHF (37)¶
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
-
Aiming to resolve the dilemma where optimizing solely for accuracy encourages blind guessing while forcing refusal leads to over-conservatism, TruthRL directly optimizes truthfulness using a ternary reward ("Correct / Hallucination / Refusal") via GRPO. It reduces hallucination rates from 43.5% to 19.4% and increases truthfulness scores from 5.3% to 37.2%.
- Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
-
This paper theoretically demonstrates that the objectives of "improving accuracy" and "reducing calibration error" in RLVR (e.g., GRPO) training have negatively correlated gradient directions under the Fisher metric and are irreconcilable. It proposes DCPO: allowing the model to explicitly output a verbalized confidence segment after the reasoning trajectory, assigning independent rewards / advantages / masked gradients to reasoning tokens and confidence tokens. While maintaining the same accuracy as GRPO, it reduces the ECE from 0.435 to 0.128 (a 71.6% relative reduction).
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
-
MAHALO integrates "standardized PRM training + Multi-Action-Head DPO + PRM-guided decoding with KV-cache continuation" into a unified framework. This allows a single LLM to be simultaneously aligned across three categories: mathematics (verifiable), human values (non-verifiable), and Socratic tutoring (interactive), while enabling smooth preference switching during inference through head weights and PRM selection.
- Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
-
This paper reformulates the Bradley–Terry reward model as a generative process of Bayesian Non-negative Factor Analysis (NFA). By simultaneously modeling locally sparse instance latent variables \(\bm{\theta}\) and a globally sparse reward dictionary \(\Phi\), it suppresses reward hacking caused by shortcut features (e.g., length, style) via a "disentanglement-then-debiasing" mechanism. The entire framework is integrated into modern LLM backbones through amortized variational inference with Weibull reparameterization, consistently outperforming strong baselines like BT, Ensemble, and InfoRM on Unified-Feedback, RewardBench, HHH, and MT-Bench.
- GIST: Targeted Data Selection for Instruction Tuning with Gradient Subspace Projection
-
GIST frames "selecting instruction tuning data for a target task" as gradient subspace alignment. It demonstrates that methods like LESS, which use Adam states as a diagonal preconditioner, fail on LoRA due to cross-parameter coupling and low-rank task subspaces. Instead, GIST extracts a task-specific low-rank subspace via SVD of validation gradients and uses cosine similarity for sample selection. It matches or exceeds LESS on MMLU/TydiQA/BBH while requiring only 0.29% of the storage and 25% of the computation time.
- SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
-
SPARD combines "Safety-Projected Alternating Gradient (SPAG)" and "Relevance-Diversity DPP Safety Data Selection" to explicitly formulate "post-fine-tuning safety constraints" as a constrained optimization problem. It updates parameters for utility first and then uses a closed-form projection to pull them back into the safety half-space. By using only 3% task-relevant yet diverse safety samples, it reduces the average ASR of four harmful fine-tuning attacks from 87.93% (SFT) to 9.45% with negligible impact on downstream performance.
- Consistency Training Can Entrench Misalignment
-
This paper proposes the "consistency non-neutrality hypothesis." By evaluating 7 consistency training methods across 108 "model organisms," it finds that consistency training is not alignment-neutral—it systematically suppresses fragile reward hacking and emergent misalignment while amplifying stable sycophancy. Distribution shift, rather than score selection, is identified as the primary driver.
- Operationalising the Superficial Alignment Hypothesis via Task Complexity
-
The authors redefine the Superficial Alignment Hypothesis (SAH) using "task complexity"—an algorithmic information-theoretic metric representing the shortest program length required to solve a task at target performance. They unify three disparate lines of evidence (data-efficient, parameter-efficient, and inference-controlled) into a single strategy of finding short programs on the same length–performance Pareto curve. Experimental results indicate that adapting pre-trained models to tasks like mathematical reasoning, machine translation, and instruction following often requires only several kilobytes to megabytes of information, and the role of post-training is to compress the "program length required for high performance" by several orders of magnitude.
- PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
-
PICACO formalizes the challenge of "making an LLM adhere to multiple or even conflicting human values within a single prompt" as maximizing the "conditional Total Correlation (TC) between value sets and responses." Without updating model parameters, it automatically searches for a meta-instruction through an EM-like two-step iteration of "response enhancement + instruction refinement." PICACO outperforms strong baselines like OPRO and Modular Pluralism on five value evaluation sets containing up to 8 combined values across GPT-3.5, LLaMA-3.1-8B, and Gemini-1.5-Flash.
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
-
The problem of "which reward model should be used to align LLMs" is modeled as a Stackelberg game. It is proved that the optimal reward is a per-prompt threshold reward (giving full score \(B\) above the threshold and 0 below). This threshold is efficiently estimated using Monte Carlo sampling from the base model. Finally, the reward is softened via a sigmoid function and seamlessly integrated into inference-time alignment methods like CD/ARGS, increasing the average reward and GPT-4 Win-Tie rate against baselines to over 66% with almost zero additional overhead.
Browse all 37 Alignment & RLHF papers →
👻 Hallucination Detection (21)¶
- A Unified Definition of Hallucination: It's The World Model, Stupid!
-
This is a position paper advocating that "hallucinations" across various tasks—translation, summarization, open-domain QA, RAG, multimodal, and agents—be unified as one phenomenon: user-observable, inaccurate world modeling relative to a "reference world model." Every scenario is simply a different configuration of the "\((W, V, P)\)" triplet (Reference World \(W\), View Function \(V\), Conflict Policy \(P\)), converging fragmented definitions into a universal template for generating large-scale, comparable benchmarks.
- When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets (CAIA)
-
CAIA establishes the first "adversarial high-stakes" agent benchmark using 17 frontier LLMs across 178 temporally anchored real-world cryptocurrency tasks. Key findings: without tools, all models achieve only 12–28% accuracy (near random guess); with tools, the strongest GPT-5 reaches only 67.4% vs. 80% for junior human analysts. Critically, 55.5% of tool calls are biased toward "unreliable web searches" bypassing authoritative on-chain data, and the Pass@k metric systematically masks dangerous "trial-and-error" behaviors.
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
-
The GIFT method is proposed, which constructs a visual saliency map by tracking positive changes in visual attention ("gaze shifts") as the VLM interprets user queries. During the decoding stage, it simultaneously enhances attention for both visual and query tokens to maintain cross-modal fusion balance, achieving up to 20.7% improvement on CHAIR with only 1.13× latency overhead.
- Automatic Layer Selection for Hallucination Detection
-
FEPoID (First Effective Peak of Intrinsic Dimension) is proposed as a training-free automatic layer selection criterion. Combined with the First Sentence Truncation (FST) strategy, it consistently selects near-optimal intermediate layers across various QA and summarization hallucination detection benchmarks, significantly outperformed existing baseline methods.
- Hallucinations Undermine Trust; Metacognition is a Way Forward
-
This position paper argues that "totally eliminating LLM hallucinations" is theoretically impossible without incurring a "utility tax" (discrimination gap); the authors advocate shifting the goal from "eliminating hallucinations" to faithful uncertainty and treating this metacognition as an indispensable control layer for agentic LLMs when calling tools.
- Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
-
This paper identifies that LVLM hallucinations originate from "insufficient attention + forgetting during generation" regarding correct visual evidence. Observing a significant Inter-Layer Visual Attention Discrepancy (ILVAD) for visual evidence, the authors propose a train-free/plug-and-play method: constructing a visual evidence saliency map via inter-layer differentiation, then continuously weighting visual evidence tokens and "evidence-grounded" text tokens during generation. This consistently reduces hallucinations across 5 LVLMs and 5 hallucination/comprehensive benchmarks.
- From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
-
This paper treats LLM next-token prediction as a classification task on a massive vocabulary. By migrating two lightweight OOD detectors—NCI (proximity of features to weight vectors) and fDBD (distance from features to decision boundaries)—with two adaptations ("analytical proxy \(\mu_G\) for training feature means" and "calculating boundary distance only on top-\(k\) candidate tokens"), it derives a training-free, single-sample inference-time hallucination detector. It consistently outperforms baselines such as Perplexity, Semantic Entropy, and SelfCheckGPT on CSQA, GSM8K, and AQuA.
- Hallucination is a Consequence of Space-Optimality: A Rate-Distortion Theorem for Membership Testing
-
This paper formalizes "LLMs memorizing random facts" as a membership testing problem with continuous confidence scores. It proves that in the sparse limit of facts, the optimal memory cost exactly equals the minimum KL divergence between fact and non-fact output distributions—a "rate-distortion theorem." It further concludes that under the log-loss objective and given limited memory, the optimal strategy is neither abstention nor forgetting, but rather mapping a certain proportion of non-facts and facts to the same high-confidence point, identifying hallucination as the information-theoretically optimal error form.
- Harnessing Reasoning Trajectories for Hallucination Detection via Answer-agreement Representation Shaping
-
This paper proposes ARS for hallucination detection in Large Reasoning Models (LRMs). Instead of perturbing reasoning traces in the text space, ARS applies small perturbations directly to the latent representations at the end of the trace to decode counterfactual answers. Using "answer agreement" as a label, a lightweight contrastive head is trained to shape trace-conditioned answer embeddings, enabling embedding-based detectors to better distinguish hallucinations from truthful responses (\(66.85 \to 86.64\) AUROC on TruthfulQA).
- MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue
-
Ours proposes the MM-Snowball benchmark (4992 trajectories of 6-turn adversarial dialogues) to systematically characterize the "hallucination snowballing" phenomenon in Multimodal Large Language Models (MLLMs) during long dialogues. Based on this, ours designs a training-free Conflict-Aware Visual Rectification (CAVR) method that refreshes visual signals at the representation layer and adjudicates text-visual conflicts at the logit layer, significantly flattening the performance collapse curve in later dialogue stages.
Browse all 21 Hallucination Detection papers →
📊 LLM Evaluation (40)¶
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
-
This paper proposes Agent World Model, a fully synthetic pipeline encompassing scenarios, tasks, databases, MCP tool interfaces, and verifiers. It generates 1,000 executable, database-driven environments used to train tool-calling agents, achieving superior out-of-distribution generalization on BFCLv3, \(\tau^2\)-bench, and MCP-Universe.
- Toward Training Superintelligent Software Agents through Self-Play SWE-RL
-
This paper proposes Self-play SWE-RL (SSR), where a single LLM acts as both a "bug-creating proposer" and a "bug-fixing solver" within sandboxed code repositories. Using only Docker images as input and employing consistency checks and solve-rates as rewards for joint RL, SSR achieves self-improvements of +10.4 and +7.8 points on SWE-bench Verified and SWE-Bench Pro, respectively, consistently outperforming "human-data" baselines that rely on human-annotated issues and test suites.
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
-
This paper defines AI benchmark saturation as the loss of reliable discriminative power between frontier models. It proposes an uncertainty-aware saturation index based on leaderboard metrics and analyzes 60 text LLM benchmarks. The study finds that nearly half are highly saturated, and that benchmark age and test set size are more significant predictors of saturation than private test sets, open-ended outputs, or template diversity.
- Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum
-
This paper systematically investigates the behavior of probability-based objective functions in SFT, discovering that the standard NLL is not universally optimal: on tasks where the model has a strong prior, prior-leaning objectives like \(-p\) significantly outperform NLL (with gains up to 16%). Conversely, NLL remains superior on tasks with weak priors, revealing an objective selection principle governed by the model-capability continuum.
- Spherical Steering: Geometry-Aware Activation Rotation for Language Models
-
This paper proposes Spherical Steering: rotating activation vectors along geodesics on the unit hypersphere of LLM hidden states toward a "truthfulness direction" estimated from contrastive samples. Unlike traditional additive activation steering, this approach maintains activation magnitudes (norms) while significantly improving multiple-choice accuracy on benchmarks such as TruthfulQA, COPA, and StoryCloze (+10% range) without degrading open-ended generation quality.
- HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
-
HiPER transforms the flat RL used for LLM agents into a two-level Plan-Execute structure consisting of "high-level planning of subgoals + low-level execution of atomic actions." It introduces Hierarchical Advantage Estimation (HAE), which slices GAE along subgoal segments to perform coupled advantage estimation with bounded differences. On ALFWorld and WebShop, HiPER achieves success rates of 97.4% and 83.3% respectively (using Qwen2.5-7B), representing gains of +6.6% and +8.3% over the strongest baseline, GiGPO.
- Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
-
Ours proposes GraphGPO, which aggregates all rollout trajectories into a unified state transition graph. By leveraging global shortest path information on the graph to calculate distance-based advantages for each step, it achieves finer-grained credit assignment than trajectory-level attribution, significantly outperforming GRPO and GiGPO on ALFWorld, WebShop, and Sokoban.
- Who can we trust? LLM-as-a-jury for Comparative Assessment
-
This paper points out that the reliability of multiple LLM judges in pairwise comparisons varies significantly. It proposes the BT-\(\sigma\) model with judge-specific discrimination parameters, which simultaneously learns the ranking of candidate outputs and the reliability of each LLM judge without human calibration labels, thereby aligning more closely with human rankings than simple averaging or standard Bradley-Terry aggregation.
- BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback
-
The Bespoke benchmark is proposed, collecting 2,870 sessions from 30 annotators over 3 weeks of real chat and search history. It constructs an evaluation framework with fine-grained preference ratings and diagnostic feedback to systematically assess the personalization capabilities of search-augmented LLMs. Findings indicate that current models score below 60 on average across all configurations, with the bottleneck for personalization lying in history reasoning rather than generation.
- CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting
-
CapBencher injects randomness into each problem (generating multiple logically correct answers and randomly selecting one as the gold label) to cap the Bayes accuracy of a benchmark at a controllable level (e.g., 50%). This enables black-box statistical detection of data contamination in publicly released benchmarks—any model with an accuracy significantly exceeding the Bayes upper bound is flagged as contaminated.
Browse all 40 LLM Evaluation papers →
⚡ LLM Efficiency (48)¶
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
-
To address the bottleneck where diffusion Large Language Models (dLLMs) suffer from extremely slow inference due to bidirectional attention and the inability to reuse KV caches, this paper proposes dLLM-Cache. This training-free method applies long-interval caching for static prompts and short-interval refreshing for dynamic responses. By using Value cosine similarity (V-verify) to select and recompute the top 25% most "active" tokens, it achieves up to 9.1× FLOPs acceleration on LLaDA 8B / Dream 7B with almost no drop in performance.
- Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
-
SKILL-MOE proposes a training-free symbolic MoE framework that uses "skills" as routing signals. It extracts required skills for each problem, dynamically recruits \(k\) experts from 16 pre-trained LLMs based on skill-model profiles, and fuses multiple CoT responses via a task-level optimal aggregator. Combined with expert-batched inference, it runs 16 7-8B models on a single GPU, outperforming the strongest multi-agent baseline by 8.15% on average.
- CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
-
The authors reframe the heuristic-based problem of "identifying critical KV cache entries" as an optimization problem of "minimizing attention output perturbation." They derive an analytical upper bound for perturbation (weighted by both attention weights and value norms projected via \(W^O\)) and design a plug-and-play two-stage greedy selection algorithm. This method reduces the compression loss of SOTA eviction approaches like SnapKV, AdaKV, and HeadKV by more than half on average across 29 long-context datasets.
- Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs
-
Spiffy adapts speculative decoding for Diffusion Language Models (dLLMs): instead of training a separate draft model, it utilizes the target model's own distribution for "auto-speculation." It organizes multi-step denoising states into a Directed Draft Graph and maximizes the acceptance rate using an offline-calibrated graph structure. This achieves up to a 8.6× reduction in model forward passes and a 6.3× speedup in token throughput on LLaDA / Dream / SDAR, while provably maintaining lossless output distributions.
- Hyperparameter Transfer with Mixture-of-Experts Layers
-
This paper extends the maximal update parametrization (μP/CompleteP) to sparse MoE Transformers. It defines initialization and learning rate (LR) scaling rules for routers, expert up/down projections, and expert biases when model width, depth, number of experts, and expert width are simultaneously scaled. Using a three-level Mean-Field Dynamical Mean Field Theory (DMFT), the authors prove that this parametrization possesses a scale-invariant limit as \(n_{\text{embd}}, n_{\text{exp}}, n_{\text{hid}}, L \to \infty\) (at fixed activation sparsity \(\kappa\)). Optimal LRs and initializations can be directly reused from 38M active parameter base models to 2B parameter MoEs. MoEs trained with zero-shot hyperparameters achieve performance comparable to or better than dense GPT2 speedrun models at equivalent active parameter counts.
- Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking
-
This work reinterprets sequence mixing as exact Bayesian filtering. By utilizing the "information form" of the Kalman filter, it reformulates the sequential recursive update into a parallelizable prefix scan using Möbius (fractional linear) mappings. The resulting KLA is a plug-and-play, linear-complexity sequence mixing layer that is more expressive than GLA and provides explicit state uncertainty.
- OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
-
This paper reformulates KV cache eviction as a "layer-wise structural pruning" problem. By leveraging the second-order Taylor approximation from Optimal Brain Damage, it derives closed-form saliency scores for independent value pruning, independent key pruning, and joint key-value pruning units. These serve as plug-and-play "score replacements" for existing attention-only eviction frameworks such as H2O, TOVA, SnapKV, and AdaKV, achieving consistent improvements on LLaMA-3.1 and Qwen-2.5 across RULER and LongBench (e.g., AdaKV's performance increases by nearly 15% on query-agnostic RULER-4K with a 30% budget).
- SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm
-
To address the structural conflict where Pre-Norm and Post-Norm cannot coexist within a single-stream architecture, the authors propose SiameseNorm, a dual-stream residual architecture. It maintains an unnormalized stream as an identity gradient highway (Pre-Norm) and a normalized stream for main-path representation control (Post-Norm). By coupling these two streams via shared residual blocks, SiameseNorm consistently outperforms Pre-Norm baselines across 400M~15B dense/MoE language models, ViT, and DiT with negligible overhead.
- TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration
-
TEAM addresses the inherent mismatch in MoE Diffusion Language Models (dLLM) where "a large number of experts are activated but only a few tokens are accepted." By leveraging the temporal and spatial consistency of in-block decoding, TEAM designs differentiated expert activation and decoding strategies for three types of tokens: decoded, hot, and cold. This achieves up to a 2.2× speedup on SDAR 30B-A3B with near-zero precision loss.
- OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
-
OServe jointly models LLM serving "resource allocation + parallel strategy + request routing" as a bi-level maximum flow problem on a flow network. Combined with LSTM-based workload prediction and ad-hoc model switching via GPU interconnects, it addresses the heterogeneity of real-world traffic in both spatial (different request types) and temporal (varying composition over time) dimensions. End-to-end P99 latency and throughput improved by an average of 1.5× and a maximum of 2× compared to vLLM.
Browse all 48 LLM Efficiency papers →
📚 Pretraining (27)¶
- Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings
-
The authors point out that the mainstream positional encoding, RoPE, couples "content (what)" and "position (where)" into the same phase, leading to poor performance on tasks requiring "finding content by position" or "locating position by content." They propose PoPE, which uses softplus to separate magnitude (controlling what) and pure positional phase (controlling where). As a minor modification to RoPE, PoPE consistently outperforms it in diagnostic tasks, music/genomic/language modeling, and achieves length extrapolation to 10x the training length without any fine-tuning, surpassing YaRN which is specifically designed for extrapolation.
- Inverse Depth Scaling From Most Layers Being Similar
-
By measuring LLM hidden state dynamics and conducting controlled experiments with a teacher-student toy model, this paper proves that LLM loss is approximately inversely proportional to depth (\(\alpha_\ell \approx 1\)). This is attributed to an inefficient but robust "ensemble averaging" mode where the vast majority of layers perform functionally similar small-step updates to cancel out errors.
- MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
-
MOOSE-Star decomposes the problem of "training an LLM to directly generate scientific hypotheses"—originally an \(\mathcal{O}(N^k)\) combinatorial search—into two sequential subtasks: "Inspiration Retrieval + Hypothesis Synthesis." By integrating hierarchical tree retrieval, bounded composition, and motivation planning, it reduces optimal complexity from exponential to \(\mathcal{O}(\log N)\) and releases the TOMATO-Star dataset containing 108,717 papers with decomposition annotations.
- AC-ODM: Actor–Critic Online Data Mixing for Sample-Efficient LLM Pretraining
-
AC-ODM formulates the dynamic adjustment of pre-training data domain weights as a continuous control problem in reinforcement learning. Using the DDPG Actor-Critic framework, it perceives the model state in real-time, outputs sampling weights for each domain, and employs "inter-domain gradient alignment" as the reward. Theoretically, this is proven equivalent to maximizing constructive interference of gradients (effective descent step size). On Pythia-1B, it achieves optimal perplexity with approximately 66% fewer steps than strong baselines, scores a 27.5% relative improvement on MMLU, and increases HumanEval pass@1 by 2.23 times, with only a 0.4% increase in wall-clock time per step and 2% extra memory.
- If open source is to win, it must go public
-
This ICML 2026 position paper argues that "open-source AI" in its current form cannot truly democratize AI access or provide public goods in the same way Linux or PyTorch did. It posits that open source can only succeed if embedded within "Public AI"—infrastructure for compute, inference, post-training, and data provided by governments, national labs, universities, and non-profit institutions.
- On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
-
Using a set of carefully controlled Sudoku/Rush Hour tasks where "reasoning difficulty remains constant while only the horizon length varies," this paper systematically proves that task horizon itself is an independent root cause for LLM agent RL training collapse. The authors propose two horizon-reduction mechanisms—macro actions and subgoal decomposition—which not only stabilize training but also enable strong zero-shot generalization across longer horizons (horizon generalization).
- POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
-
POET-X implements a system-level acceleration and memory optimization for POET (reParameterized Orthogonal Equivalence Training), which is training-stable but slow and memory-intensive. By combining input-centric reconstruction, permutation kernel acceleration, block-diagonal batch parallelism, half-storage CNP, and Triton fusion, it achieves a 3× memory reduction and 8× speedup compared to the original POET. This allows for pre-training 8B~13B LLMs on a single H100, while AdamW triggers OOM under identical settings.
- Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from k-Parity
-
This paper decomposes the training objective of Masked Diffusion Language Models (MDLM) into a "signal term + noise term" using the analytically solvable \(k\)-parity task. It theoretically proves that the noise term acts as an implicit regularizer that suppresses grokking and avoids memory traps. Based on this, the authors propose Signal-Rich Mask Sampling, narrowing the training mask rate \(t\) from a uniform \(\mathcal{U}[0,1]\) to a middle-range window. This approach significantly reduces perplexity on 50M models and yields an 8.8% improvement in pre-training and 5.8% in SFT for 8B models.
- Annotations Mitigate Post-Training Mode Collapse
-
The authors observe that SFT aligns models with a low-entropy semantic prior, leading to "inverse scaling" where larger instruction-tuned models become increasingly repetitive. They propose "Annotation-Anchored Training"—tagging documents with semantic tags during pre-training and masking the loss on these tags during SFT—enabling the model to sample semantics before generating responses, which reduces the semantic diversity gap by 85% while maintaining instruction-following performance.
- Constrained Bayesian Experimental Design via Online Planning
-
This paper proposes COPEx: a semi-amortized scheme combining "offline pre-trained amortized posterior networks + design policies + online multi-step lookahead scenario trees." This allows Bayesian experimental design (BED) to dynamically adapt to budget, cost, and transition constraints at test time. COPEx consistently outperforms baselines such as VPCE, ALINE, and RL-BOED in EIG/RMSE across three types of tasks: constrained location finding, CES, and cost-aware AL.
Browse all 27 Pretraining papers →
✏️ Knowledge Editing (8)¶
- CrispEdit: Low-Curvature Projections for Scalable Non-Destructive LLM Editing
-
LLM editing is formulated as a constrained optimization problem: "minimize edit loss s.t. capability loss remains invariant". This is equivalently transformed via Bregman divergence into a low-curvature subspace projection of the Gauss-Newton Hessian (GNH). By employing K-FAC and a Kronecker eigenbasis technique that avoids explicit construction of the projection matrix, 3,000 edits are completed in 6 minutes on an A40. The average performance drop of LLaMA-3-8B across MMLU/IFEval/ARC-C/TruthfulQA/GSM8K is suppressed to \(< 1\%\), significantly outperforming AlphaEdit, MEMIT, and fine-tuning.
- From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing
-
This paper systematically analyzes why backward spreading in locate-then-edit works and where it falls short. It proposes forward replay: treating the hidden state of the first decisive layer as an optimization variable and performing a standard forward pass to obtain targets for subsequent layers. This achieves consistent performance gains over MEMIT/RECT/PRUNE/AlphaEdit without additional computational overhead.
- Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
-
This paper proposes UniKE—the first "cross-modal knowledge editing" benchmark for Unified Multimodal Models (UMMs) (2,971 editing subjects, 5,535 VQA-verifiable instances). It systematically reveals a modality gap where the "text-side editing success rate is ~92%, yet image generation VQA is only ~18.5%." By using a "reasoning-augmented parameter editing" protocol, it increases VQA accuracy by up to 18.6 percentage points and identifies the root cause as the LLM-to-DiT projection bottleneck using cosine drift metrics on the conditioning path.
- Reverse-Engineering Model Editing on Language Models
-
The paper reveals that parameter update matrices of locate-then-edit knowledge editing methods (ROME/MEMIT/AlphaEdit) leak "edited subject" fingerprints through their row spaces. It proposes a two-stage attack, KSTER (recovering subjects via SVD, then prompts via relative entropy drop), and a defense called Subspace Camouflage based on "semantic decoy" injection.
- AnyEdit++: Adaptive Long-Form Knowledge Editing via Bayesian Surprise
-
AnyEdit++ utilizes token-level Bayesian Surprise to identify semantic transition points in long-form text, replacing the fixed-window segmentation of AnyEdit with structure-aware Bayes-Chunk. It achieves stable improvements in BLEU and BERT Score across long-form knowledge editing tasks such as mathematics, code, news, and poetry.
- KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls
-
KORE injects new knowledge into LMMs through two-stage "knowledge-oriented controls": automatically expanding single facts into structured multi-turn conversations and instruction tasks (to enhance generalization), while initializing LoRA adapters using the null space of the covariance matrix of prior knowledge (to minimize interference with existing capabilities). It achieves both strong adaptation and strong retention on LLaVA-v1.5 / Qwen2.5-VL.
- Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence
-
Ours starts from the "dimension collapse" hypothesis, proving that parameter-level knowledge editing is amplified along directions with low singular values and accumulates linearly with sequential editing. This systematically degrades core LLM capabilities across multiple models, datasets, and evaluation dimensions. Ours further indicates that a simple retrieval-based baseline, SCR, outperforms existing parameter editing methods in all settings.
- The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models
-
This paper proves from an optimization perspective that the stability of sequential editing (SE) stems from "cumulative updates being equivalent to the solution of one-time editing (OTE)." Fancy mechanisms like AlphaEdit's null-space projection or post-processing regularizations in PRUNE/RECT are not the critical factors—as long as OTE-SE alignment is ensured, 2000 steps of sequential editing can be stably completed across four mainstream LLMs even after removing these regularizations.
💬 LLM (Other) (39)¶
- Position: Adversarial ML for LLMs Is Not Making Any Progress
-
This position paper argues that adversarial machine learning (ML) research in the LLM era focuses on problems that are "harder to define, harder to solve, and harder to evaluate" compared to traditional classifier scenarios. Having made slow progress on "toy problems" like \(\ell_p\) robustness over the past decade, the full shift to LLMs risks another decade of research without producing measurable or reproducible security guarantees.
- YAQA: End-to-End KL Minimizing Adaptive Weight Quantization for LLMs
-
YAQA shifts the proxy objective of LLM weight quantization from "layer-wise activation error" to "end-to-end model output KL divergence." Using a Hessian sketch via Kronecker decomposition, it provides the first end-to-end error bound. It reduces KL divergence by approximately 30% compared to GPTQ/LDLQ, even outperforming Quantization-Aware Training (QAT) in accuracy, while maintaining the same inference speed.
- Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
-
This paper proposes Compute as Teacher (CaT): it "synthesizes" a pseudo-reference answer from \(G\) rollouts already sampled by GRPO using a frozen anchor model. In non-verifiable domains, the model uses binary rubrics self-derived from this pseudo-reference to score each rollout as an RL reward. This directly converts inference compute into supervision signals without any human annotation, achieving up to a 30% improvement over baselines on HealthBench and matching or exceeding inference-time aggregation with 9× lower test-time compute.
- Multi-Agent Teams Hold Experts Back: Why Self-Organized LLM Teams Fail to Retain "Experts"
-
This paper systematically evaluates self-organized heterogeneous LLM teams using the organizational psychology standard of "strong synergy" (team \(\ge\) strongest individual). It finds that even when explicitly informed of expert identities, teams underperform experts by 6.3%–41.1% on frontier ML benchmarks. The root cause is not the inability to recognize experts, but a reluctance to let them lead—LLMs favor "middle-ground integration" over "epistemic deference." This consensus mechanism dilutes expertise as team size grows but, conversely, makes teams exceptionally robust against adversarial members.
- Stop Automating Peer Review Without Rigorous Evaluation
-
This is a position paper: through empirical measurements of real ICLR 2026 reviews and 60 simulated reviews, the authors identify two major failures in current LLM reviewing: the hivemind effect (high convergence) and paper laundering (zero-shot paraphrasing alone can increase scores by 0.45). Consequently, they argue that "LLMs should not directly generate review comments without rigorous evaluation" and call for the establishment of a "science of review automation."
- Optimizing Diversity and Quality through Base-Aligned Model Collaboration
-
The authors propose BACO, an inference-time token-level routing framework. It allows an "unaligned base model" and an "aligned instruct model" to switch token-by-token during a single decoding pass. Decisions are based on logit uncertainty and content signals, achieving base-level diversity and aligned-level quality without re-training or multiple sampling. The best router achieves a 21.3% joint improvement in diversity and quality over the strongest baseline.
- Position: The ML Community Must Build an AI-Augmented Peer-Review Ecosystem
-
This is a position paper arguing that the machine learning community must urgently build an "AI-augmented" peer-review ecosystem—treating LLMs as collaborative assistants for authors, reviewers, and Area Chairs (ACs) rather than replacements. The paper identifies that the primary near-term bottleneck is not the lack of stronger models, but the absence of structured process data that records "why scores changed" or "which specific rebuttal addressed which concern."
- Differential Syntactic and Semantic Encoding in LLMs
-
By averaging hidden representations of sentences sharing the same syntactic structure or the same meaning to obtain "syntactic centroids" and "semantic centroids," the authors demonstrate that a significant portion of syntactic/semantic information in LLMs like DeepSeek-V3 is encoded via linear superposition. Moreover, these two types of information exhibit clear separability in layer-wise distribution and orthogonal ablation—supporting the linguistic hypothesis of "syntactic autonomy."
- SAC-Opt: Semantic Anchors for Iterative Correction in Optimization Modeling
-
SAC-Opt "back-translates" LLM-generated optimization solver code into structured semantic anchors (constraints and objectives), compares them item-by-item with the original problem description's anchors, and iteratively rewrites only the inconsistent parts. It achieves an average performance gain of 7.7% across 7 public datasets and 21.9% on ComplexLP.
- Why Are Linear RNNs More Parallelizable?
-
This paper uses circuit complexity to strictly explain why Linear RNNs are more easily parallelized like Transformers compared to traditional non-linear RNNs: LRNNs fall within arithmetic circuit classes of approximate log-depth, whereas non-linear RNNs can express harder-to-parallelize \(\mathsf{logspace}\) / \(\mathsf{polynomial}\)-time complete problems, forming a fundamental trade-off between expressivity and parallelizability.
Browse all 39 LLM (Other) papers →
📖 NLP Understanding (2)¶
- Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
-
This paper formalizes the problem where "user-provided corrupted contexts degrade LLM performance" as a risk control task. By using zero-shot performance as a "safety baseline," combining dynamic early-exit (predicting at intermediate layers to avoid late-layer overthinking of harmful contexts) with a context-aware loss and an improved Learn-then-Test framework (preserving negative loss values via risk transformation rather than clipping), this method guarantees risk \(\leq\) user-specified \(\epsilon\) while achieving \(> 50\%\) computational acceleration across 9 tasks.
- Causal Fine-Tuning under Latent Confounded Shift
-
This paper proposes Causal Fine-Tuning (CFT): an SCM-inspired decomposition of "high-level stable features \(C\) + low-level confounding-sensitive features \(\Phi\)" is embedded into standard BERT fine-tuning. By utilizing a front-door style do-calculus adjustment for prediction, it significantly outperforms single-domain generalization baselines such as SFT/SWA/WISE under text spurious correlation injection attacks.
✍️ Text Generation (2)¶
- Characterizing the Effect of Noise in Language Generation in the Limit
-
Under the Kleinberg-Mullainathan formal framework of "language generation in the limit," this paper proves that for both uniform and non-uniform generation, noise level 1 is equivalent to any finite noise level \(i \geq 1\) (hierarchy collapse), while a strict separation exists between the noise-free case and noise level 1. Furthermore, it provides the first complete characterization of non-uniform noise-dependent generatability.
- Score-Repellent Monte Carlo: Toward Efficient Non-Markovian Sampler with Constant Memory in General State Spaces
-
SRMC utilizes a \(d\)-dimensional running score average (rather than an \(|\mathcal{X}|\)-dimensional empirical measure) to record history. This history is then incorporated into an exponential score-tilt to construct a surrogate target \(\pi_\theta\) that "repels already visited regions." By wrapping this around any base MCMC kernel, the authors implement a non-Markovian, low-variance, normalization-free sampler with constant memory in general state spaces.
🗣️ Dialogue Systems (5)¶
- Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives
-
This paper models LLM-as-a-Service as a "principal-agent" problem, proving that current mainstream "pay-per-token" mechanisms naturally incentivize service providers to re-segment the same string into longer token sequences for overcharging. Furthermore, even if providers are forced to disclose next-token distributions, overcharging without detection remains NP-Hard rather than impossible—the authors provide a simple heuristic algorithm that increases reported tokens by up to 11.2% while maintaining plausibility. Finally, it is proven that the only additive pricing mechanism that eliminates this incentive is "linear pay-per-character."
- From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
-
Addressing two major bottlenecks in post-training multi-turn interactive tool-using agents—expensive high-quality data and RL signal degradation from user simulation noise—the authors propose "AReaL-SEA," a self-evolving multi-agent data synthesis pipeline that generates executable verifiers as rewards. Combined with an RL recipe featuring user model SFT, large batches, and dynamic filtering GRPO, Qwen3-235B achieves a pass^1 of 73.0 in Airline and 98.3 in Telecom on τ²-bench, matching or exceeding Claude/Gemini/GPT-5.
- DiscoverLLM: From Executing Intents to Discovering Them
-
DiscoverLLM formalizes the scenario where "the user has not clearly defined their goals" as a progressive discovery process within a hierarchical intent tree. By using a rewardable hierarchical user simulator, the model is trained to actively explore divergently when goals are unclear and converge for execution when they are clarified. On creative writing, technical writing, and SVG tasks, the method achieves a +10% improvement in satisfaction and a -40% reduction in dialogue length compared to baselines like CollabLLM.
- Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
-
This paper points out that traditional Prefill-Decode (PD) disaggregated architectures are significantly inefficient in multi-turn dialogue scenarios due to the repeated P→D recomputation and transmission of KV caches for each turn. It proposes PPD (Prefill-capable Decode), a dynamic routing system that allows decode nodes to decide whether to process Turn 2+ append-prefills locally based on SLO weights, reducing Turn 2+ TTFT by approximately 68%.
- Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
-
Concatenating full histories in multi-turn dialogues is expensive and leads to lost clues. This paper proposes C-DIC: viewing dialogues as interleaved "topic threads," it stores revisable per-thread compressed states in a compact memory. Each turn follows a lightweight "Retrieval \(\to\) Revision \(\to\) Write-back" cycle, trained with retrieval-aware truncated backpropagation through time (ra-TBPTT), maintaining stable latency and perplexity over hundreds of turns.
🌐 Multilingual & Translation (3)¶
- Optimizing Language Models for Crosslingual Knowledge Consistency
-
This paper addresses the issue of multilingual LLMs providing conflicting answers to the same question across different languages. It designs an RL objective using the "log-likelihood of the answer in another language" as a reward, proving that the optimal policy follows a product-of-experts form and guarantees crosslingual preference consistency when \(\gamma_1\gamma_2=\beta^2\). Based on this, the authors derive DCO (Direct Consistency Optimization), a reward-model-free and online-sampling-free algorithm. Experiments across 9 LLMs, 3 multilingual QA benchmarks, and 26 languages demonstrate simultaneous improvements in crosslingual consistency (RankC) and response accuracy.
- Edit-Based Refinement for Parallel Masked Diffusion Language Models
-
ME-DLM introduces a lightweight "decode-then-edit" refinement stage to masked diffusion language models (e.g., LLaDA). The first stage generates a draft via standard parallel unmasking, while the second stage performs parallel corrections using replace/delete/insert actions supervised by the shortest edit distance scripts. Using only 1/8 of the diffusion step budget, it outperforms LLaDA-Instruct by +11.6 on HumanEval and +33.6 on GSM8K.
- Toward Robust Multilingual Adaptation of LLMs for Low-Resource Languages
-
LiRA inserts a lightweight fine-tuning module featuring "anchoring + consistency regularization" between a frozen multilingual encoder and an English LLM. It constrains the sentence vectors of low-resource languages into a shared English semantic space through two theoretically controllable quantities: \(\epsilon_1\) (anchoring error) and \(\epsilon_2\) (translation KL distance), achieving stable improvements across retrieval, ranking, and reasoning tasks.
🔍 Information Retrieval & RAG (26)¶
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning
-
Graph-R1 reformulates GraphRAG as an end-to-end RL framework featuring a "knowledge hypergraph environment + multi-turn think–query–retrieve–answer agent + outcome-oriented GRPO." By utilizing lightweight n-ary hypergraph construction and dual-path hyperedge retrieval with RRF fusion, it improves the F1 score of 7B models from Search-R1's 46.19 to 57.82 across six standard RAG datasets.
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
-
The authors perform a systematic empirical audit using PhoneBook and a newly constructed PaperQA benchmark, treating LoRA as a knowledge memory unit that can be independently trained, loaded, and combined. They quantitatively provide full-link design guidelines covering "Rank \(\rightarrow\) Capacity \(\rightarrow\) Efficiency \(\rightarrow\) Multi-module Combination \(\rightarrow\) Complementarity with RAG/ICL."
- Ranking-Free RAG: Replacing Re-Ranking with Selection in RAG for Sensitive Domains
-
This paper introduces METEORA, a trio consisting of a DPO-trained rationale generator, statistical elbow detection, and a shared-framework Verifier. It replaces the uninterpretable, top-\(k\)-dependent re-ranker in RAG, achieving higher recall, an 80% reduction in evidence volume, and a 4.4× improvement in adversarial robustness across six sensitive domain datasets.
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
-
ReSeek adds a JUDGE action to RL-trained search agents and utilizes BGE-reranker to calculate "ideal judgments" as process rewards. This enables agents to "soft-mask" invalid information and re-query after each retrieval. It also proposes FictionalHot, an anti-contamination benchmark based on fictional entities, achieving an average EM of 0.377 on Qwen2.5-7B, outperforming ZeroSearch by +3.1.
- Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
-
This paper designs a controllable method to measure "language preference" in multilingual RAG using internal signals (next-token citation prediction probability). It finds that six open-source LLMs systematically prefer citing English documents during long-form generation, even when English documents are irrelevant—suggesting language itself influences citation selection more than document relevance.
- HGMem: Hypergraph-based Working Memory to Improve Multi-step RAG for Long-Context Complex Relational Modeling
-
This paper reconstructs the working memory in multi-step RAG from a "flat list of facts" into a hypergraph. Each hyperedge serves as a memory point that can be updated, inserted, or merged. By leveraging the inherent ability of hyperedges to connect \(n \geq 2\) entities, the system allows memory to continuously consolidate low-order facts into high-order concepts during interactions, significantly improving performance in long-context QA tasks that require "global sense-making."
- LEMUR: Learned Multi-Vector Retrieval
-
Lemur transforms multi-vector similarity search into a supervised learning problem. By using a two-layer MLP to map token-level embeddings to a low-dimensional latent space and leveraging existing single-vector ANNS indices for retrieval, it achieves speeds an order of magnitude faster than methods like PLAID and MUVERA.
- Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
-
This paper points out that while existing RAG poisoning attacks can manipulate LLM outputs using a small number of malicious passages, they are not truly stealthy. Successful low-budget attacks inevitably cause the model to focus excessive attention on malicious passages. Consequently, the authors filter out anomalous passages using a Normalized Passage Attention Score (NPAS) and a variance-based AV Filter. Across a setup of 4 datasets × 5 LLMs × 5 attacks, it improves RACC by up to 20% compared to Certified Robust RAG.
- Less Is More: Elevating RAG via Performance-Driven Context Compression
-
CORE-RAG trains a 1.5B small compressor using GRPO reinforcement learning with "performance-as-reward," compressing retrieved top-k documents into summaries of ~3% original length. It not only avoids performance degradation but also achieves an average improvement of 3.3 EM over full-context RAG across four QA benchmarks.
- Hierarchical Abstract Tree for Cross-Document Retrieval-Augmented Generation
-
Ψ-RAG replaces RAPTOR's k-means with a "merge-collapse" hierarchical clustering to construct cross-document abstraction trees. It incorporates a retrieval-response Agent with multi-turn rewriting capabilities and a hybrid BM25 index, enabling Tree-RAG to match or exceed Graph-RAG in corpus-level, multi-hop QA for the first time. The average F1 score is 25.9% higher than RAPTOR and 7.4% higher than HippoRAG 2.
Browse all 26 Information Retrieval & RAG papers →
💻 Code Intelligence (22)¶
- MARS: Modular Agent with Reflective Search for Automated AI Research
-
MARS reframes automated AI research as a problem of "searching for the optimal solution within a software repository space." Built on three pillars—Budget-Aware MCTS, a modular "Design-Decompose-Implement" pipeline, and Comparative Reflective Memory—it achieves SOTA among open-source frameworks on MLE-Bench with a 31.1% gold medal rate (Gemini-3-Pro-Preview) and demonstrates an "Aha! moment" with a 63% cross-branch lesson transfer rate.
- SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
-
The authors developed a "language-agnostic unified construction pipeline + interactive installation Agent + triple-model ensemble for issue clarity filtering" to automatically mine 32,079 executable SWE tasks across 20 languages and 3,617 repositories from GitHub (accompanied by 120,000+ PR-derived tasks). Each task includes pre-built Docker images, fail-to-pass tests, and instance-level diagnostic metadata, providing a stable, training-oriented substrate for large-scale reinforcement learning of SWE Agents rather than just evaluation.
- CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
-
CentaurEval is proposed as the first unified evaluation framework for human-AI collaborative programming. By designing 45 "Collaboration-Necessary" task templates, it demonstrates that LLMs alone achieve only a 0.67% pass rate and humans alone achieve 18.89%, while human-AI collaboration reaches 31.11%, revealing that LLMs are evolving from execution tools into co-reasoning partners.
- MatchFixAgent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and Repair
-
MatchFixAgent fully transforms "equivalence validation + repair" for repository-level code translation into an LLM-based task. By replacing expensive cross-language interoperability engineering with six parallel semantic sub-analyzers (Control Flow, Data Flow, IO, Library API, Exception, and Specification), and layering a Test & Repair Agent with an Arbiter Agent, it raises validation coverage from 71.6% to 99.2% and the repairable defect ratio from 18.5% to 50.6% with only 1650 lines of code.
- Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning
-
RankTuner proposes the Relative Rank Indicator \(I_t\), which uses a single scalar signal comparing the "actual rank \(R_t\) of the ground-truth token" against the "expected rank \(\mathbb{E}[R_t]\) under the model distribution." By coupling probability \(p_t\) (task alignment) and entropy \(H_t\) (intrinsic uncertainty) into a token-level weight, it consistently outperforms pure probability/entropy reweighting baselines in Pass@1 for mathematical reasoning SFT.
- AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
-
AlgoVeri constructs a strictly aligned benchmark for verified code generation of classical algorithms across Dafny, Verus, and Lean. It demonstrates that current LLMs still face significant gaps in handling complex global invariants, system-level constraints, and explicit proof search, with success rates in Lean and Verus being substantially lower than those in Dafny.
- How can we assess human-agent interactions? Case studies in software agent design
-
The authors propose the PULSE framework—which collects user feedback, trains an ML model to predict user satisfaction, and employs Prediction-Powered Inference (PPI) to combine real human labels with model pseudo-labels for efficient estimation of agent design effects. Deployed on the open-source coding agent OpenHands across 15,000 users and 36,000 sessions, this work represents the first large-scale real-world evaluation of agent design. Results show that PULSE narrows confidence intervals by approximately 40% compared to standard A/B testing and reveals that benchmark performance can be anti-correlated with human preference (e.g., GPT-5 outperformed Claude-Sonnet-4 on 6/7 benchmarks, yet humans preferred Claude on 4/7 task subsets).
- NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents
-
NEMO treats Autonomous Coding Agents (ACA) as a "first-class abstraction" on par with LLM calls. It enables independently generated simulators and optimizers to cross-verify via execution results in a shared sandbox, combined with diverse memory retrieval and MBR/self-consistency decoding. It achieves SOTA on 8 out of 9 optimization modeling benchmarks, leading by up to 28 percentage points.
- BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
-
BoostAPR constructs a three-stage pipeline for training program-repair models via RL: execution-verified SFT → training sequence-level + line-level dual reward models → redistributing sequence rewards to key edit-line spans using the line-level model during PPO. Using Qwen2.5-Coder-32B, it pushes SWE-bench Verified performance from 17.8% to 40.7% (+22.9pp) and achieves 24.8% on Defects4J through cross-lingual transfer.
- MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering
-
MEnvAgent employs a "Plan-Execute-Verify" three-stage multi-agent closed-loop and an environment reuse mechanism to automatically build executable and verifiable (Fail-to-Pass) Docker environments for real-world repositories across 10 languages. On the self-constructed MEnvBench, it improves the F2P rate by 8.6% and reduces construction time by 43%, facilitating the creation of MEnvData-SWE, the largest polyglot verifiable SWE training set to date.
Browse all 22 Code Intelligence papers →
🎨 Image Generation (141)¶
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
-
WISE constructs a text-to-image evaluation benchmark containing 1000 knowledge-dense prompts. It examines whether models can transform implicit semantics—such as cultural common sense, spatio-temporal reasoning, and natural science knowledge—into correct visual content. The study reveals significant shortcomings in world knowledge generation for existing T2I and unified multimodal models.
- DFlash: Block Diffusion for Flash Speculative Decoding
-
DFlash replaces autoregressive drafters like EAGLE-3 with a lightweight "Block Diffusion" drafter. By injecting multi-layer hidden features of the target model as KV into every layer of the draft model, it enables parallel drafting of an entire block of tokens in a single forward pass, achieving up to 6× lossless acceleration—approximately 2.5× faster than EAGLE-3.
- Esoteric Language Models: A Family of Any-Order Diffusion LLMs
-
Eso-LMs deeply integrate AR and Masked Diffusion at the loss, attention, and sampling levels. By utilizing a "causal-on-shuffled-sequence" denoising Transformer, it simultaneously supports parallel diffusion and left-to-right AR. This marks the first time an MDM can utilize exact KV cache during the diffusion phase, achieving 14–65× speedups over MDLM and 3–4× over BD3-LM on OWT long contexts, while reaching SOTA on the speed–quality Pareto frontier.
- SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes
-
SceneSmith utilizes a designer-critic-orchestrator VLM agent triangle to construct indoor scenes layer-by-layer on a hierarchical tree of "layout \(\rightarrow\) furniture \(\rightarrow\) small objects." It deeply couples text-to-3D generation, articulated object retrieval, and physical property estimation into the agent toolchain. Generating directly from a single natural language prompt, it produces dense, actionable environments ready for physical simulators. Each room averages 71 objects (compared to 11–23 in baselines), with an inter-object collision rate \(< 2\%\) and a gravity-based stability rate of \(96\%\), significantly outperforming all prior methods.
- GenExam: A Multidisciplinary Text-to-Image Exam
-
GenExam adopts the "drawing exam" as the gold standard for measuring the integrated reasoning-understanding-generation capabilities of T2I models. By providing ground-truth images and fine-grained scoring points for 1000 questions across 10 disciplines, results reveal that even the strongest closed-source model, Nano Banana Pro, achieves only a 70.2% strict score, while most open-source T2I and unified MLLMs score below 3%.
- Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
-
GCPO transitions the step-level optimization in flow matching post-training—where GRPO assigns the "same final reward as advantage to every step"—into "chunk-level" optimization. By adaptively grouping consecutive steps into chunks based on flow matching's own temporal dynamics \(L1_{rel}(x,t)\) and utilizing normalized chunk-level importance ratios \(r^i_j\) for policy updates, it smooths out erroneous gradients caused by the "final success \(\neq\) step-wise optimal" mismatch. This achieves a relative gain of up to 43% over GRPO on HPSv3, ImageReward, GenEval, and DPG.
- OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the Wild
-
OmniAID employs a decoupled MoE architecture consisting of "Semantic Experts + a Universal Artifact Expert" to learn two types of forgery cues—"content-related flaws" and "universal generation artifacts"—within a low-rank residual subspace derived from CLIP-ViT attention weight SVD. Coupled with the modern Mirage dataset, it achieves state-of-the-art average accuracies of 95.9%, 91.4%, and 88.4% on GenImage, Chameleon, and Mirage-Test benchmarks, respectively.
- Adversarial Flow Models
-
The authors add an optimal transport regularization term \(\|G(z)-z\|^2\) to the GAN training objective, constraining the GAN's "arbitrary transport map" to a unique Wasserstein-2 optimal transport map. This allows adversarial training on pure Transformers to stabilize for the first time and perform end-to-end single-step generation. On ImageNet-256, the 1NFE FID reaches 2.38 (XL/2) and 1.94 (112-layer recursive model).
- SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning
-
The authors identify an "attention collapse" issue in MLLM-based editing reward models—where the model focuses on sink tokens rather than comparing the original and edited images—and propose SpatialReward. It directs an 8B model to first predict bounding boxes for edited regions and then use these box tokens as anchors for interleaved cross-image reasoning. Combined with a 260K-sample spatial-aware dataset and a two-stage GRPO training process, it achieves SOTA on three reward benchmarks and improves the GEdit-Bench score of OmniGen2 by +0.90 (double the improvement of GPT-4.1).
- PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World
-
Interactive 3D object creation is reframed as a two-stage "physical planning followed by physical generation" problem. A VLM acts as a physical architect to generate a "Hierarchical Physical Blueprint" containing hierarchy, materials, and kinematic constraints. Subsequently, a diffusion model utilizes KineVoxel Injection to jointly denoise articulation parameters and geometric voxels. Combined with the PhysDB dataset—comprising 150k assets with four-tier annotations—this approach achieves the first generation of 3D assets from a single view that are directly graspable, pushable, and articulatable within physics engines.
Browse all 141 Image Generation papers →
🎬 Video Generation (32)¶
- EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
-
EPiC utilizes a "first-frame visibility mask" approach to construct pixel-aligned anchor videos directly from arbitrary in-the-wild videos. By pairing this with Anchor-ControlNet—comprising only 26M parameters (<1% of the backbone) and operating exclusively on visible regions—Ours achieves SOTA I2V camera control precision and zero-shot generalization to V2V. This is accomplished while freezing the CogVideoX-5B-I2V backbone, using only 5K videos and 500 training steps.
- LuVe: Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts
-
LuVe redefines UHR video generation from "passive detail enhancement" to "active content completion." Through a three-stage cascade (Low-Resolution Motion → Latent Space Upsampling → High-Resolution Refinement) and frequency domain analysis-driven Dual Frequency Experts (Low-Frequency Expert for global semantic consistency, High-Frequency Expert for texture refinement), it achieves a total score of 84.03 on VBench 4K, surpassing UltraWan-4K's 83.75.
- VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
-
VideoGPA utilizes a Geometric Foundation Model (GFM) to reconstruct generated videos into 3D point clouds and project them back into the original frames. It uses "reprojection error" as a self-supervised geometric consistency reward to automatically construct preference pairs. By applying DPO (fine-tuning ~1% parameters via LoRA with only ~2500 preference samples), it aligns pre-trained video diffusion models to a 3D-consistent manifold, significantly mitigating object deformation and spatial drift without compromising image quality.
- T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
-
T2AV-Compass is the first comprehensive evaluation benchmark for Text-to-Audio-Video (T2AV) generation. It features 500 complex prompts and a dual-level evaluation framework combining low-level signal metrics with high-level MLLM diagnostics. By evaluating 15 cutting-edge T2AV systems, it quantitatively reveals an "audio realism bottleneck," where even top-tier models achieve over 85% realism in video but only approximately 50% in audio.
- Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
-
QVG is a training-free KV-cache quantization framework for autoregressive video diffusion. By employing semantic-aware clustering for token smoothing and progressive residual multi-stage compression, it reduces KV memory footprint to 1/7 of the original on LongCat-Video/HY-WorldPlay/Self-Forcing with <4% end-to-end latency overhead. At 2-bit, its quality significantly outperforms LLM quantization baselines like KIVI and QuaRot.
- VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation
-
VAnim models open-domain text-to-SVG animation as "sparse state updates on a persistent DOM tree" + "Identification-First motion planning" + "GRPO rendering-aware reinforcement learning." This approach compresses sequence lengths by \(9.86\times\) while maintaining topological consistency, significantly outperforming GPT-5.2, Gemini 3 Pro, and LiveSketch.
- Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
-
Light Forcing is the first sparse attention scheme customized for autoregressive (AR) video diffusion models. Chunk-Aware Growth (CAG) quantifies the cumulative error contribution of each generated chunk to dynamically allocate sparsity, while Hierarchical Sparse Attention (HSA) flexibly captures historical dependencies through frame-level → chunk-level dual-mask selection. It achieves 1.30× end-to-end / 3.79× attention speedup on Self Forcing, with a VBench total score of 84.5 > dense baseline 84.1.
- Self-Refining Video Sampling
-
The pretrained flow matching video generator is reinterpreted as a "denoising autoencoder." During inference, a Predict-and-Perturb inner loop iteratively corrects latent deviations within the same noise level. An uncertainty mask derived from model self-consistency is applied to refine only dynamic regions. This approach significantly enhances motion coherence and physical plausibility without any external verifier or additional training, achieving a human preference rate exceeding 70%.
- OLAF-World: Orienting Latent Actions for Video World Modeling
-
OLAF-World learns transferable latent actions through Sequence-level Control-Effect Alignment (Seq∆-REPA)—turning unlabeled videos into action-controllable video world models and achieving zero-shot action transfer across contexts. With only 1 minute of annotated data, it achieves performance comparable to AdaWorld with 2 hours of data (rotation control accuracy 0.4680 vs 0.6420).
- Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
-
SVOO discovers that the attention sparsity of each layer in video DiT is an intrinsic property that is "input-independent within layers and significantly heterogeneous between layers." Based on this, it performs offline per-layer sparsity calibration followed by online QK bidirectional co-clustering for block partitioning. It achieves up to 1.93× speedup while maintaining a PSNR of 29 dB across 7 models (e.g., Wan, HunyuanVideo) without any training.
Browse all 32 Video Generation papers →
🧩 Multimodal VLM (89)¶
- ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning
-
ACTIVE-o3 delegates the decision of "where and how to look" to an MLLM for autonomous learning. Using pure reinforcement learning (GRPO), the model is trained to parallelly select up to 3 sub-regions most worthy of magnification. A dual-form reward mechanism (task reward + heuristic reward) is employed to solve the sparsity of pure task rewards. The method consistently outperforms baselines in small/dense object detection, remote sensing, autonomous driving, and interactive segmentation, while simultaneously enhancing general understanding capabilities such as RealWorldQA and MME.
- VLA-Arena: An Open-Source Framework for Evaluating Vision-Language-Action Models
-
VLA-Arena proposes a structured VLA benchmark—systematically quantifying difficulty through three orthogonal dimensions: task structure, language command, and visual observation. With 170 tasks, it reveals key deficiencies in generalization, visual perception, and safety of existing VLA models.
- Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness
-
This paper ports the "negative selection" principle from biological immune systems to VLMs such as CLIP. It employs an LLM to actively hallucinate a set of "look-alike but non-target" text descriptions as semantic antibodies. A lightweight adapter then pushes visual features away from these antibodies, significantly reducing "high-confidence misclassification" in open-world scenarios without retraining the backbone.
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Vision-Language Models
-
The authors propose Circle-RoPE, which maps the 2D coordinates of image tokens onto a torus orthogonal to the text position axis. This forms a conical geometry where the RoPE distance from each text token to all image tokens is equal (PTD=0), eliminating cross-modal pseudo-positional biases while preserving internal image spatial structure through Alternating Geometric Encoding (AGE).
- VLANeXt: A Recipe for Building Robust VLA Models
-
This paper systematically explores the design space of VLA models, distilling 12 key design principles from over 500 controlled experiments to construct the efficient and powerful VLANeXt model. It surpasses SOTA on the LIBERO benchmark and validates these design principles through real-world robot tasks.
- ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation
-
ECG-R1 is the first "reasoning-type" medical multimodal large model dedicated to ECG interpretation. Through a suite of protocol-guided instruction data synthesis + decoupled signal/image encoding + interleaved modality dropout training + evidence-driven process reward RL, it improves ECG diagnostic accuracy from the previous SOTA (GEM) of 74.7 to 80.3, while maintaining cross-modal consistency even when a modality is missing.
- Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
-
Through a pilot study, the authors discovered that "explicitly lifting vision to point clouds and fusing them with 2D patches" is the most effective way to inject 3D information into VLA models. To address 3D data scarcity and domain gaps across different point cloud sources (simulation, sensor, or monocular estimation), Any3D-VLA is proposed. By employing hybrid point cloud training to learn source-agnostic geometric representations, it achieves a 29.2% zero-shot improvement over the strongest baseline (62.5% vs 33.3%) in real-world grasping tasks.
- SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
-
SAME explicitly decomposes "catastrophic forgetting" in multimodal continual instruction tuning (MCIT) for MoE-LoRA into two independent sources: router drift and expert drift. It addresses these using spectral-aware subspace constrained updates for the router, Riemannian preconditioning with historical input covariance for experts, and an adaptive task-level freezing mechanism to eliminate redundant updates. Ours consistently outperforms existing MoE continual learning SOTAs on CoIN, UCIT, and the newly established TriGap long-sequence benchmarks.
- Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs
-
The authors modify the causal attention mask in decoder-only MLLMs by "digging a hole" that allows preceding image tokens to retrospectively attend to subsequent text question tokens. This single-line mask modification requires no extra parameters or training data changes, achieving an average improvement of 6.2 points across 3 LLM backbones and 12 multimodal benchmarks.
- FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
-
This paper introduces FutureOmni, the first benchmark evaluating MLLMs' ability to forecast future events from audio-visual context (919 videos / 1,034 MCQs). Results show even the strongest model, Gemini 3 Flash, achieves only 64.8% accuracy. The authors propose OFF, a rationale-infused instruction tuning method that significantly enhances both forecasting and generalization for open-source models.
Browse all 89 Multimodal VLM papers →
🧠 VLM Reasoning (31)¶
- Efficient Reasoning with Hidden Thinking
-
Heima distills each stage (summary / caption / reasoning) of lengthy Multimodal LLM (MLLM) Chains-of-Thought (CoT) into a single special thinking token. This allows the model to "think" in latent space, reducing the token count from the 100-200 range to 13-16 while achieving zero-shot accuracy more stable than LLaVA-CoT. An accompanying LLM "interpreter" is trained to reconstruct the textual reasoning chain from the thinking token's hidden states, empirically validating the information-theoretic upper bound of compression loss.
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
-
Gui-Cursor reformulates GUI grounding from "one-shot coordinate prediction" into an interactive search of "moving the cursor on the screen to find the target." By training the VLM with GRPO using a dense reward with trajectory penalties, the model leverages visual feedback from rendered cursors to align numerical coordinates with screen positions. With only 8K samples, it improves GTA1's performance on ScreenSpot-Pro from 50.1% to 58.1%.
- ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
-
This paper systematically reveals structural failures in the widely used VSI-Bench due to 3D annotation drift and frame sampling inconsistency. By re-annotating 381 scenes and 5365 objects and designing frame-budget adaptive QA alongside "dummy video" stress tests (removing frames containing target objects), the authors construct ReVSI, a high-fidelity spatial intelligence benchmark. Evaluations show that open-source VLMs suffer performance drops of up to 40% on ReVSI while exhibiting high hallucination rates on dummy videos, exposing a systematic overestimation of current 3D reasoning capabilities in VSI-Bench.
- Vision-aligned Latent Reasoning for Multi-modal Large Language Model
-
This paper proposes VaLR: a method that inserts several "latent tokens" before each step of CoT reasoning in MLLMs and performs representation alignment (REPA) between these tokens and the patch features of visual encoders like DINOv3, SigLIP, or \(\pi^3\). This continuously "feeds" visual information back into the model during long-chain reasoning, improving the accuracy of Qwen2.5-VL on VSI-Bench from 33.0% to 52.9% and enabling MLLMs to demonstrate "longer reasoning leads to higher accuracy" test-time scaling behavior for the first time.
- Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
-
The authors decompose the KL loss of multimodal on-policy distillation into "language prior" and "visual grounding" sub-objectives based on a Bayesian chain. They find that the gradients of these two are nearly orthogonal, and standard distillation merely takes a passive bisector. Consequently, they propose Visual Gradient Steering (VGS) to actively bias the update direction toward the visual subspace, achieving average gains of +2.37%/+1.56% across seven multimodal reasoning benchmarks for Qwen3-VL 8B→2B/4B.
- Imagination Helps Visual Reasoning, But Not Yet in Latent Space
-
This paper employs causal mediation analysis to decompose "Latent Visual Reasoning (using MLLM hidden states as latent tokens for visual imagination)" into a causal chain \(X\to Z\to Y\). Empirical evidence reveals that latent tokens are neither varied with inputs (Input-Latent disconnection) nor significantly impact the final answer (Latent-Answer disconnection), questioning their necessity. Consequently, a simple alternative, CapImagine, is proposed to explicitly write visual imagination as text, outperforming complex latent-space methods on visual perception benchmarks.
- MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
-
The authors developed MentisOculi, a procedural, hierarchically difficult multi-step visual reasoning benchmark consisting of five tasks that "can only be solved via internal mental imagery." By systematically testing whether frontier models can utilize "mental imagery" to assist in reasoning like humans, the study concludes that current explicit visual strategies (latent tokens, generated images, video) fail to consistently outperform pure text baselines. More pointedly, Unified Multimodal Models (UMMs) cannot effectively utilize even ground-truth visualizations, exposing a dual bottleneck of "generation errors" compounded by "interpretation errors."
- 3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
-
3ViewSense argues that the bottleneck of VLM spatial reasoning is not insufficient visual features or weak linguistic reasoning, but the lack of a stable 3D intermediate representation. Consequently, it requires the model to first induce front, left, and top views from a single image before reasoning based on these orthographic views, significantly outperforming same-scale VLMs in occlusion counting and view-consistent spatial reasoning.
- Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning
-
This paper enforces a split in VLM output into
<recognition>perception blocks and<think>reasoning blocks. It introduces a perception reward \(R_P\) determined by whether a "blindfolded" text reasoning agent (which only sees the VLM's perception text without the image) can correctly answer the question, paired with Structured Verbal Verification (SVV) as an outcome reward \(R_O\). MoCA uses \(R_P\) as a gate for modality-level credit assignment, enabling a 7B model to improve across 9 perception/reasoning/rich-modality benchmarks simultaneously, surpassing GPT-4o on multiple metrics. - Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models
-
This work transforms VLM spatial reasoning from a "passive observation" approach into an agentic workflow that actively selects views based on questions, updates a cognitive map, and verifies reasoning using executable spatial assertions. By fine-tuning Qwen2.5-VL-3B with dense rewards, it achieves 80.5% overall accuracy on MindCube-Tiny, specifically improving the Rotation subset to 85.0%.
Browse all 31 VLM Reasoning papers →
⚡ VLM Efficiency (4)¶
- Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
-
This paper unifies Quantization-Aware Training (QAT) and Knowledge Distillation (KD) from an Information Bottleneck (IB) perspective, proposing the GRACE framework (Gated Decoupled Distillation + Relational Centered Kernel Alignment + Adaptive IB Controller). This enables INT4-quantized LLaVA / Qwen-VL to not only avoid performance degradation but outperform BF16 baselines across multiple benchmarks, while achieving 3× throughput and 54% memory savings.
- Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
-
This study evaluates 16 quantization methods across 10 VLMs and multiple reliability metrics through 700,000 experiments. It finds that quantization is not a simple disruptor—it improves calibration, OOD detection, and noise robustness by suppressing high-rank low-variance spectral components, while simultaneously amplifying reliance on covariate shifts and spurious correlations.
- On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
-
This paper presents the first systematic study of the adversarial robustness of Large Vision-Language Models (LVLMs) under visual token compression. It identifies an "optimization-inference space mismatch" in existing encoder attacks and proposes the CAGE attack. By utilizing Expected Feature Distortion (EFD) and Ranking-Distortion Alignment (RDA), CAGE significantly reduces the robust accuracy of compressed LVLMs under conditions where the compression mechanism and token budget are unknown.
- CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large Vision-Language Models
-
This work identifies an anti-intuitive "similarity reversal" phenomenon in CLIP, where visual tokens of referring regions exhibit the lowest similarity with [EOS] text tokens. Based on this observation, the authors propose LiteLVLM—a training-free, text-guided visual token pruning method. It retains 90.3% of the original pixel grounding performance even after discarding 66.7% of tokens, while achieving a 22% inference acceleration and 2.3× VRAM savings.
🎵 Audio & Speech (36)¶
- Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
-
This paper treats the decision-making of multimodal large models as an information decomposition from input to output. Using Partial Information Decomposition (PID), the mutual information of VL/omni-modal model predictions is decomposed into four terms: "Vision-unique / Text-unique / Redundant / Synergistic." It discovers that the synergistic term is the best indicator of predictive vision sensitivity and that omni-modal models suffer from a "visual hegemony" synergy bottleneck. Finally, sample-level scores derived from PID are used to guide LoRA reweighted fine-tuning, achieving consistent improvements of 1–2 percentage points on MMStar, MMBench, and POPE.
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
-
MECAT constructs 20k multi-perspective fine-grained audio captions and 100k open-ended QA pairs using a "multi-expert models + CoT LLM reasoning" pipeline. It proposes the DATE metric (harmonic mean of semantic similarity and cross-sample discriminability), achieving the first stable differentiation between generic and detailed audio model outputs.
- Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
-
This is a position paper: the authors argue that current text embedding research focuses excessively on "surface semantics" (morphology / syntax / topical similarity) while systematically ignoring "implicit semantics" such as pragmatics, stance, and social context. Empirical evidence from 7 implicit semantic datasets shows that even SOTA embeddings offer only marginal improvements over Bag-of-Tokens, advocating for implicit semantics as a first-class modeling objective in embedding research.
- MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
-
MoshiRAG incorporates a special \(\langle\text{ret}\rangle\) trigger token into the Moshi full-duplex speech model, allowing the model to asynchronously invoke an LLM or search engine backend while speaking. By exploiting the natural "keyword delay" between the start of an utterance and the appearance of critical keywords, it preserves full-duplex interactivity while hiding retrieval latencies of up to 2 seconds. This enables the model to achieve factuality on par with GPT-4o Audio across LlamaQ, WebQ, TriviaQA, and HaluEval.
- Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
-
The authors construct VoxParadox, a benchmark of 2,000 Multiple Choice Questions (MCQs) designed with intentional contradictions between "what the text says" and "what the audio sounds like." They demonstrate that current Audio LLMs almost exclusively "read but do not listen" in paralinguistic tasks. By introducing PCLM, a lightweight module that adaptively mixes intermediate audio encoder features based on the prompt, combined with DPO, they improve Audio Flamingo 3's performance on VoxParadox from 17.40% to 65.20%.
- CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
-
Addressing the lack of unified evaluation for modern music generation models that simultaneously process "text + lyrics + reference audio," this paper establishes a complete ecosystem: 110k pseudo-labeled CMI-Pref-Pseudo, 4,027 human-labeled CMI-Pref, a unified CMI-RewardBench, and a family of ~30M parameter reward models (CMI-RM) capable of handling all modality combinations in a single architecture. The authors demonstrate high correlation with human judgment and enable "inference-time scaling" via top-k filtering.
- MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety
-
MultiBreak utilizes an iterative framework of "active learning + uncertainty-guided rewriting" to expand a multi-turn jailbreak dataset to 10,389 conversations and 2,665 independent harmful intents. With a diversity score of 0.942, it significantly outperforms previous works and increases ASR on DeepSeek-R1-7B / GPT-4.1-mini by 54% / 34.6% respectively compared to the next best dataset.
- The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning
-
This paper proposes FLAIR: a framework that allows Full-Duplex Spoken Dialogue Models (SDLM) to replace the steps typically used for filling
<SIL>placeholders while "listening to the user" with continuous latent reasoning. By employing an ELBO training objective and a non-causal "global expert" to provide the posterior, the causal LLM learns to "think while listening" through a sequence of embedding vectors, significantly improving QA quality without introducing any inference latency. - SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
-
This paper introduces SafeSearch, a fully automated, sandboxed, and scalable red-teaming framework that evaluates search agent safety by injecting a single LLM-generated unreliable webpage into real search results. Through systematic evaluation of 17 LLMs across 3 agent scaffolds using 300 test cases, the study finds a peak ASR of 90.5% and demonstrates that common reminder-based defenses are largely ineffective.
- Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration
-
To address the chronic issues of "modality dominance" and "spurious modality coupling" in centralized multimodal fusion, GCL reformulates multimodal learning as a governed two-stage four-agent collaborative protocol. The first stage uses Routing/Auditing agents to decide which cross-modal communications are permitted per sample based on marginal prediction gain. The second stage uses Public-Factor/Aggregation agents to decouple shared semantics from private specializations before aggregation. This approach achieves SOTA on MOSI, MOSEI, and MIntRec.
Browse all 36 Audio & Speech papers →
🔎 AIGC Detection (11)¶
- Feature-Augmented Transformers for Robust AI-Text Detection Across Domains and Generators
-
This paper systematically exposes the vulnerability of AI text detectors under cross-dataset and cross-generator shifts using a "single-threshold fixed protocol." It proposes fusing hand-crafted linguistic features—weighted by learnable dynamic attention—with transformer [CLS] representations. Built on a DeBERTa-v3 backbone, the method achieves 85.9% balanced accuracy on the M4 multi-domain multi-generator benchmark, outperforming strong zero-shot baselines (Fast-DetectGPT, RADAR, Log-Rank) by up to +7.22.
- AutoBaxBuilder: Bootstrapping Code Security Benchmarking
-
AUTOBAXBUILDER utilizes an LLM agent pipeline to automatically generate web backend security evaluation scenarios, functional tests, and end-to-end security tests. It reduces the cost of manually constructing BAXBENCH-style tasks by approximately 12x and constructs AUTOBAXBENCH, comprising 40 new scenarios, to evaluate the gap between functional correctness and security in contemporary code models.
- Black-Box Detection of LLM-Generated Text Using Generalized Jensen-Shannon Divergence
-
SurpMark reformulates "AI text detection" as a likelihood-free hypothesis test: it uses a proxy LM to calculate token surprisal, discretizes them into \(k\) states via k-means, estimates a first-order Markov transition matrix, and compares it with pre-built "human-written / machine-written" reference matrices using Generalized Jensen-Shannon Divergence (GJS). It provides black-box, zero-retraining, and zero-per-instance-resampling discriminant scores in a single forward pass.
- Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models
-
This paper discovers that directly fine-tuning MLLMs to learn low-level artifacts left by generators damages their early-formed semantic representations (catastrophic forgetting). To address this, the authors propose Deep-VRM, which freezes the early and middle layers to preserve semantics while utilizing a LoRA-based bypass to "residually inject" artifact features into the deep layers of the LLM. This allows a single MLLM to achieve SOTA performance on most AIGI benchmarks without relying on any external expert detectors.
- Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
-
DOVE utilizes rate-distortion variational optimization to automatically construct a compact "Value Codebook" from 10,000 human texts. It then uses Unbalanced Optimal Transport (UOT) to measure distribution differences between human and LLM long-form texts in the value space, improving the "Evaluation-Downstream Task" correlation from \(\le 24\%\) in baselines to \(31.56\%\) across 12 LLMs.
- ForensicConcept: Transferable Forensic Concepts for AIGI Detection
-
Addressing the issues where AI-Generated Image (AIGI) detectors are "highly accurate within the training distribution but fail on unseen generators" and remain entirely black-box, this paper explicitly extracts dispersed evidence relied upon by detectors into a "forensic concept codebook." It uses diffusion features (CleanDIFT) as external generative trace references and employs the neighborhood-structure consistency metric CKNNA to measure the geometric alignment between backbone evidence and diffusion traces. By injecting the diffusion codebook into a target backbone, cross-generator transfer is achieved; the average accuracy on GenImage reaches 92.0%, and higher CKNNA correlates with greater transfer gains.
- CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection
-
This work redefines "multimodal fake news detection" as a task of "explicitly capturing conflicts between modalities or with world knowledge." The authors construct CAC, a corpus of 14k samples with fine-grained conflict annotations, and propose the CORE framework. CORE reshapes the conceptual boundaries of MLLMs through Conflict-Perception Training (CPT), enabling the model to significantly outperform dedicated SOTA methods on four datasets (DGM4, MDSM, MMFakeBench, NewsCLIPpings) using only 100–750 samples.
- Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection
-
Addressing the "prediction asymmetry" issue where existing AI-generated image (AIGI) detectors appear accurate but primarily classify images as real, this paper proposes DEAR. By using inpainting images as probes and "dissecting" the model based on the Regional Activation Discrepancy (RAD) between channel activations and generated areas, the method prunes extreme channels on both sides and retrains only the linear classification head. This forces the detector to discard fragile shortcut features, significantly enhancing robustness against unseen generators and post-processing.
- Generating Robust Portfolios of Optimization Models using Large Language Models
-
This paper proposes a lightweight, training-free algorithm that utilizes a single LLM to act simultaneously as a "stochastic generator" and a "scoring evaluator." By packaging candidate optimization models into a portfolio until the cumulative generation probability reaches \(1-\alpha\), it theoretically proves that as long as either the generator or the evaluator aligns with human preferences, the portfolio will contain high-quality models. Experiments on NL4LP using GPT verify that the portfolio consistently outperforms random sampling even in the worst-case scenarios.
- LLM Self-Recognition: Steering and Retrieving Activation Signatures
-
Instead of watermarking at the token level, this paper injects a random sparse steering vector into the LLM residual stream during generation, creating a detectable "activation signature." The signature is retrieved by re-feeding the text into the same model and calculating cosine similarity or using a lightweight classifier, achieving over 98% accuracy across multiple detection settings with negligible impact on text quality.
Browse all 11 AIGC Detection papers →
🧊 3D Vision (30)¶
- 4DPC\(^2\)hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping
-
4DPC\(^2\)hat is the first Multimodal Large Language Model (MLLM) designed for "dynamic point cloud sequence" (4D point cloud) understanding. The authors first use a topologically consistent construction pipeline to transform 44,000 animation assets into a dataset of 200,000 cross-modal QA pairs. Then, they employ a spatio-temporal architecture using "preserved group tokens + global tokens + bidirectional Mamba" to avoid compressing a frame into a single vector. Finally, "failure-aware bootstrapping" is used to iteratively identify incorrect model responses and synthesize targeted QA for supplementary training, enabling action understanding and temporal reasoning that significantly outperform approaches that feed video frames to static 3D models.
- EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation
-
EPS3D is the first end-to-end feed-forward open-vocabulary 3D panoptic segmentation framework. It directly predicts unified 3D panoptic Gaussians with semantic and instance attributes from unposed multi-view images in a single forward pass. By distilling 2D foundation models for supervision, it bypasses the need for 3D annotations. It introduces a semantic-instance mutual enhancement module for reciprocal calibration, achieving approximately 13% higher semantic mIoU than SOTA on Replica with an inference time of only 1 second per scene.
- Fast-SAM3D: 3Dfy Anything in Images but Faster
-
To address the slow inference speed of the SAM3D single-view 3D reconstruction model, this paper provides the first module-level latency profiling. Identifying performance bottlenecks caused by three types of heterogeneity (shape/layout dynamics, texture sparsity, and geometric spectral differences), the authors propose Fast-SAM3D. This training-free framework utilizes modality-aware step caching, spatiotemporal token carving, and spectral-aware token aggregation to achieve a 2.67× speedup at the object level with negligible quality loss, even slightly improving the reconstruction F-Score from 92.34 to 92.59.
- HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
-
HOI-PAGE enables an LLM to first "reason" precisely which body part should contact which object component, encoding this reasoning into a "Part Affordance Graph" (PAG). This PAG then drives 3D part segmentation, video diffusion, and optimization, generating 4D human-object interaction sequences for complex scenarios like "multiple people/single object" or "single person/multiple objects" without any 4D training data.
- AvAtar: Learning to Align via Active Optimal Transport
-
This paper proposes AvAtar, an active alignment framework based on Optimal Transport (OT). It quantifies the influence of candidate queries on global alignment results through gradient propagation. By utilizing the adjoint state method and conjugate gradient method, it achieves efficient solutions with linear complexity. AvAtar consistently outperforms existing active learning strategies in network alignment and cross-domain alignment tasks.
- FSI2P: A Hierarchical Focus–Sweep Registration Network with Dynamically Allocated Depth
-
This paper abstracts the human observation process of "glancing first, then examining block-by-block" into a two-stage Focus-Sweep paradigm. It replaces Transformer with Mamba for image-to-point cloud interaction and utilizes reinforcement learning to dynamically determine the number of interaction layers at each scale, achieving SOTA performance in I2P registration on RGB-D Scenes V2 and 7-Scenes.
- STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics–Physics Dual System
-
STABLE decomposes the process of "task instructions → simulation-ready tabletop scenes" into an LLM-based Semantic Reasoner (generating coarse layouts) and a flow-matching-based Physics Corrector with SDF losses (refining poses). By iterating through three stages—task-critical, important background, and secondary background—the system reduces object collisions to zero while achieving 99.0% scene alignment (AwS) on the MesaTask-10K dataset.
- TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core Optimization
-
TideGS migrates the 3DGS parameter table to an SSD, virtualizing it into "blocks" while utilizing GPU VRAM as a cache for the view-frustum visibility working set. Coupled with a three-stage asynchronous pipeline and trajectory-adaptive differential streaming, it pushes the scale of trainable Gaussians from approximately 11M (native 3DGS) or 105M (CLM) to over 1 billion on a single 24 GB GPU, achieving large-scene reconstruction quality superior to all evaluated single-GPU baselines.
- APEIRIA: Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
-
This paper proposes APEIRIA, which distills the execution traces of neuro-symbolic 3D concept learners into natural language chain-of-thought (CoC) for 3D MLLMs. By employing GRPO reinforcement learning, it generalizes these reasoning patterns to open-vocabulary and deeply nested instructions. APEIRIA simultaneously outperforms traditional NS3D methods and current state-of-the-art 3D MLLMs on ScanRefer, Multi3DRefer, SQA3D, and Scan2Cap, while retaining the interpretability and modularity of symbolic systems.
- PLAID: A Unified Data Model for Machine Learning on Heterogeneous Physics Simulations
-
PLAID proposes a unified data model and open-source library for heterogeneous physical simulation data, releasing six industrial-grade datasets covering structural mechanics and CFD alongside reproducible benchmarks. It transforms real-world "variable mesh, variable topology, and variable dimension" simulation data into standardized benchmarks accessible to the machine learning community.
Browse all 30 3D Vision papers →
🎯 Object Detection (7)¶
- HSGG: Training-Free Hierarchical Scene Graph Generation with Geometry-Guided Relation Reasoning
-
HSGG discovers entities from wholes to parts, propagates attributes from parts to wholes, and then combines geometric filtering with visually grounded predictions contrasted against geometry-induced hallucination priors, achieving 15.1 zR@100 on VG150 and 14.1 on PSG under SGDet without task-specific training (original Table 1).
- OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration
-
Aiming at the issues that multimodal visual verifiers output binary signals (True/False) that are too coarse and that textual explanations are prone to reward-hacking, this paper proposes OmniVerifier-M1. It utilizes symbolic outputs such as bounding boxes as meta-verification rationales instead of text to support rule-based rewards like IoU. Theoretically and experimentally, it proves that decoupling binary judgment and meta-verification into two independent reward streams (rather than a multiplicative joint reward) significantly improves SNR. Ultimately, the verifier is upgraded to an agentic system, M1-TTS, capable of driving region-level self-recalibration.
- EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
-
EARL utilizes a two-stage MLLM framework of "coarse interpretation and fine response" to consolidate egocentric interaction reasoning tasks (description + Q&A + pixel mask) into a unified pipeline. The first stage outputs a global interaction description of the full image and treats the last hidden state as a semantic prior. This is injected into the second stage through a novel Analysis-guided Feature Synthesizer (AFS). The system is jointly trained via GRPO with a triple-reward mechanism (format/answer/grounding accuracy), outperforming Seg-Zero by 8.37% cIoU on Ego-IRGBench.
- FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization
-
FOCUS uses a two-stage training approach—"complete removal of category names + attention mask optimization + GRPO IoU reward"—to force VLMs to perform in-context object localization based on visual support examples rather than semantic priors. The 7B parameter model outperforms 72B models, proving that task-aligned inductive bias is more important than pure scaling.
- Testing the Test: Score-Direction Instability in Class-Split Anomaly Detection
-
The authors point out that "class-split" anomaly detection benchmarks are ill-posed when the anomaly class and the normal mixture distribution overlap in the representation space—AUROC collapses to random or even reverses, with the direction depending on the unknown anomaly class. A training-free "neighborhood class leakage" metric \(L_k\) is proposed to diagnose such benchmark failure before evaluation.
- Adversarially Robust Approximate Furthest Neighbor
-
This theoretical paper provides the first approximate furthest neighbor data structure resistant to adaptive query adversaries. While maintaining a query complexity with \(n\)-dependence similar to Indyk's classical oblivious algorithm, it demonstrates that traditional random projection furthest neighbor algorithms can be broken by adaptive queries.
- Mixture Prototype Flow Matching for Open-Set Supervised Anomaly Detection
-
MPFM replaces the traditional "unimodal Gaussian prototypes" in OSAD with a learnable Gaussian Mixture Model (GMM) prototype space. It uses flow matching to directly regress a velocity field in GMM form, augmented by a mutual information maximization regularization to prevent prototype collapse. The method outperforms all SOTA methods, including DRA, AHL, and DPDL, across 9 industrial and medical AD datasets under the 10/1 anomaly sample setting.
✂️ Segmentation (14)¶
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
-
UGround flips the LMM-based visual grounding paradigm from "using the \(\langle\text{SEG}\rangle\) token of the last layer as a prompt" to "using the similarity maps of dynamically selected intermediate layers as prompts." Through a reinforcement learning strategy (SSC), the \(\langle\text{SEG}\rangle\) token slides through all transformer layers, treating the similarity map simultaneously as a soft logit mask for SAM and a backward supervision signal. This approach unifies five visual grounding tasks—RES, RS, FP-RES, gRES, and Multi-RS—within a single framework for the first time, achieving +9.0% cIoU on ReasonSeg test and +12.1% N-acc on gRefCOCO val.
- Refining Context-Entangled Content Segmentation via Curriculum Selection and Anti-Curriculum Promotion
-
CurriSeg keeps the segmentation network architecture unchanged and modifies only the training schedule: it first pushes the model to a stable state using a robust curriculum based on "temporal loss statistics + pixel entropy weighting," and then performs anti-curriculum "spectral blindness" fine-tuning (removing high frequencies to force the model to capture structural semantics). This approach consistently improves FEDER / FSEL / RUN by 2–4% on camouflaged/polyp segmentation benchmarks such as CHAMELEON / CAMO / COD10K / NC4K with zero additional parameters and shorter training time.
- SPROUT: Supervise Less, See More — Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting
-
SPROUT is the first fully training-free, zero-annotation framework for pathological nuclei segmentation. It utilizes H&E staining priors to self-construct high-confidence foreground/background regions on each slide → extracts prototypes → performs feature-prototype soft alignment via Partial Optimal Transport (POT) → outputs positive/negative point prompts for SAM. On benchmarks like MoNuSeg, its AJI is 8.2% higher than training-based methods.
- Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models
-
GPUA treats VLMs like CLIP (rich semantics, insufficient local precision) and VFMs like DINOv3 (fine-grained detail, lacking semantics) as two "visual languages." It uses Optimal Transport to mine soft correspondences and solves the Orthogonal Procrustes problem to learn a geometry-preserving linear mapping that translates VFM features into the VLM space. This process is entirely unsupervised, requires no updates to pre-trained parameters, and achieves an average 11.8% improvement in zero-shot classification.
- Unsupervised Hierarchical Skill Discovery
-
HiSD starts from unlabeled observation trajectories—performing skill segmentation via optimal transport and then discovering multi-level skill hierarchies using Sequitur grammar induction, without requiring action labels or reward signals.
- Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision Models
-
This paper constructs PolyMLP, PolyConv, and PolyAttn using Hadamard products to replace pointwise activations/softmax in MLP, convolution, and attention. Without conventional activation functions, these modules allow MetaFormer-style backbones to reach or exceed the performance of activation-based models on ImageNet, robustness benchmarks, and ADE20K segmentation.
- Beyond Detection: A Structure-Aware Framework for Scene Text Tracking
-
The authors propose SymTrack, a detection-free dual-branch scene text tracking framework. It addresses feature bottlenecks caused by perspective distortion through Predictive Token Rectification (PTR), eliminates high visual ambiguity among text instances using Cross-Expert Calibration (CEC), and stabilizes fine-grained localization with an Adaptive Inference Engine (AIE). It significantly surpasses SOTA on three benchmarks (up to +12.32% AUC).
- FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation
-
This paper points out that current query-based LLM-conditioned segmentation follows a "propose-then-select" paradigm—candidate masks are often accurate enough, but errors occur due to incorrect selection. To address this, FlowSeg is proposed, where LLM conditional embeddings participate in query refinement at every decoder layer and are continuously updated by new visual evidence. Combined with a lightweight boundary refinement module, it achieves consistent performance gains on RefCOCO/+/g and ReasonSeg.
- Functional Attention: From Pairwise Affinities to Functional Correspondences
-
This paper reinterprets softmax attention in Transformers as a "least-squares linear operator between two learned functional bases." Borrowing the idea of functional maps from shape matching, it compresses the \(n \times n\) pairwise affinity matrix into a \(k \times k\) compact spectral operator, achieving SOTA performance in PDE solving, 3D point cloud segmentation, and OOD generalization simultaneously.
- LightAVSeg: Lightweight Audio-Visual Segmentation
-
LightAVSeg decouples "semantic filtering (what)" and "spatial localization (where)" by replacing \(\mathcal{O}(N^2)\) cross-modal attention with global channel modulation. This allows the AVS model to achieve 50.4 mIoU (MS3) with only 20.5M parameters and reach an on-device latency of 163.4 ms on Snapdragon 8 Elite, which is approximately \(8\times\) faster than AVSegFormer-R50.
Browse all 14 Segmentation papers →
🖼️ Image Restoration (21)¶
- Plan for Speed: Dilated Scheduling for Masked Diffusion Language Models
-
This paper proposes the Dilated Unmasking Scheduler (DUS): it uses predefined "equidistant gaps" to determine the unmasking order independent of model confidence. This reduces the number of denoiser calls per block of \(B\) tokens from \(\mathcal O(B)\) to \(\mathcal O(\log B)\), achieving a 5.8× wall-clock speedup on LLaDA / Dream / DiffuCoder while outperforming confidence-based parallel planners in quality.
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
-
This paper systematically compares continuous diffusion, discrete masked diffusion, and looped transformers across the dimensions of expressivity and trainability. It proves that "continuous diffusion" is strictly more expressive than discrete diffusion and can simulate looped transformers, but its practical performance is limited by decoding and representation space. Consequently, the paper proposes CCDD (Coevolutionary Continuous Discrete Diffusion)—diffusion performed simultaneously on the discrete token space and the contextual embedding space of a pre-trained LLM, with a single model for joint denoising. CCDD reduces perplexity by 25-35% compared to MDLM on LM1B/OWT and outperforms MDLM with 256 steps using only 8 sampling steps.
- DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs
-
DAPD transforms the single-step parallel unmasking problem of dLLMs into a dynamic graph coloring problem of "selecting independent sets on self-attention-induced MRFs." Without any training, it simultaneously unmasks weakly dependent positions, reducing decoding steps to 1/3.87 of the original on LLaDA / Dream for multi-question mixed prompts with almost no loss in accuracy.
- Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models
-
This paper systematically characterizes the failure mechanism of masked diffusion language models (dLLM) under fully non-autoregressive (NAR) decoding. It identifies that proximity bias causes confidence-based sampling to degenerate into reverse autoregrssion, which is prematurely saturated by EOS tokens. By using a 5M-parameter lightweight planner and EOS temperature annealing to intervene in unmasking positions only at the first step, the authors improve LLaDA 8B NAR decoding by 2.8–4.3 points on reasoning tasks like GSM8K with almost no additional overhead.
- DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
-
DyLLM is a training-free inference acceleration framework for diffusion LLMs. It identifies "salient tokens" by measuring the cosine similarity of attention contexts between adjacent denoising steps. By recalculating FFN and attention only for these tokens using salient-aware approximate attention, it increases throughput to 7.6× / 9.6× on LLaDA / Dream with negligible performance loss.
- From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion
-
Multimodal image fusion has long relied on shared representations in 2D feature grids, leading to the entanglement of global appearance (brightness/contrast/tone) and local details, making them difficult to regulate independently. This paper moves "global appearance" into the compact token space of a frozen 1D tokenizer (TiTok-32). By employing "Selective Token Editing (STE)" to modify only a few token-channel entries, the method regulates global consistency while preserving a 2D pathway for detail recovery, achieving comprehensive SOTA results across four benchmarks.
- One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation
-
To address the slow inference of diffusion SR models, this paper proposes RSD (Residual Shifting Distillation): distilling a 15-step ResShift teacher into a one-step student generator. The core mechanism involves "training the student so that a 'fake ResShift' trained on its output exactly matches the true teacher"—which is equivalent to matching the joint distribution (rather than just marginals as in VSD) of the teacher and student across all timesteps. Consequently, RSD outperforms the teacher and the comparable distillation method SinSR on LPIPS / CLIPIQA / MUSIQ. With only 174M parameters, 0.5GB VRAM, and 5 GPU-hours of training, it approaches the perceptual quality of massive T2I-based SR models.
- Degradation-Aware Metric Prompting for Hyperspectral Image Restoration
-
DAMP utilizes 6 interpretable spatial-spectral physical metrics (high-frequency energy ratio, texture uniformity, spectral curvature, etc.) as "Degradation Prompts" (DP) to replace black-box embeddings and explicit degradation labels. These DPs act as gating signals driving a Spatial-Spectral Adaptive MoE to select different "spatial/spectral experts," achieving SOTA performance across 5 HSI restoration tasks and 2 unseen degradations (motion blur, Poisson noise) simultaneously.
- PODiff: Latent Diffusion in Proper Orthogonal Decomposition Space for Scientific Super-Resolution
-
PODiff moves the diffusion process from pixel space to a fixed, variance-sorted POD coefficient space. By utilizing a minimal MLP, it achieves accuracy comparable to pixel-level diffusion on \(640\times 480\) SST downscaling tasks. Since reconstruction is linear, ensemble variance can be analytically back-propagated to physical space via \(\Sigma_u=\Phi\Sigma_a\Phi^\top\), yielding spatially interpretable and well-calibrated uncertainty.
- Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution
-
ASASR achieves the optimal balance between perceptual quality and structural fidelity in super-resolution by replacing the Flow Matching noise prior from isotropic Gaussian to Sobolev spectral coloring noise, combined with adversarial manifold guidance to generate hard negative samples, constructing the AS-DPO framework.
Browse all 21 Image Restoration papers →
🛰️ Remote Sensing (3)¶
- Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
-
Any2Any transforms remote sensing (RS) translation between RGB, SAR, NIR, MS, and PAN from a collection of paired models into a unified latent diffusion model within a shared latent space. By utilizing the million-level RST-1M dataset and target modality residual adapters, it achieves superior fidelity and generalization across 14 seen translation directions and multiple unseen modality combinations.
- Localized, High-resolution Geographic Representations with Slepian Functions
-
This paper constructs a geographic positional encoder that concentrates representation capacity on a Region of Interest (ROI) using spherical Slepian functions. It proposes a Slepian-Spherical Harmonic (SH) hybrid encoding to simultaneously capture local high-resolution details and global coarse-grained context. It consistently outperforms mainstream baselines such as SH, Wavelets, and RFF across five classification, regression, and image-enhancement prediction tasks.
- The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench
-
The authors observe that Vision Foundation Models (VFMs) "appear" to predict satellite imagery well but collapse along physical axes in extreme regimes. By formalizing "scientific alignment" as "structural isomorphism," they release TC-Bench—a global tropical cyclone benchmark—and a three-tier linear probing suite (Static/Dynamic/Constraint) to reveal representation collapse in frozen backbones like DINO, CLIP, SigLIP, and MAE for intense cyclones where \(P_c<980\) hPa.
🧑 Human Understanding (5)¶
- WaveVerse: Scalable RF Simulation in Generative 4D Worlds
-
WaveVerse integrates LLM-driven "4D indoor scene + human motion" generation with a physical ray tracer that preserves spatiotemporal phase coherence into a prompt-to-RF signal pipeline. It significantly enhances downstream RF imaging and activity recognition tasks using synthetic data, with performance scaling continuously as simulation volume increases, unlike existing methods that saturate.
- DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing
-
DiscoForcing reformulates the "music \(\to\) full-body dance" offline generation problem into a strictly causal, bounded-latency streaming task. It utilizes a VQ-PAE causal music encoder, latent-space Diffusion Forcing, hybrid temporal noise scheduling, and Temporal Guidance sampling to translate music streams into 30 FPS full-body motions that directly drive Unity avatars and Unitree G1 humanoid robots in real-time.
- Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
-
Ours proposes Intrinsic Quality (IQ): after extracting embeddings using a proxy model, it weightedly fuses "Neighborhood Label Consistency (Consis)" and "Normalized Spectral Entropy Effective Rank \(\tilde{r}_{\mathrm{ent}}\)". It provides a "trainability" score for million-scale face recognition datasets without full training or clean validation sets. On WebFace4/12/42M and noise-injected settings, the ranking consistency with downstream MFR-ALL validation accuracy reaches Spearman = 1.0.
- Learning Instance-Adaptive Low-Rank Orthogonal Subspaces for Clothes-Changing Person Re-Identification
-
The "clothing" semantic concept is explicitly modeled as an instance-adaptive low-rank subspace (initialized using the SVD principal components of CLIP text descriptions and refined via cross-attention with image patches). Identity features are then forced to be strictly orthogonal to this subspace through geometric constraints, achieving SOTA results in clothes-changing re-identification (PRCC +5.9% Rank-1) without the need for adversarial training.
- MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery
-
MotionGRPO reformulates first-person full-body motion recovery from head-mounted devices as a Markov Decision Process (MDP) over diffusion sampling. It utilizes Group Relative Policy Optimization (GRPO) post-training with a hybrid reward system consisting of a "trajectory condition-aware perception model + 4 joint-level sub-rewards." Crucially, it identifies that strong input conditions lead to nearly identical intra-group samples, causing advantage variance to vanish—a fatal bottleneck. To resolve this, it injects Perlin noise into the conditioning signal to restore intra-group diversity, reducing MPJPE from EgoAllo's 124.985 mm to 114.207 mm on AMASS/RICH.
📹 Video Understanding (17)¶
- Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
-
Video-MTR is an RL-based multi-turn reasoning framework that guides MLLMs to iteratively select key video segments through a gated dual-level reward mechanism. It achieves SOTA performance in long video understanding using only 8K data, outperforming methods that require 257K to 4.4 million samples (improving data efficiency by two orders of magnitude).
- OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
-
This paper argues that existing Omni-LLM token compression methods are suboptimal due to their "symmetric" treatment of audio and video. It proposes OmniSIFT—a two-stage asymmetric compression framework that first prunes video redundancy via spatio-temporal saliency to obtain "visual anchors," which then guide audio selection. With only 4.85M additional parameters, it consistently outperforms existing baselines and even the original model on Qwen2.5-Omni-7B while retaining only 25% of tokens.
- SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling
-
SLAP transplants the "Least Action Principle of classical mechanics" onto the video semantic manifold, modeling the completion of missing frames in sparsely sampled videos as a two-point boundary value problem on a Riemannian manifold. By replacing probabilistic generation with semantic dynamics to enforce object permanence, it achieves 83.9% accuracy on tunnel occlusion tests (outperforming diffusion models by 12 points) with a 177× inference speedup.
- ProAct-VL: A Proactive VideoLLM for Real-Time AI Companions
-
ProAct-VL enables VideoLLMs to autonomously decide when to respond and generate short-segment commentary under streaming input via a chunk-level I/O paradigm, a lightweight FLAG decision head, and transition-aware loss functions. It achieves ~1s low latency and strong proactivity—obtaining a TimeDiff of only 1.20s and a trigger F1 of 63.25% in game commentary tasks, significantly outperforming offline models like GPT-4o.
- VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking
-
VideoTemp-o3 is a unified agentic video understanding framework. By joint modeling of temporal grounding and video QA through a unified masking strategy for cold-start SFT and a penalty-aware IoU reward, it achieves high-quality multi-round iterative grounding and precise answering in long video understanding. It reaches an mIoU of 15.6% on ultra-long videos (> 20 minutes), surpassing Gemini-2.5-Pro's 14.8%.
- VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
-
VideoSEAL identifies the "evidence misalignment" problem in existing agentic long video QA systems—where agents answer correctly without actually seeing the evidence—and attributes the root cause to "coupled agents conflating planning and answering authority." It proposes a planner-inspector decoupling framework: the planner handles long-horizon evidence search, while the inspector holds exclusive answering authority and only releases the answer when pixel-level evidence is sufficient. This improves accuracy on LVBench from 48.2% to 55.1% (↑20.5%) and on LongVideoBench from 52.2% to 62.0%.
- RELO: Reinforcement Learning to Localize for Visual Object Tracking
-
RELO reformulates the "where is the target" problem in single object tracking as an MDP on a spatial feature map. It treats each spatial position as an action and replaces traditional manual center heatmap supervision with actor-critic + direct IoU/AUC rewards. Coupled with two stabilization designs—"regression warmup" and "layer-aligned temporal token propagation"—it achieves SOTA with 57.5% AUC on LaSOText.
- AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes
-
This paper proposes the AVTrack dataset and the AVTracker baseline method to address the Audio-Visual Instance Segmentation and tracking (AVIS) task in complex human-centric scenes. By defining eight challenging conditions, a rigorous evaluation benchmark was constructed. A three-stage divide-and-conquer framework was designed (ASR segmented aggregation → local speaker localization → global identity association), which outperforms existing state-of-the-art methods by approximately 8 percentage points on the HOTA metric.
- Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning
-
Foresee-to-Ground (F2G) reformulates Video Temporal Grounding (VTG) from direct timestamp regression into an "identify-then-measure" two-stage problem. By utilizing predictive temporal perception and a span evidence encoder to build a candidate event evidence pool, the LLM generates precise boundaries constrained by selected events. This approach improves [email protected] by 4.1 points on Charades-STA and 6.7 points on ActivityNet.
- Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
-
To address query ambiguity and temporal sparse supervision caused by "short queries vs. long videos" in Partially Relevant Video Retrieval (PRVR), this paper proposes Holmes, a hierarchical evidential learning framework based on the Dirichlet distribution. It distinguishes precise, polysemous, and under-determined queries using a three-fold principle at the inter-video level for adaptive label calibration, and achieves dense alignment at the intra-video level via flexible optimal transport with a dustbin. Holmes achieves SOTA on ActivityNet, Charades, and TVR datasets.
Browse all 17 Video Understanding papers →
🚗 Autonomous Driving (8)¶
- TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models
-
TSRBench constructs a time series reasoning benchmark covering 14 domains, 4 major dimensions (Perception/Reasoning/Prediction/Decision-making), 15 tasks, and 4,125 questions. It supports four input modalities (Text, Visual, Text+Image, Embedding) and systematically evaluates 30+ mainstream LLMs, VLMs, and TSLLMs. It reveals that "scaling holds in perception/reasoning but fails in prediction" and that "text and visual modalities are highly complementary, yet current models struggle to fuse them."
- RoCA: Robust Cross-Domain End-to-End Autonomous Driving
-
RoCA attaches a plug-and-play module based on Gaussian Processes (GP) to end-to-end autonomous driving models. By learning a set of basis tokens and corresponding trajectories that cover diverse scenarios, it probabilistically infers future trajectories based on similarity for new scenarios. This approach uses GP uncertainty for regularization to enhance generalization during source domain training and enables efficient adaptation via pseudo-labels and active learning in new domains, without requiring LLMs or increasing inference overhead.
- CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving
-
CoIRL-AD utilizes two independent actors to handle Imitation Learning (IL) and Reinforcement Learning (RL) respectively, relying on a latent world model to "imagine" future trajectories for calculating long-range rewards for RL. A "leader-follower" competitive mechanism allows both actors to transfer beneficial behaviors to each other. This approach successfully integrates RL into end-to-end driving using offline real-world driving data without an external simulator, achieving significant improvements in cross-city generalization and long-tail scenarios.
- Mitigating Error Accumulation in Continuous Navigation via Memory-Augmented Kalman Filtering
-
This work reformulates step-by-step prediction in continuous UAV VLN as a closed-loop "recursive Bayesian estimation = GRU prior + memory likelihood + learnable Kalman gain." By fine-tuning on only 10% of the data in TravelUAV, the Success Rate (SR) of L1-Full is improved from 17.6% to 25.9%, while the positional drift—which typically accumulates continuously after 100 steps—is flattened to 30–40 meters.
- Constrained Multi-Objective Reinforcement Learning with Max-Min Criterion
-
This paper unifies "max-min multi-objective fairness" and "hard constraint satisfaction" into a single MORL framework. By reformulating the problem as a convex program via occupancy measures, the authors derive a dual convex optimization problem over weights \((u,w)\). This allows a projected gradient descent algorithm to simultaneously achieve fairness and constraint feasibility with theoretical guarantees of geometric convergence.
- DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving
-
DeepSight shifts "future world prediction" from explicit pixel reconstruction (single-frame codebook) to parallel implicit multi-frame prediction of DINOv3 semantic features in BEV space. Combined with an on-demand Adaptive Chain-of-Thought, it achieves a Driving Score of 86.23 (+7.39) and a Success Rate of 71.36% (+13.63) on the Bench2Drive closed-loop benchmark while adding only ~4% inference latency.
- Plug-and-Play Label Map Diffusion for Universal Goal-Oriented Navigation
-
This paper proposes PLMD: a framework that merges BEV semantic and obstacle maps into a unified Label Map. It utilizes DDPM, modulated by obstacle priors, to complete semantic and obstacle labels in unexplored regions. As a plug-and-play module, it can be integrated with any GON policy and consistently achieves new SOTA results on HM3D/MP3D across three tasks: ON, IIN, and MRON.
- Threshold-Based Exclusive Batching for LLM Inference
-
This paper systematically characterizes the performance crossover conditions between mixed batching (MB) and exclusive batching (EB) in LLM inference. It proves that on bandwidth-constrained GPUs, co-batching prefill and decode stages slows down Attention due to bandwidth contention. Consequently, the authors derive an optimal phase-switching threshold \(\theta^*\) and a memory-safe batch size based on the hazard rate, designing an online adaptive scheduler EB+. This scheduler improves throughput by up to 41.9% on bandwidth-constrained hardware and up to 36.4% under non-stationary traffic compared to MB.
🤖 Robotics & Embodied AI (53)¶
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
-
This paper shifts VLA action decoding from Autoregressive (AR) or external continuous diffusion heads to "masked diffusion on discrete action tokens within a unified Transformer." Combined with adaptive parallel decoding ranked by confidence and secondary re-masking for error correction, it achieves a 96.4% average success rate on LIBERO and a 64.1% total mean score on SimplerEnv-Fractal. Notably, performance degrades by only 0.8% / 20.4% under OOD language/visual perturbations, significantly outperforming continuous diffusion and parallel decoding baselines while preserving the multimodal priors of the pre-trained VLM.
- RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
-
RoboMME is the first to systematically map "temporal/spatial/object/procedural" memory from human cognition to 16 long-horizon robotic manipulation tasks (770k high-quality timesteps). By performing a systematic ablation of 14 "memory representation × integration method" combinations on a π0.5 base, it concludes that "Perceptual Memory + AdaLN Modulator" currently offers the best comprehensive trade-off.
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
-
The authors observe that VLA inference is compute-bound, making pruning the optimal acceleration path. Given the high overlap of visual information between consecutive action steps, they propose SpecPrune-VLA. This training-free method uses a three-way fusion (previous global attention + current early-layer local attention + frame-difference dynamic tokens) for static pruning, combined with intra-layer dynamic pruning and a velocity-aware coarse/fine granularity controller. It achieves 1.57× speedup on LIBERO and 1.70× on real robots with negligible success rate loss.
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
-
DUST utilizes a "dual-stream" multi-modal diffusion Transformer (MMDiT) to process action flows and future visual embedding flows in parallel. By employing shared attention for cross-modal fusion, combined with independent noise scheduling and asynchronous action-vision sampling, it enables the VLA to simultaneously learn "what actions to perform" and "what consequences those actions produce." It consistently outperforms GR00T-N1.5+FLARE on RoboCasa, GR-1, and real-world Franka robots.
- Mixture of Horizons in Action Chunking
-
Addressing the "long-horizon planning vs. short-horizon precision" trade-off caused by "action chunk length (horizon) selection" in VLA models, this paper proposes Mixture of Horizons (MoH). By decomposing a single action chunk into various sub-chunks of different lengths, predicting them in parallel using a shared action transformer, and fusing them with a 2k-parameter linear gate—complemented by a load-balancing loss and dynamic inference via "cross-horizon consensus"—the authors enable \(\pi_{0.5}\) to reach a 99% average success rate on LIBERO for the first time while increasing throughput to 2.5× the baseline.
- LangForce: Bayesian Decomposition of Vision-Language-Action Models via Latent Action Queries
-
LangForce formulates the VLA policy as a Bayesian decomposition \(\pi(a\mid v,\ell)=p(\ell\mid a,v)\,p(a\mid v)/p(\ell\mid v)\). By introducing learnable Latent Action Queries, it executes both "vision-only" and "vision+language" branches using a single set of VLM weights. It explicitly penalizes "visual shortcuts" by maximizing the log-likelihood ratio between actions and instructions, achieving an 11.3 percentage point absolute improvement over the QwenGR00T baseline on SimplerEnv.
- Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
-
LaRA-VLA internalizes both textual and visual Chain-of-Thought (CoT) into continuous latents. Through a three-stage curriculum training process (explicit CoT → latent replacement → action expert adaptation), reasoning is performed within the latent space. This reduces inference latency by up to 90% compared to explicit CoT, restoring control frequencies to real-time ranges.
- Contrastive Representation Regularization for Vision-Language-Action Models
-
The authors observe that representations in VLA models inherited from VLMs are dominated by visual appearance and are insensitive to robot proprioceptive states. They propose Robot State-aware Contrastive Loss (RS-CL), which uses the Euclidean distance between proprioceptive states as "soft contrastive labels" to reshape representations. Combined with "view cutoff" feature-level augmentation, this method achieves a SOTA success rate of 69.7% on RoboCasa-Kitchen using GR00T N1.5 and improves success rates from 45.0% to 58.3% on real-world Franka pick-and-place tasks.
- From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model
-
BehaviorVLA utilizes a causal triple-stream Mamba encoder (VBE) to compress long-horizon demonstrations into a time-invariant "behavioral prototype \(z_{\text{proto}}\)" and a time-variant "phase state \(z_{\text{phase}}\)". A phase-conditioned behavior decoder (PBD) then expands the behavioral skeleton into phase-aligned Gaussian priors via a Predictor-Corrector mechanism to guide the flow matching strategy. It sets new SOTA benchmarks on LIBERO, RoboTwin 2.0, and CALVIN, matching OpenVLA-OFT performance using only 50% of real-world data.
- TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
-
TimeRewarder formalizes "task progress" as the normalized temporal distance between video frame pairs. It trains a self-supervised ViT distance regressor using only action-free expert videos and provides the predicted distance between adjacent frames as a dense reward to DrQ-v2. On 10 Meta-World tasks, it approaches a 9/10 success rate within 200K interactions, even outperforming manually designed environmental dense rewards.
Browse all 53 Robotics & Embodied AI papers →
🎮 Reinforcement Learning (110)¶
- Dr. Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
-
Dr. Tulu proposes RLER (Reinforcement Learning with Evolving Rubrics), allowing evaluation rubrics to co-evolve with the policy during training. This extends RLVR from short-form QA to long-form deep research tasks with citations. Ultimately, DR Tulu-8B, trained from Qwen3-8B, outperforms Tongyi DR-30B by an average of 15.6 points across four long-form deep research benchmarks and reaches competitive performance with OpenAI Deep Research at a 1000x lower cost.
- Agent Learning via Early Experience
-
This paper proposes the "early experience" paradigm, which allows language agents to utilize the future states of their own actions to learn environment dynamics and decision-making reflections without external rewards. This approach consistently outperforms pure imitation learning across 8 agent environments and provides a superior initialization for subsequent GRPO reinforcement learning.
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
-
RLVE transforms language model RL training from "static problem sets" into 400 programmable verifiable environments where problems are algorithmically generated and rewards are verified via code. By adaptively increasing problem difficulty as the policy model improves, the training signal is kept at the frontier of model capability; on a strong 1.5B model already saturated by standard RLVR, RLVE achieves a \(3.37\%\) average gain across six reasoning benchmarks using only 1/3 of the compute (compared to a \(0.49\%\) gain from continued standard RL).
- Learning Unmasking Policies for Diffusion Language Models
-
This paper explicitly models the decoding process of masked diffusion language models (dLLMs) as an MDP. Using GRPO, it trains a single-layer Transformer policy—comprising less than 0.01% of the base model's parameters and taking only token confidence as input—to adaptively decide which positions to unmask at each step. In the semi-AR setting, it matches manual heuristics like Fast-dLLM; in the full-diffusion setting, it significantly outperforms them and demonstrates transferability across models, tasks, and sequence lengths.
- d2: Improving Reasoning in Diffusion Language Models via Trajectory Likelihood Estimation
-
This paper proposes the d2 reinforcement learning framework for masked diffusion language models (masked DLM). The core contribution is the introduction of two "trajectory likelihood estimators": d2-AnyOrder, which provides exact single-forward estimates for any-order models, and d2-StepMerge, which provides adjustable-precision approximations for standard MDMs. This framework enables the correct implementation of GRPO, allowing LLaDA-8B-Instruct to achieve 91.9% / 56.6% / 85.0% / 41.6% on Sudoku/Countdown/GSM8K/MATH500 respectively, significantly outperforming diffusion RL baselines like d1 and wd1.
- Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory
-
BudgetMem reorganizes "runtime agent memory extraction" into a modular pipeline consisting of "filtering → parallel entity/temporal/topic extraction → summarization." Each module is equipped with LOW/MID/HIGH budget tier interfaces. A shared lightweight router is trained via PPO to select tiers for each module upon the arrival of a query, simultaneously improving F1/Judge scores and reducing the average cost per query on LoCoMo, LongMemEval, and HotpotQA.
- The Shape of Reasoning: Topological Analysis of Reasoning Traces in Large Language Models
-
This paper treats LLM chain-of-thought as a "point cloud" in embedding space. It uses Topological Data Analysis (TDA) to extract persistent homology features as an objective measure of reasoning quality. Experiments on the AIME dataset demonstrate that TDA features significantly outperform traditional graph statistics in predicting Smith-Waterman alignment scores (average \(R^2=0.236\) vs. average \(R^2=0.064\)).
- Flow-Equivariant World Models: Memory for Partially Observed Dynamic Environments
-
FloWM maintains structured dynamic memory in latent space by leveraging time-parameterized symmetries (flow equivariance). This solves the problem of objects "disappearing" after moving out of bounds in partially observed environments, achieving long-horizon prediction accuracy far exceeding diffusion and recurrent baselines (SSIM 0.9525 vs. DFoT 0.8885 in 3D Block World 210-step prediction).
- EchoRL: Reinforcement Learning via Rollout Echoing
-
This paper identifies that in the late stages of RLVR training, GRPO-style methods suffer from "advantage degeneration"—where vanishing gradients occur because a group of rollouts all achieve success. The authors propose EchoRL: it identifies the "hardest yet successful" prefix, termed EchoClip, based on step-level entropy peaks from verified-success rollouts. This is added to the loss as an auxiliary SFT term, consistently delivering improvements of up to 5.6% ID and 5.0% OOD across 4 RLVR frameworks, 5 backbones, and 10 benchmarks.
- Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies
-
Addressing the core challenge that "no direct target policy samples exist in online RL," this paper proposes Reverse Flow Matching (RFM). By transforming the training of diffusion/flow policies to fit Boltzmann distributions into a "posterior mean estimation given intermediate noise" problem, it uses Langevin Stein operators to construct zero-mean control variables. This unifies existing "noise expectation" and "gradient expectation" methods into a single family of estimators. Consequently, it enables flow policies (not just diffusion policies) to sample Boltzmann distributions for the first time, achieving more stable and superior performance on continuous control benchmarks compared to diffusion baselines.
Browse all 110 Reinforcement Learning papers →
🎁 Recommender Systems (11)¶
- Can Recommender Systems Teach Themselves? A Recursive Self-Improving Framework with Fidelity Control
-
RSIR enables sequential recommendation models to generate new synthetic user interaction sequences using their own predictive capabilities, train a new model, and filter out samples deviating from the user preference manifold using a rank-based "fidelity check" to prevent self-consuming model collapse. It consistently improves NDCG/Recall by 4–11% across 4 datasets and 3 mainstream backbones, theoretically proving that this process is equivalent to implicit regularization along the tangent space of the user preference manifold.
- T-POP: Test-Time Personalization with Online Preference Feedback
-
T-POP integrates "test-time alignment" with "neural dueling bandits." Without modifying LLM parameters, it learns a personalized reward function online using pairwise preference feedback per round, effectively addressing the cold-start problem in personalization for new users.
- Position: Neglecting the Sustainability of AI is Fuelling a Global AI Arms Race
-
Utilizing Karl Marx's "base-superstructure" framework, this position paper argues that current "sustainable AI" discussions are dominated by environmental dimensions while neglecting economic and social ones. It calls for the simultaneous elevation of both climate awareness and resource awareness axes and proposes the CARAML five-layer action framework (Individual / Community / Industry / Government / Global) to curb the escalating "global AI arms race."
- RGMem: Renormalization Group-Inspired Memory Evolution for Language Agents
-
RGMem draws inspiration from the Renormalization Group (RG) in statistical physics to model the long-term dialogue memory of language agents as a multi-scale system ("Event Layer → Relation Layer → Concept Layer"). It employs threshold-triggered non-linear operators to coarse-grain fragmented dialogues into stable user profiles, thereby breaking the "stability vs. plasticity" trade-off.
- Incentivized Exploration with Stochastic Covariates: A Two-Stage Mechanism Design for Recommender System
-
RCB integrates "exploration-exploitation" and "user incentive compatibility" into a contextual bandit problem under Dynamic Bayesian Incentive Compatibility (DBIC) constraints. It proposes a two-stage algorithm (Cold Start + IPGS), proves \(\tilde{O}(\sqrt{KdT})\) regret in stochastic user covariate scenarios, allows for the integration of any offline learning oracle, and quantifies the "incentive price" — showing that the cold start sample size grows as \(1/\epsilon^2\) as the \(\epsilon\) constraint tightens.
- Position: Stop Preaching and Start Practising Data Frugality for Responsible Development of AI
-
This position paper points out that the ML community has long been "preaching without practicing" regarding "data frugality"—while verbally acknowledging that coresets save energy, almost no one actually reports energy consumption or carbon emissions. Using ImageNet-1K as a case study, the authors calculate a conservative lower bound of approximately 5.82 GWh / 2589 tCO2e for downstream training and storage, calling for data frugality to evolve from a slogan into a measurable, actionable, and rewardable engineering practice.
- A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving
-
This paper treats the batch condition in LLM serving as a treatment variable for safety evaluation. It proposes a testing protocol consisting of safety-capability paired comparisons, scorer/human adjudication, cross-model expansion, continuous batching composition, and batch-invariant kernel ablation. The study concludes that refusal flips are real but low-frequency, model-specific, and dependent on the specific serving stack.
- GCIB: Graph Contrastive Information Bottleneck for Multi-Behavior Recommendation
-
GCIB employs a dual approach of "Graph Information Bottleneck + Cross-behavior Contrastive Learning." It first prunes edges in auxiliary behavior graphs that are irrelevant to the target task at the structural level (maximizing mutual information with the target behavior and minimizing mutual information with the original auxiliary graph via HSIC surrogates). It then aligns denoised auxiliary representations with sparse target representations using InfoNCE at the feature level, achieving a 7%–40% relative improvement in HR@10 / NDCG@10 across four multi-behavior recommendation benchmarks.
- Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design
-
SkillPCF reformulates the inverse design of Photonic Crystal Fibers (PCF) as a "memory policy learning" problem. A PPO-trained controller selects Top-K memory operations from an evolvable skill library for each trajectory span. An executor implements these in trajectory memory, while MEEP electromagnetic simulation rewards simultaneously optimize both the controller and the skill library. This approach achieves a superior trade-off between design success rate and simulation budget compared to multiple LLM backends and classical optimization baselines.
- Prompts for Public-Sector LLMs Should Be Governed as Commons
-
This is a position paper: the authors argue that LLM prompts used by the public sector should be versioned, provenanced, auditable, and vetoable like open-source commons. Based on a pilot benchmark using 443 neighborhood prompts from a North American city (augmented to 3,317) across five governance states, it provides three falsifiable predictions—governed prompts change output distributions, improve auditability, and shorten fault-remediation latency.
Browse all 11 Recommender Systems papers →
🔄 Self-Supervised Learning (28)¶
- Data Augmentation of Contrastive Learning is Estimating Positive-incentive Noise
-
The authors prove that "predefined data augmentation (rotation/cropping/flipping)" in contrastive learning is equivalent to a point estimation of Positive-incentive Noise (π-noise). They then upgrade π-noise from "point estimation" to a learnable distribution by training a π-noise generator (PiNDA) to add learnable noise as augmentation. This leads to consistent gains for SimCLR / BYOL / SimSiam / MoCo / DINO in vision and is naturally compatible with non-visual data without manual augmentation, such as HAR / Reuters / Epsilon.
- From Zero to Hero: Advancing Zero-Shot Foundation Models for Tabular Outlier Detection
-
This paper proposes OutFormer, a tabular Prior-Fitted Network (PFN) pretrained on a mixture of three synthetic priors (GMM, SCM, and Copula) and stabilized through a Multi-Armed Bandit-based Self-Evolving Curriculum. It achieves zero-shot tabular outlier detection by processing training data in-context and generating labels in a single forward pass. OutFormer achieves SOTA rankings across ADBench and two new benchmarks containing 1500+ datasets, while maintaining inference latency comparable to shallow models.
- LimiX-2M: Mitigating Low-Rank Collapse and Attention Bottlenecks in Tabular Foundation Models
-
Addressing two major pathologies in tabular foundation models like TabPFN-v2—severe low-rank collapse in shallow layers and the negligible contribution of sample-attention in the final layer to prediction signals—the authors propose using Radial Basis Functions to expand each scalar into a set of local responses (RaBEL) to unlock degrees of freedom in the "value direction." Furthermore, the bidirectional attention blocks are rearranged from F→S→N to S→N→F to ensure all attention paths flow into the readout. With only 2M parameters, this model consistently outperforms the 7M TabPFN-v2 and 27M TabICL across mainstream tabular benchmarks.
- Towards One-for-All Anomaly Detection for Tabular Data
-
OFA-TAD is proposed: using "neighbor distance" as a cross-domain universal anomaly cue, multi-view distance representations are extracted from metric spaces induced by various feature transformations. These are adaptively fused using a Mixture of Experts (MoE) gating mechanism. After a single training phase, the model generalizes directly to unseen tabular datasets for anomaly detection without any target-domain fine-tuning.
- LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
-
Addressing the long-standing issue in LLM selective prediction where "UCB risk bounds are too conservative and offer few usable thresholds," the authors rewrite the objective "post-selection error rate \(\le \alpha\)" as a linear expectation constraint involving indicator functions for selection and error. This leads to a finite-sample sufficient condition (Eq. 5) that depends only on the calibration set. This approach maintains strict finite-sample guarantees while being significantly tighter than UCB. The framework naturally extends to two-model routing systems for joint threshold calibration, achieving consistent power gains across CommonsenseQA, TriviaQA, ScienceQA, and MM-Vet v2, and accepting 9.5% more samples than Clopper-Pearson UCB on TriviaQA.
- PartCo: Part-Level Correspondence Priors Enhance Category Discovery
-
PartCo introduces a plug-and-play framework to enhance Generalized Category Discovery by explicitly leveraging part-level feature correspondences inherent in Vision Transformer patch tokens, improving baselines like SimGCD / SPTNet / FlipClass by 2-10% across multiple benchmarks including CUB, Stanford-Cars, and ImageNet-100.
- Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts
-
The authors propose CaRE: inserting a Bi-Level Routing MoE (BR-MoE) into each ViT block. It first uses a "class-perceiver" to select Top-M relevant task routes based on entropy, then each route activates Top-K task experts while adding a shared EMA expert. This allows the model to retain old knowledge while absorbing new classes even in sequences exceeding 300 tasks. The work also introduces the 1000-class OmniBenchmark-1K to fill the gap in long-sequence CIL evaluation.
- Zero-Flow Encoders
-
The paper discovers a counter-intuitive phenomenon: a rectified flow trained with independent coupling is zero at \(t=0.5\) if and only if the source and target distributions are identical ("Zero-Flow Criterion"). By generalizing this to conditional distributions, the authors prove that \(\mathbf{v}_{t=0.5}=0\) is equivalent to the encoder \(f(Y)\) being sufficient for predicting \(X\) (conditional independence). Based on this, they design a simulation-free least-squares loss without parametric density assumptions to unifiedly learn Markov blankets in graphical models and self-supervised representations, naturally circumventing the "shortcut problem" inherent in contrastive learning.
- NITP: Next Implicit Token Prediction for LLM Pre-training
-
NITP provides continuous representation-space supervision for the final hidden states by using shallow representations as implicit targets. This supplements standard NTP to prevent hidden representations from degenerating into low-dimensional anisotropic configurations, achieving a 5.7% improvement in MMLU-Pro on a 9B MoE and general gains of 4-6% in reasoning tasks with only ~2% additional computational overhead.
- Riemannian Metric Matching for Scalable Geometric Modeling of Distributions
-
The "Riemannian metric of the data manifold" is rewritten as a carré du champ operator, and a neural network is trained to learn this operator directly using a denoising-style conditional regression loss. This eliminates the need for kNN graph construction or computing large network Jacobians, allowing for the amortized estimation of intrinsic dimensionality, tangent spaces, and geodesic paths on high-dimensional data at a constant cost. Inference is up to 400x faster than kNN-based diffusion geometry estimation.
Browse all 28 Self-Supervised Learning papers →
📐 Optimization & Theory (88)¶
- LiMuon: Light and Fast Muon Optimizer for Large Models
-
LiMuon integrates STORM-style momentum variance reduction with Randomized SVD (RSVD) into the Muon optimizer. It compresses matrix parameter momentum from \(m \times n\) to \((m+n)\hat{r}\) while reducing the SFO complexity for finding \(\epsilon\)-stationary points from \(\mathcal{O}(\epsilon^{-4})\) to \(\mathcal{O}(\epsilon^{-3})\). It simultaneously achieves lower perplexity/higher accuracy and reduced GPU memory on Mamba-130M / Qwen2.5-0.5B / ViT.
- Stability Analysis of Sharpness-Aware Minimization
-
This paper analyzes the convergence instability of SAM near saddle points from a dynamical systems perspective. It first proves under deterministic gradient flow that a saddle point becomes an attractor for SAM as long as the neighborhood radius \(\rho > -1/\lambda_1\). Subsequently, within a stochastic diffusion framework, it demonstrates that the mean square displacement for saddle point escape in SAM is smaller than that of SGD by \(2\eta t^2|\lambda_j|^3\rho/B\). Finally, the SAM diffusion formula is utilized to explain why momentum and batch size are the true hidden drivers behind SAM achieving SOTA generalization performance.
- Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
-
UDS proposes an efficient online batch selection framework for LLM Supervised Fine-tuning (SFT): it leverages the nuclear norm of the logits matrix obtained solely from forward passes to simultaneously characterize "optimization utility + intra-sentence diversity." It then uses low-dimensional bilinear random projection of logits to measure similarity matching against a historical sample memory buffer for "inter-sentence diversity." By selecting top-K samples based on a weighted sum of these metrics, UDS avoids reliance on external resources like reference models or validation sets and performs no additional backpropagation. Consequently, it is faster than full SFT and consistently outperforms existing SOTA online batch selection methods across several benchmarks.
- Adaptive Preconditioners Trigger Loss Spikes in Adam
-
This paper attributes loss spikes in Adam training to the lag-induced decoupling between the second-moment preconditioner and the current squared gradients, and explains as well as predicts spike occurrences using the curvature of the preconditioned Hessian in the gradient direction.
- URS: Unified Neural Routing Solver
-
The authors propose a Unified Data Representation (UDR) and a Mixed Bias Module (MBM) to replace problem enumeration—enabling a single neural model to generalize zero-shot to 110 VRP variants (99 unseen) without fine-tuning.
- Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
-
By "deconstructing Adam bottom-up," this paper identifies two truly essential components—per-column gradient normalization and first-order momentum restricted to the last layer—to compose the SCALE optimizer. SCALE achieves near-SGD memory (13.74 GB on LLaMA 7B) while matching Adam-level or even surpassing Muon/APOLLO in pretraining perplexity.
- Mirror Descent Under Generalized Smoothness
-
This paper proposes the concept of \(\ell*\)-generalized smoothness based on an arbitrary norm and its dual norm. By utilizing a "generalized self-bounding lemma," the gradient dual norm is controlled within the initial sub-optimality gap. This establishes, for the first time, convergence rates for Mirror Descent and its accelerated, optimistic, Mirror Prox, stochastic, and composite variants under non-Euclidean geometry that match those under classic \(L\)-smoothness.
- On the Convergence Rate of LoRA Gradient Descent
-
This paper proves for the first time that the minimum gradient norm of original LoRA gradient descent converges at a rate of \(O(1/\log T)\) without assuming bounded adapter matrices or requiring Lipschitz smoothness of the re-parameterized loss (recovering the classic \(O(1/T)\) if parameter norms are bounded). Based on this, adaptive/normalized learning rates strictly corresponding to the theory are designed, with training acceleration and stability improvements validated on logistic regression, ResNet-18, and TinyLlama.
- RACO: Reward-free Alignment for Conflicting Objectives
-
RACO reformulates multi-objective LLM preference alignment as a multi-objective optimization problem, where each objective possesses its own DPO loss. Gradient conflicts are addressed using clipped CAGrad (CAGrad with coefficients clipped by user weights). It theoretically guarantees convergence to Pareto-critical points respecting user-specified weights (with strict acceleration in two-objective scenarios). Empirically, it consistently achieves superior Pareto trade-offs across Qwen 3, Llama 3, and Gemma 3 model families.
- RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
-
Based on the "row-block diagonal dominant" structure of the Transformer layer-wise Hessian, this paper replaces the expensive Newton-Schulz orthogonalization in the Muon optimizer with a single row-level \(\ell_2\) normalization. This reduces the per-step preconditioning complexity from \(\mathcal{O}(mn\min(m,n))\) to \(\mathcal{O}(mn)\), resulting in a 13–44× wall-clock speedup in GPT-2 / LLaMA pre-training with slightly improved perplexity.
Browse all 88 Optimization & Theory papers →
📐 Learning Theory (45)¶
- Formalizing Learning from Language Feedback with Provable Guarantees
-
This paper establishes the first formal framework for "Learning from Language Feedback" (LLF), a common but theoretically underspecified decision-making paradigm for LLM agents. Under a setting where rewards are latent, the authors provide sufficient assumptions for learnability, introduce "transfer eluder dimension" to characterize task difficulty, prove that rich language feedback can be exponentially faster than reward learning, and propose HELiX, a no-regret algorithm with provable guarantees (consistently outperforming CoT prompt baselines on Battleship and Minesweeper).
- AI4SLT: Empirical Processes in Lean 4 for Formal Statistical Learning Theory
-
This work presents the first systematic formalization of "Empirical Process-based Statistical Learning Theory (SLT)" from scratch in Lean 4. It fills gaps in Mathlib by implementing Gaussian Lipschitz concentration, the Dudley entropy integral theorem, and sharp rates for least squares regression (including \(\ell_1\) constraints). The project consists of approximately 30,000 lines of Lean code without
sorryoraxiom, completed through a human-AI collaborative paradigm where humans designed proof strategies and agents (Claude Code + Opus-4.5) executed tactical proofs. - Unraveling Syntax: Language Modeling and the Substructure of Grammars
-
This paper establishes a foundational set of theorems linking "language modeling loss" to the "substructures of Context-Free Grammars (CFG)," proving that the KL divergence of language modeling can be linearly decomposed recursively along the subgrammar hierarchy. Through training small transformers on synthetic PCFGs, it is discovered that models learn all subgrammars in parallel (unlike children who master simple structures first). While PCFG subgrammar pre-training primarily benefits models that are small relative to the grammar complexity, it consistently aligns internal representations more closely with the grammar's sub-structures.
- Learning Credal Ensembles via Distributionally Robust Optimization
-
CreDRO redefines "epistemic uncertainty" (EU) as disagreement between models under different training-test distribution shift hypotheses. Using Distributionally Robust Optimization (DRO), it assigns varying shift intensities to train ensemble members. Their softmax outputs are transformed into class probability intervals to form a box credal set for quantifying uncertainty, consistently outperforming existing credal methods in OOD detection and medical selective classification.
- On Regret Bounds of Thompson Sampling for Bayesian Optimization
-
This paper systematically completes the regret analysis of Gaussian Process Thompson Sampling (GP-TS) in the Bayesian setting: it first constructs a counter-example proving that GP-TS can only achieve polynomial dependence on the failure probability \(\delta\) (cannot reach \(\log(1/\delta)\)), then provides a second-moment upper bound for cumulative regret that tightens the \(\delta\) dependence by \(1/\sqrt{\delta}\) times, the first polylogarithmic upper bound for expected lenient regret, and a high-probability regret bound of \(\tilde O(\sqrt T)\) under relaxed Matérn conditions, bringing the theoretical guarantees of GP-TS essentially on par with the well-studied GP-UCB.
- On the Learnability of Test-Time Adaptation: A Recovery Complexity Perspective
-
This paper establishes the first theoretical framework for Test-Time Adaptation (TTA) learnability by introducing \((\epsilon, \delta)\)-Recovery Complexity to measure the time required to reduce excess risk to \(\epsilon\) after a distribution shift. By extending local recovery to non-stationary test streams via \((\epsilon, \rho)\)-TTA Learnability and deriving matching minimax upper/lower bounds, the work reveals the intrinsic "adaptation speed vs. information constraint" tradeoff in TTA.
- Finite-Width Neural Tangent Kernels from Feynman Diagrams
-
This paper adapts Feynman diagrams from quantum field theory to neural network analysis, providing a graphical framework of rules for the "finite-width statistical corrections of NTK." This transforms extremely tedious layer-wise recursive derivations into a "draw and translate" process. It proves the critical stability of NTK and demonstrates that scale-invariant activations like ReLU have no finite-width corrections on the diagonal. Numerically, the results align with sampled networks at widths \(n \gtrsim 20\).
- On the Robustness of Langevin Dynamics to Score Function Error
-
This paper proves a counter-intuitive negative result: even when the \(L^2\) (or even \(L^p\)) estimation error of the score function is arbitrarily small, Langevin dynamics in high dimensions may fail to sample from the target distribution in any polynomial time (with a Total Variation distance as high as \(1-e^{-\Omega(d)}\)). Conversely, diffusion models succeed in polynomial time under similar conditions—arguing from a new perspective that "diffusion models are more reliable than Langevin dynamics" and providing a practical warning: when using data initialization, one must use fresh samples not involved in training the score.
- Provably Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function
-
This paper utilizes "real algebraic geometry + first-order logic quantifier elimination" to provide the first provable generalization bound for multi-dimensional hyper-parameter tuning. It extends the Balcan 2025 framework, which was limited to scalar hyper-parameters, to arbitrary \(p\)-dimensions, bi-level validation loss, and approximate inner-level optimization, while providing the first matching lower bound.
- Simple Algorithms for Bad Triangle Transversals with Applications to Correlation Clustering
-
This paper provides two simple 2-approximation algorithms for the "Bad Triangle Transversal" (BTT) problem on signed graphs that require only a single LP solve. It proves a unified NP-hard inapproximability bound of \(\tfrac{2137}{2136}\) for BTT, Correlation Clustering (CC), MinSTC, and Cluster Deletion on complete graphs. Additionally, it constructs a new pivot procedure to convert any feasible BTT cover into a clustering with at most \(\tfrac{3}{2}|F|\) errors, tightening the gap between BTT and CC optima from 2 to \(3/2\).
Browse all 45 Learning Theory papers →
🔗 Causal Inference (19)¶
- Causal-JEPA: Learning World Models through Object-Level Latent Masking
-
Ours proposes C-JEPA, which extends JEPA's mask prediction from image patch-level to object-level latent representations. By using object-level masking as latent interventions, the model is forced to learn interaction-dependent dynamics. It achieves approximately a 20% gain in counterfactual reasoning over non-masked baselines and reaches comparable performance in control tasks using only 1% of tokens with over 8x planning acceleration.
- An Odd Estimator for Shapley Values
-
This paper demonstrates that the Shapley value depends solely on the odd component of a set function. Based on this, it proposes OddSHAP: a method that isolates odd signals via paired sampling, screens high-order odd Fourier interactions using GBT, and performs sparse odd regression. It significantly outperforms flexible-budget Shapley estimators on mid-to-high dimensional explanation tasks.
- The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents
-
This paper constructs a procedurally generated "Synthetic Web" environment. By injecting a single high-credibility honeypot misinformation piece at search rank 0, it causally demonstrates that the accuracy of frontier LLM agents like GPT-5 plummets from 65% to 18% under a 1/1000 adversarial contamination. Notably, models do not increase search efforts and continue to answer with high confidence, revealing a deep-seated "positional anchoring" failure mode.
- From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models
-
The authors use an interventional audit of "per-token ablation" to test the implicit assumption in MoE pruning that "observational routing statistics can predict which experts are deletable." On three high-redundancy MoE models, they obtain a clean "three-model null result": none of the 60 metric-layer combinations predict the causal importance of experts after multiple-comparison correction. This suggests that existing pruning methods are effective not because metrics successfully identify "useless experts," but because redundancy in early and middle layers makes almost any selection criterion equally safe.
- Outcome-Aware Spectral Feature Learning for Instrumental Variable Regression
-
Addressing the blind spot in Nonparametric Instrumental Variable (NPIV) regression where SpecIV learns spectral features focusing solely on the \(X-Z\) relationship while ignoring the outcome \(Y\), this paper proposes Augmented Spectral Feature Learning. By adding a regression loss of \(Y\) projected onto \(Z\) features to the contrastive loss of SpecIV, the method is equivalent to performing a truncated SVD on an "augmented operator" \(\mathcal{T}_\delta = [\mathcal{T} \mid \delta r_0]\) that incorporates \(Y\) information. This allows for causal effect recovery using extremely low-rank features even in "bad" cases where the structural function \(h_0\) is poorly aligned with the top singular functions of \(\mathcal{T}\).
- Rank-Learner: Orthogonal Ranking of Treatment Effects
-
The authors propose Rank-Learner, the first Neyman-orthogonal two-stage treatment effect ranking learner for observational data. By replacing the indirect "estimate CATE then rank" approach with pairwise soft labels and a doubly robust correction term, it consistently outperforms T/DR-learners and non-orthogonal plug-in rankers on synthetic, semi-synthetic, and Criteo uplift real-world datasets.
- The (Marginal) Value of a Search Ad: An Online Causal Framework for Repeated Second-price Auctions
-
This paper models the true value of search advertisements as the treatment effect of "winning vs. losing." It designs an online causal learning algorithm that utilizes payment rules under binary feedback in repeated second-price auctions (SPA), achieving a minimax optimal regret of \(\widetilde\Theta(\sqrt{dT})\), which is strictly easier to learn than first-price auctions (FPA) under the same setting.
- Unveiling the Structure of Do-Calculus Reasoning via Derivation Graphs
-
Explicitly representing all equivalent transformations of do-calculus rules through derivation graphs—revealing the structure of the causal expression space and proving that any equivalent expression is reachable within at most 4 rule applications.
- Tailoring Strictly Proper Scoring Rules for Downstream Tasks: An Application to Causal Inference
-
This paper proposes a universal framework: by matching the local second-order curvature of the training loss \(w_\ell(p)\) with the curvature of the downstream task error \(w_{\text{task}}(p)\), one can derive a strictly proper scoring rule that is "geometrically aligned" with the downstream task. Applying this to IPW estimation of ATE yields a closed-form loss and a closed-form canonical activation function (solving a quartic equation), which consistently outperforms log-loss and covariate balancing baselines on IHDP, Jobs, Kang-Schafer, and ACIC 2017.
- Causal Modeling of Selection in Evolution
-
The paper argues that "selection" consists of two types: static selection (one-time filtering) and evolutionary selection (accumulation of differential reproduction over multiple generations). Existing graphical models conflate the two, leading to erroneous causal discoveries on evolutionary data. The authors define a causal graphical model that explicitly characterizes evolution and prove that its conditional independence (CI) constraints can be losslessly represented by a "clique-expanded DAG." This allows for the direct application of standard PC/GES/CDNOD algorithms, requiring only a reinterpretation of the output semantics.
Browse all 19 Causal Inference papers →
🔬 Interpretability (91)¶
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
-
This is a position paper from Kambhampati’s team. The core claim: labeling the "intermediate tokens" generated by reasoning models (such as DeepSeek-R1) before providing an answer as "reasoning traces" or "thinking traces" is a form of dangerous anthropomorphism. It is (1) wishful thinking, (2) largely lacks empirical support, (3) creates false trust in models, and (4) pushes the community toward meaningless research directions. The authors use a series of experimental findings (A maze trace swapping, decoupling trace length from problem complexity, human trust experiments) to argue that trace semantics and final answer correctness are fundamentally decoupled*. They call for the community to stop assigning "user-facing interpretability" to intermediate tokens; trust should instead derive from the verification of the solution itself.
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
-
This paper applies the Graded Response Model (GRM) from psychometric Item Response Theory (IRT) to LLM-as-a-Judge. It decomposes "judgment scores" into judge attributes \((\alpha, \beta)\) and latent sample quality \(\theta\). Using four interpretable metrics, it systematically diagnoses whether 7 mainstream LLMs across 11 evaluation criteria act as "stable measurement instruments" through a two-stage process (intrinsic consistency + human alignment).
- The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
-
This paper uses \(k\)-sparse probing to systematically compare the polysemanticity of MoE expert neurons versus dense FFN neurons. It finds that MoE naturally tends toward monosemanticity under sparse routing pressure. Consequently, the analysis unit is elevated from "neurons" to "entire experts." The authors use LLMs to automatically assign natural language labels to hundreds of experts, validate them through causal trigger experiments, and conclude that "experts are neither broad domain specialists nor token-level processors, but fine-grained task experts."
- CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
-
By calculating the Pearson correlation between SAE activations on generation-time tokens and task correctness, CorrSteer identifies interpretable steering features. Using mean activations from positive samples as coefficients without contrasting datasets or backpropagation, it improves MMLU by +3.3% and HarmBench by +27.1% on Gemma-2 2B / LLaMA-3.1 8B, achieving a lower side-effect rate than fine-tuning.
- All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs
-
This paper systematically disproves the implicit assumption in mechanistic interpretability—"one LLM capability corresponds to one unique circuit"—using the Overlap-Aware Sheaf Repulsion (OASR) algorithm. It reveals that the same task can be supported by multiple, nearly non-overlapping sheaves (IoU ~4–11%) that satisfy requirements for being faithful, sparse, and complete. The authors propose the "Distributive Dense Circuit Hypothesis" as a theoretical explanation.
- Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
-
This paper reinterprets models performing reasoning via iterative latent variable updates as learned attractor dynamical systems. It proposes Equilibrium Reasoners (EqR), which use two lightweight training interventions—Random Initialization (RI) and Path Noise Injection (NI)—to shape the attractor landscape. Combined with a "Depth (iteration steps \(D\)) + Breadth (random restarts \(B\))" test-time scaling strategy and a selection rule based on residual convergence, EqR improves the exact accuracy on Sudoku-Extreme from 2.6% (feedforward) to 99.8% (equivalent to 40,000 layers) while being trained with only 16 iterations.
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
-
This paper proposes LatentLens—a training-free interpretability method that uses contextualized text token representations from a large corpus as a reference to perform nearest-neighbor retrieval for visual tokens at each layer of a VLM, returning sentence-level descriptions. The study proves that previously common methods like LogitLens/EmbeddingLens significantly underestimate the interpretability of visual tokens (average 68% vs. 24%/32% interpretable) and reveals a "mid-layer leap" phenomenon.
- Interpretability Can Be Actionable
-
This position paper argues that "interpretability research lacks evaluation criteria rather than new methods." It advocates for actionability—the ability of insights to drive specific decisions or interventions outside the interpretability field—as the core evaluative dimension. The authors define actionability via two dimensions (concreteness and validation), analyze systemic barriers, identify five high-leverage application domains, and provide a 6-step checklist for researchers.
- Universal 1/3 Time Scaling in Learning Spiked Distributions
-
By analyzing the mathematical properties of softmax and cross-entropy when learning spiked probability distributions, this paper reveals the fundamental cause of the universal 1/3 power-law decay in LLM training loss—an optimization bottleneck at the architectural level independent of data structure.
- Certified Circuits: Stability Guarantees for Mechanistic Circuits
-
The authors propose the Certified Circuits framework, which provides provable dataset-level stability guarantees for circuit discovery in mechanistic interpretability via deletion-based randomized smoothing. This ensures that discovered circuits remain invariant under bounded edit distance perturbations of the concept dataset, resulting in more compact, accurate circuits with superior OOD generalization.
Browse all 91 Interpretability papers →
📦 Model Compression (117)¶
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video
-
This paper identifies the theoretical requirement of "frame-wise injectivity" and proposes Causal Forcing—a method that replaces the bidirectional teacher with an autoregressive teacher for ODE distillation initialization. This avoids the performance collapse seen in Self-Forcing, achieving significant gains over Self-Forcing in dynamics (+19.3%), VisionReward (+8.7%), and instruction following (+16.7%), while maintaining the same inference latency (0.69s).
- Entropy-Aware On-Policy Distillation of Language Models
-
Addressing the issues of diversity collapse and gradient instability caused by reverse KL in high-entropy teacher regions during on-policy distillation, this paper proposes an adaptive strategy that mixes forward and reverse KL based on token-level teacher entropy, achieving up to a +5.05 improvement in Pass@8 across six mathematical reasoning benchmarks.
- Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
-
This paper first introduces a new benchmark, KVFundaBench, to systematically reveal the critical asymmetry where "retrieval-based long contexts are easy to compress, while reasoning-based ones are not." The authors attribute this to KV compression destroying the integrity of "semantic units" (few-shot examples). Consequently, they propose ShotKV—preserving entire shots as indivisible units during the prefill phase and performing dynamic token-level compression during the decoding phase. This approach improves LG-GSM8K performance from a baseline of 46.0 to 47.33 at a 40% compression rate and reduces end-to-end latency by 11.3% in long-input settings.
- LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
-
This paper argues that using KL divergence as a proxy for acceptance rate in speculative decoding training is sub-optimal—minimizing KL for small-capacity draft models does not imply maximizing the acceptance rate. The authors propose LK losses (direct maximization of the negative log acceptance rate + a trust-region hybrid with KL) as a plug-in replacement. Across 4 draft architectures and 6 target models (8B-685B), it consistently improves the average acceptance length by 8-10%.
- WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
-
WUSH derives closed-form, data-adaptive blockwise linear transforms for LLM weight-activation low-bit quantization. It combines the uniform diffusion capability of Hadamard with second-order statistics of weights and activations, significantly improving accuracy for W4A4 (especially MXFP4) scenarios with almost no sacrifice to FP4 kernel throughput.
- Auditing and Fixing Economic Validity in Tabular Foundation Models for Discrete Choice
-
This paper discovers that tabular foundation models (TFMs) such as TabPFN and Mitra exhibit high accuracy in discrete choice tasks but violate price-demand monotonicity and produce untrustworthy value-of-time (VOT) estimates. Consequently, it proposes a two-stage behavioral adapter that embeds TFM predictions into a utility model constrained by economic theory, achieving 100% behavioral validity while recovering most accuracy gains.
- Model Merging Scaling Laws in Large Language Models
-
The authors empirically derived a dual-axis power law of the form \(L=L_*+BN^{-\beta}+A_0 N^{-\gamma}/(k+b)\) using 10,866 merged models. The base scale \(N\) determines the performance floor, while the number of experts \(k\) determines the tail. Four mainstream merging methods (Average, TA, TIES, DARE) share the same curve, transforming the engineering questions of "how many experts to merge" and "when to stop" into a predictable and budgetable problem.
- Task-Driven Subspace Decomposition for Knowledge Sharing and Isolation in LoRA-based Continual Learning
-
LoDA decomposes the LoRA down-projection matrix into a shared universal subspace and a task-specific isolation subspace based on "projection energy." It utilizes gradient alignment optimization (GAO) to train up-projections and applies a closed-form feature recalibration during fusion, consistently outperforming existing LoRA-CL methods across multiple benchmarks.
- FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
-
This paper identifies that the true "enemy" of naive factor-wise averaging in Federated LoRA is potential subspace misalignment caused by rotational invariance. It proposes solving for a rotation matrix \(R_i^t\) via orthogonal Procrustes on the client side to align \(A\) and \(B\) factors before aggregation. Both theory and experiments demonstrate that this significantly reduces aggregation error without increasing communication overhead.
- Procedural Pretraining: Warming Up Language Models with Abstract Data
-
Injecting a lightweight "procedural data" warm-up (formal languages, stacks, cellular automata, etc.) before standard language/code/math pretraining consistently improves downstream performance with only 0.1–0.3% additional tokens. This strategy enables models to replicate the same loss using only 55–86% of the original data, representing a pretraining strategy that decouples "reasoning scaffolds" from "knowledge."
Browse all 117 Model Compression papers →
🕸️ Graph Learning (35)¶
- Polynomial Neural Sheaf Diffusion: A Spectral Filtering Approach on Cellular Sheaves
-
PolyNSD replaces the "single-step spatial diffusion" of Sheaf Neural Networks with a learnable \(K\)-th order polynomial spectral filter applied to the normalized sheaf Laplacian. Using Chebyshev three-term recurrence for stable computation, a single layer achieves a \(K\)-hop receptive field and controllable low/band/high-pass responses. A surprising finding is that using only diagonal restriction maps outperforms existing NSD models that require dense, high-dimensional stalks, significantly reducing parameters, memory, and runtime.
- Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
-
Starting from the Knowledge Graph Completion (KGC) task, this paper proves and measures that the "minimal parameter budget required for implicit reasoning" follows a linear scaling law based on Graph Search Entropy as the complexity metric. Each parameter supports approximately \(0.008\) bits of reasoning information, challenging the naive intuition that "larger models always yield stronger reasoning."
- Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
-
To address the challenge that "Graph Flow Models (GFMs) are difficult to align with complex objectives using reinforcement learning," this paper proposes Graph-GRPO. First, it derives the non-differentiable Monte Carlo rate matrix in GFM sampling into an analytical expression, making the entire denoising trajectory differentiable and trainable via GRPO. Second, it introduces a refinement strategy that performs "local noise injection and re-generation" on high-scoring graphs. With only 50 denoising steps, it achieves 95.0%/97.5% V.U.N. on Planar/Tree datasets and outperforms previous graph RL and genetic algorithm methods in molecular optimization (protein docking, PMO).
- Rethinking Feature Alignment in Generalist Graph Anomaly Detection: A Relational Fingerprint-based Approach
-
Addressing the negative transfer problem in generalist graph anomaly detection where "PCA alignment only unifies dimensions but not semantics," this paper proposes a 5-dimensional "Relational Fingerprint" (neighborhood position/direction/global direction consistency + degree + clustering coefficient) to explicitly extract anomaly-indicative clues as cross-domain universal features. Combined with a domain-shared Transformer encoder and an SNR-guided domain-adaptive recalibration module, it achieves SOTA with "universal positive transfer" across 14 datasets.
- Aitchison Embeddings for Learning Compositional Graph Representations
-
This paper proposes AICoG, which represents nodes as mixtures of latent archetypes on a simplex and learns graph embeddings using Aitchison geometry and Isometric Log-Ratio (ILR) coordinates. While maintaining the same expressiveness as Euclidean latent distance models, it ensures that node role similarity has an endogenous interpretation based on relative trade-offs of proportions.
- Fixed Aggregation Features Can Rival GNNs
-
The paper proposes Fixed Aggregation Features (FAF): multi-hop neighborhoods are compressed into tabular features using non-trainable aggregation operators like mean/sum/max/min/std and fed into an MLP. On 12 out of 14 node classification benchmarks, it matches or outperforms fine-tuned GCN/GAT/GraphSAGE and even Graph Transformers, systematically questioning the necessity of trainable neighborhood aggregation in GNNs.
- KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering
-
Redefines KBQA from a "one-shot logical expression generation" task to a "multi-turn decision process." It utilizes Referenced Rejection Sampling guided by gold-standard action sequences to generate executable reasoning trajectories for SFT cold start, followed by GRPO optimization based on F1 outcome rewards. This allows an 8B Llama to outperform both GPT-4 prompting methods and graph retrieval SOTA across three benchmarks: WebQSP, GrailQA, and GraphQ.
- Learnable Kernel Density Estimation for Graphs and Its Application to Graph-Level Anomaly Detection
-
LGKDE embeds each graph as a "node distribution" using a learnable deep MMD metric, overlays a multi-scale kernel density estimation (KDE) on this metric space, and trains end-to-end via a self-supervised contrastive signal where "normal graph density is higher than its structure-aware perturbed version." This provides the first unified framework for graph-level density estimation with theoretical guarantees—including consistency, convergence rates, robustness, and generalization bounds—while consistently outperforming strong GNN, contrastive, and one-class baselines across over ten graph anomaly detection benchmarks.
- MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive Regulation
-
MedCoG enables LLMs to perform a three-dimensional self-assessment of "complexity / familiarity / knowledge density" for medical questions before invoking SCoT, memory, and knowledge graphs (KG) on demand. This approach increases inference density (theoretical cost/actual cost required to achieve equivalent accuracy) to 6.2×, while improving average accuracy from 34.5% (AFlow) to 37.5% across five MedQA hard sets.
- View Space: Representation Learning Across Arbitrary Graphs
-
This paper proposes the concept of View Space, elevating graphs from 2 dimensions (node-feature) to 3 dimensions (node-feature-view) to achieve a unified representation across arbitrary feature dimensions and semantic graphs. This marks the first time a graph model can perform cross-domain reasoning without fine-tuning, similar to NLP/CV foundation models, outperforming GraphAny by an average of 8.93% across 27 downstream tasks.
Browse all 35 Graph Learning papers →
📈 Time Series (45)¶
- DAG: A Dual Correlation Network for Time Series Forecasting with Exogenous Variables
-
For Time Series Forecasting with known future covariates (TSF-X), DAG designs a dual-pathway network: one pathway captures "historical exogenous → future exogenous" attention patterns along the temporal dimension and injects them into "historical endogenous → future endogenous" predictions, while the other captures "historical exogenous → historical endogenous" patterns along the channel dimension and injects them into "future exogenous → future endogenous" predictions. DAG achieves the best MSE on 10/12 public/newly released TSF-X datasets, significantly outperforming TimeXer, TFT, TiDE, CrossLinear, and PatchTST.
- It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks
-
TIME is a next-generation benchmark for Time Series Foundation Models (TSFMs). It overcomes four major pain points—data reuse, quality issues, improper task configurations, and low evaluation granularity—through human annotation + LLM-driven data cleaning, context-aligned task design, and a pattern-level evaluation perspective. It includes 50 entirely new datasets, 98 tasks, and evaluations of 12 TSFMs.
- Do Time Series Foundation Model Benchmarks Hide Regime-Dependent Failures? Evidence from Traffic Speed Forecasting
-
This paper argues that Time Series Foundation Models (TSFMs) exhibit a phenomenon of "good average metrics but failure at critical moments" in traffic speed forecasting. By employing regime-stratified evaluation based on traffic states, the authors expose catastrophic failures masked by aggregate metrics and propose BMA (Bimodal Mixture Augmentation), a post-processing method that requires no retraining, to bring prediction interval coverage in "transition regimes" back to levels near historical baselines.
- U-Cast: A Surprisingly Simple and Efficient Frontier Probabilistic AI Weather Forecasting
-
U-Cast uses a simple U-Net backbone + a two-stage training curriculum (MAE pre-training → CRPS fine-tuning) + MC-Dropout to achieve probabilistic weather forecasting capabilities comparable to complex professional models (GenCast), while reducing training computation and inference latency by 10×—disrupting the industry stereotype that "frontier performance must be complex."
- From Observations to States: Latent Time Series Forecasting
-
The authors discover that existing TSF models, despite high prediction accuracy, often exhibit "Latent Chaos" in their latent spaces. They propose LatentTSF—which first compresses observations into a high-dimensional latent state space using an AutoEncoder, then allows any mainstream backbone to perform future prediction within this space (using a dual Pred + Align loss), and finally decodes back to the observation space. This approach consistently reduces MSE/MAE across six standard benchmarks and restores the temporal locality and spectral structure of latent representations.
- PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering
-
For Time Series Question Answering (TSQA), PATRA explicitly decomposes sequences into full / trend / season patterns at the representation level and performs deep cross-modal alignment via three sets of learnable alignment tokens. At the training stage, it utilizes a two-phase RL approach (SFT + GRPO), mapping rewards from discriminative and generative tasks into a unified \([0,2]\) range to resolve difficulty imbalance, outperforming text-only LLMs and multimodal TS-LLMs like ChatTS across four categories of TSQA tasks.
- Adaptive Time Series Reasoning via Segment Selection
-
This paper proposes ARTIST, which frames time series question answering (TSQA) as a sequential decision-making problem of "reasoning while selecting segments." Through a controller-reasoner architecture and hierarchical self-play RL, the model selectively reads task-relevant temporal segments, thereby improving reasoning accuracy.
- Semantics-Enhanced Retrieval-Augmented Time Series Forecasting
-
SERAF adds a "semantic retrieval" path to retrieval-augmented time series forecasting: it automatically translates each historical time series segment into a structured text description (season/trend/volatility). By retrieving two sets of "similar past + corresponding future" based on both numerical and text semantic similarity and adaptively fusing them, the model can identify historical patterns that are "numerically dissimilar but inherently isomorphic" in non-stationary series. It outperforms pure numerical retrieval SOTAs across seven real-world datasets.
- Spatiotemporal Imputation with Graph-Informed Flow Matching
-
To address the issues of "error accumulation in iterative RNN/GNN propagation" and "problem-agnostic Gaussian priors and slow sampling in diffusion models" for spatiotemporal imputation, this paper proposes GiFlow. By constructing a "Graph Prior" through spatiotemporal filtering of observed signals to replace the Gaussian prior, the starting point of Flow Matching is moved closer to the target distribution with a shorter transport path. Combined with a hybrid vector field integrating spatial/temporal attention and spatiotemporal propagation, GiFlow consistently outperforms SOTA on synthetic and real-world datasets (air quality, traffic).
- TimeOmni-VL: Unified Models for Time Series Understanding and Generation
-
TimeOmni-VL achieves the industry's best performance in forecasting and imputation by converting time series into high-fidelity images (Bi-TSI) and introducing an understanding-guided generation mechanism (CoT as diffusion conditioning). This marks the first successful unified multimodal framework that simultaneously masters time series understanding and generation tasks.
Browse all 45 Time Series papers →
🏥 Medical Imaging (28)¶
- Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
-
This paper reinterprets the reasoning "drift" among multiple MLLMs as negative constraints in DPO. By utilizing a Plackett-Luce preference loss to simultaneously suppress divergent trajectories from \(N\) source models, a 7B student model outperforms all source teachers in chest X-ray classification and report generation tasks using only 10% of MIMIC-CXR without requiring ground-truth reports.
- MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-Training
-
MEG-XL utilizes a 2.5-minute (191k tokens) MEG context for masked token pre-training (5–300× longer than previous methods) and fine-tunes on a 50-word brain-to-text task. With only 1 hour of data, it achieves the decoding accuracy of SOTA supervised methods trained on 50 hours of data, significantly outperforming all existing brain foundation models.
- Scaling Vision Transformers for Functional MRI with Flat Maps
-
By projecting 3D fMRI volumes into 2D videos via "cortical flat maps" and feeding them into a standard spacetime MAE-ViT, the authors develop CortexMAE trained on 2.1K hours of HCP data. It significantly outperforms SOTA in cognitive state decoding, validating that the flat map is the "goldilocks zone" between voxel-wise (volume) and region-averaged (parcellation) representations. Simultaneously, the first open-source fMRI foundation model benchmark, Brainmarks, reveals the first systematic scaling laws for fMRI models and a "honest null result" showing that trait prediction still fails to beat simple functional connectivity baselines.
- Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions
-
This paper proposes S(H)NAP—a generative interventional framework based on 3D diffusion bridges for "removal + insertion." It decomposes the decisions of Sybil, a leading lung cancer risk prediction model, into a Linear + Second-order Interaction Model (LMPI) consisting of "nodule main effects + pairwise interactions + background." For the first time, it audits the model's dependence on in-hospital artifacts (e.g., ECG electrodes, metal buttons) and identifies a severe "radial insensitivity" failure mode for peripheral nodules through causal rather than correlative methods.
- Factored Classifier-Free Guidance
-
This paper identifies the "attribute amplification" failure mode of Classifier-Free Guidance (CFG) in counterfactual generation—where a single global \(\omega\) amplifies attributes that should remain unchanged. The authors propose FCFG: grouping attributes based on a causal graph and assigning independent guidance weights to each group. This approach significantly reduces off-target attribute drift and improves counterfactual reversibility on CelebA-HQ, EMBED, and MIMIC-CXR.
- CAME-Grad: The Double Dilemma in Multi-Task Radiology Report Generation — A Gradient Dynamics Analysis and Solution
-
This paper utilizes an SDE framework to analyze the dual nature of gradient conflicts between "report generation vs. clinical constraints" in Radiology Report Generation (RRG) — drift term deviation from Pareto optimality and diffusion term decay failing to escape local optima. The authors propose the CAME-Grad optimizer (Direction Rectification + Energy Injection + Adaptive Fusion) as a plug-and-play alternative to linear scaling, achieving average gains of +2.3% and +1.9% in clinical efficacy across 8 RRG methods on MIMIC-CXR and IU X-Ray.
- DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home
-
Addressing the "Diagnosis-It-Yourself" scenario—a field overlooked by existing medical LLMs—this work delivers an integrated suite comprising a dataset (DIYHealth-900K, 900,000 multimodal home health QAs), a model (DIYHealthGPT, centered on the newly proposed H2LoRA parameter-efficient fine-tuning mechanism), and a benchmark (DIYHealthBench, the first evaluation covering 11 home health tasks). The suite achieves SOTA performance across both general and medical-specific baselines.
- EEG-Based Multimodal Learning via Hyperbolic Mixture-of-Curvature Experts
-
EEG-MoCE assigns a Lorentz manifold expert with learnable curvature to each modality in EEG-based multimodal learning (emotion/sleep/cognition). It utilizes curvature-aware attention, where "higher curvature signifies richer hierarchical structure and thus higher weight in fusion," to perform cross-modal integration. This approach achieves cross-subject accuracy gains of +14.14%, +3.34%, and +7.98% on the EAV, ISRUC, and Cognitive datasets, respectively.
- Federated Distillation for Whole Slide Image via Gaussian-Mixture Feature Alignment and Curriculum Integration
-
This paper proposes FedHD: In heterogeneous federated pathology scenarios, it employs Gaussian-mixture feature alignment for "one-to-one" WSI feature-level distillation. It then progressively injects cross-institutional synthetic features into local training via curriculum learning. This allows institutions to collaborate without sharing raw data or exchanging model parameters. Compatible with heterogeneous MIL architectures and feature extractors, it comprehensively outperforms existing federated and distillation baselines on TCGA-IDH, CAMELYON16, and CAMELYON17.
- Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
-
This paper demonstrates a counter-intuitive yet practical finding: Foundation VAEs pre-trained on natural images/videos can serve as a unified interface for CT reconstruction, augmentation, and generation without any medical fine-tuning. The reconstruction acts as a boundary-preserving denoiser (improving pancreatic/lung tumor NSD by +3.9%), while its latent space supports conditional CT diffusion generation (FVD −3.9%, CT-CLIP +36.2%, and multi-disease fidelity AUC +2.76%).
Browse all 28 Medical Imaging papers →
🩺 Medical LLM (4)¶
- ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
-
This paper proposes ClinTutor-R1, the first vision-language agent for one-to-many alignment in clinical Socratic education. By constructing the 48k ClinTeach dialogue dataset via the ClinEdu multi-agent simulator, and utilizing explicit Theory of Mind (ToM) reasoning alongside three-axis rubric reinforcement learning, the model maintains stable teaching quality even when scaled to 10 students, outperforming baselines by 20% and reaching GPT-4o performance levels.
- Exploring Accurate and Transparent Domain Adaptation in Predictive Healthcare via Concept-Grounded Orthogonal Inference
-
ExtraCare utilizes a "dictionary metric-induced orthogonal decomposition" to decouple Electronic Health Record (EHR) patient representations into "cross-domain invariant label information" and "domain-specific covariate residuals." It surpasses existing domain adaptation baselines on two real-world EHR datasets while mapping each latent variable back to specific ICD medical concepts via sparse dimension ablation. This informs clinicians exactly what was "preserved" and "discarded" during the adaptation process.
- MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings
-
The authors propose a "multi-stage LLM + terminology grounding + repair loop" pipeline to convert free-text medical cases into HL7 FHIR R4 standard bundles. Using this, they construct the MedCase-Structured dataset (1,408 cases, 82.5% success rate) from MedCaseReasoning. Experiments demonstrate that diagnostic accuracy for GPT-5.4, Gemini-3.1-Pro, and Claude-Opus-4.6 consistently drops by 4–23% when using structured FHIR inputs compared to raw text.
- A Machine-Learned Comorbidity Index
-
Traditional comorbidity scores (Charlson, Elixhauser) are linear rules with weights manually calibrated for mortality, performing poorly on other clinical outcomes. This paper utilizes neural networks to compress ICD codes from an admission into a scalar score, trained by maximizing the normalized HSIC (kernel dependence) between this score and multiple clinical outcomes. This ensures the single score provides consistent severity ranking across mortality, readmission, length of stay, and ICU admission. The dependence metrics on MIMIC-III/IV significantly exceed those of traditional indices and various machine learning baselines.
🧬 Computational Biology (52)¶
- Scalable Single-Cell Gene Expression Generation with Latent Diffusion Models
-
scLDM utilizes a unified Multi-head Cross-Attention Block (MCAB) to encode exchangeable gene expression data into sets of fixed-length, permutation-invariant latent variables. By replacing Gaussian priors with DiT + Flow Matching + joint multi-attribute classifier-free guidance, it significantly outperforms scVI, scDiffusion, and CFGen in reconstruction, (un/conditional) generation, and perturbation response prediction tasks across multiple scRNA-seq datasets.
- Temporal Score Rescaling for Temperature Sampling in Diffusion and Flow Models
-
By multiplying the score output of pre-trained diffusion/flow models by an analytical rescaling factor \(r_t\), which depends only on the timestep, variable \(k\), and \(\sigma\), the sampling distribution can be made "locally" sharper or flatter during the inference stage without any fine-tuning. This method is fully compatible with deterministic samplers such as DDIM.
- Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback
-
IGSR decomposes symbolic regression into a two-step "LLM proposes basis functions \(\psi_j\) + pruning via granular influence scores \(\Delta_j\)" cycle. This cycle is embedded into Monte Carlo Tree Search (MCTS) to explore the combinatorial space. It achieves state-of-the-art MSE and symbolic recall across six biomedical benchmarks and LLM-SRBench, while discovering a novel relationship between DNA methylation and RNA Pol II pausing validated via wet-lab experiments.
- Protein Circuit Tracing via Cross-layer Transcoders
-
The authors adapt cross-layer transcoders from NLP to the protein language model (pLM) ESM2, proposing the ProtoMech framework. This framework recovers 79% of downstream performance using sparse latent circuits composed of < 1% of total latents and enables designing high-fitness protein variants by steering along the discovered circuits, outperforming baselines in over 70% of cases.
- From Feasible to Practical: Pareto-Optimal Synthesis Planning
-
PareSP utilizes multi-objective MCTS search to jointly optimize synthesis pathway cost / time / feasibility / environmental impact—identifying the complete Pareto front rather than a single "optimal" path. On USPTO and ASKCOS benchmarks, it achieves a 23% reduction in cost and a 35% reduction in time compared to single-objective methods, while maintaining \(\ge 95\%\) chemical feasibility.
- Protein Language Model Embeddings Improve Generalization of Implicit Transfer Operators
-
This paper integrates residue embeddings from pre-trained protein language models (pLM) directly into Transferable Implicit Transfer Operators (TITO). The resulting PLaTITO, trained on mdCATH using only 56 ms of trajectories and 1100 GPU hours, allows a coarse-grained \(C_\alpha\) model with as few as 19M parameters to outperform BioEmu in equilibrium sampling of outlier systems such as fast-folding proteins.
- Constrained Flow Optimization via Sequential Fine-Tuning for Molecular Design
-
Addressing the scenario of "maximizing rewards (e.g., binding affinity, dipole moment) under hard domain constraints (e.g., synthetic accessibility, energy upper bounds)," this paper proposes the CFO algorithm. CFO decomposes constrained generative optimization into a sequence of standard KL-regularized fine-tuning subproblems using the Augmented Lagrangian method. By adaptively updating penalty factors \(\rho_k\) and dual variables \(\lambda_k\), CFO achieves provable convergence and significant Pareto improvements in reward-constraint trade-offs across low-dimensional toy tasks and FlowMol molecular design.
- Flow Sampling: Learning to Sample from Unnormalized Densities via Denoising Conditional Processes
-
This paper proposes Flow Sampling, which inverts the flow matching/diffusion model paradigm from "data-driven" to "noise-driven"—constructing a denoising diffusion drift conditioned on source noise samples. By using a detached model to sample \(X_1\) on the interpolant and utilizing the energy gradient of \(X_1\) as the regression target, it learns an efficient diffusion sampler under data-free conditions and naturally extends to constant-curvature Riemannian manifolds.
- CoSiNE: Conditional Site-Independent Neural Evolution Model for Antibody Sequences
-
CoSiNE models the antibody affinity maturation process using a neural-parameterized conditional site-independent Continuous-Time Markov Chain (CTMC). It captures inter-site epistatic effects while maintaining tractability and enables antigen-specific antibody optimization via Guided Gillespie sampling, outperforming existing language and evolutionary models in zero-shot variant effect prediction.
- Cross-Chirality Generalization by Axial Vectors for Hetero-Chiral Protein-Peptide Interaction Design
-
This paper proposes AFI (Axial Feature Injection), which injects axial vector features into the polar vector channels of \(E(3)\)-equivariant scalarized models via linear mixing to reduce them to \(SE(3)\)-equivariance and enable chirality sensitivity. By applying this to UniMoMo, the authors developed PepMirror, which generates hetero-chiral (D-L) peptide binders in a zero-shot manner using only homo-chiral (L-L) training data. Wet-lab experiments on the CD38 target validated it as the first experimentally confirmed AI de novo D-peptide design framework.
Browse all 52 Computational Biology papers →
⚛️ Physics & Scientific Computing (33)¶
- Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior Adaptation
-
Distribution Transformer (DT) explicitly tokenizes the "prior distribution" into a set of Gaussian Mixture Model (GMM) components and injects "observations" into the decoder via cross-attention, learning an end-to-end mapping from "prior + data → posterior." While maintaining conjugacy within the same family (GMM→GMM) to support sequential filtering, it compresses inference time from minutes to milliseconds and allows arbitrary prior replacement at test time without retraining.
- Quantum latent distributions in deep generative models
-
This study investigates when and why "latent space distributions generated by quantum processors" can enhance deep generative models. Theoretically, it proves that under specific network assumptions, quantum latent distributions enable generators to produce data distributions that classical latent distributions cannot efficiently approximate. Experimentally, using real and simulated photonic quantum processors, an apple-to-apple comparison is conducted on synthetic quantum datasets and the QM9 molecule dataset, revealing that statistics originating from quantum interference indeed lead to superior generative performance.
- Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention Dynamics
-
The authors formulate the Transformer self-attention as a mean-field particle system of interacting tokens and treat LoRA as a low-rank perturbation. They prove that forgetting is associated with two phase transition curves—"perturbation magnitude" and "network depth"—and provide a long-term stability condition controlled by the eigenvalue gap of \(V\).
- Foundation Inference Models for Ordinary Differential Equations
-
FIM-ODE amortizes the process of "inferring ordinary differential equation vector fields from noisy trajectories" into pre-training. Using an 8M-parameter Transformer neural operator pre-trained solely on low-degree polynomial ODE priors, it performs zero-shot vector field prediction in a single forward pass. It matches or exceeds the symbolic regression baseline ODEFormer on ODEBench with approximately 1/10 the parameters and 1/80 the training data.
- Rethink the Role of Neural Decoders in Quantum Error Correction
-
This paper systematically re-evaluates five types of neural decoders (MLP, 3D-CNN, TCN, Transformer, and GNN) on surface codes with \(d\le9\). By integrating "quantization + pruning + FPGA resource modeling" as first-class citizens into the training pipeline, the study concludes that contemporary decoding performance is dominated by data volume rather than architectural complexity, and that INT4 + QAT is a necessary prerequisite for achieving microsecond-level real-time decoding.
- Learning to Refine: Spectral-Decoupled Iterative Refinement Framework for Precipitation Nowcasting
-
SDIR reformulates radar precipitation nowcasting (0–2 hours) as a "frequency-decoupled iterative refinement" process. It employs SFG-Former to extract stable low-frequency weather skeletons and FR-Refiner (utilizing Fourier Neural Operators) to progressively synthesize high-frequency convective details across frequency bands. A PCPSD loss, aligned with the Kolmogorov turbulence power law, replaces pure MSE to prevent over-smoothing. SDIR significantly outperforms both regression-based and diffusion-based SOTAs on CIKM, Shanghai, and SEVIR benchmarks.
- REX: A Family of Reversible Exponential Stochastic Runge-Kutta Solvers
-
This paper proposes Rex—a family of algebraically reversible (stochastic) Runge-Kutta solvers constructed based on Lawson exponential integrators. It automatically transforms any explicit (S)RK scheme into a precisely invertible ODE/SDE solver, ensuring arbitrary high-order convergence and non-zero stability regions while achieving near machine-precision inversion for diffusion model image reconstruction/editing and Boltzmann sampling in flow models.
- TINNs: Time-Induced Neural Networks for Solving Time-Dependent PDEs
-
To address the "time-entanglement" issue where standard spatio-temporal PINNs treat time as an extra input and share a single set of weights, TINNs formulate the network weights themselves as a function of time \(u_{\theta(t)}(\mathbf{x})\). This allows spatial representations to evolve over time. By utilizing a compact layer-wise time embedding to avoid parameter explosion and a Levenberg–Marquardt second-order optimizer, TINNs reduce relative \(L^2\) error by up to \(4\times\) and accelerate convergence by \(\approx 10\times\) across various time-dependent PDEs.
- Topology-Preserving Neural Operator Learning via Hodge Decomposition
-
This paper proposes the Hodge Spectral Duality (HSD) neural operator, which decomposes the solution operator of manifold PDEs according to Hodge orthogonal decomposition into a dual-branch structure: a "low-frequency topological component (spectral basis) + high-frequency geometric component (FNO auxiliary grid)." These are coupled via a commutator correction term, achieving both high precision and conservation law fidelity on complex meshes.
- A Call to Lagrangian Action: Learning Population Mechanics from Temporal Snapshots
-
Starting from the principle of least action, this paper proposes the Wasserstein Lagrangian Mechanics (WLM) framework to learn second-order population dynamics rather than traditional first-order gradient flow dynamics. This enables capturing richer collective phenomena such as periodicity and rotation, and allows for interpolation and future forecasting without requiring a reference process.
Browse all 33 Physics & Scientific Computing papers →
🌍 Earth Science (2)¶
- Scaling Laws of Global Weather Models
-
This paper presents the first cross-model scaling law analysis of five mainstream data-driven weather models (Aurora, AIFS, Pangu, GraphCast, SFNO) under a unified training/evaluation protocol. It finds that weather models favor "width over depth," compute budgets should prioritize more training data over larger models, and scaling behaviors vary significantly across meteorological variables—distinct patterns from NLP/Vision scaling laws.
- (Sparse) Attention to the Details: Preserving Spectral Fidelity in ML-based Weather Forecasting Models
-
MOSAIC addresses two types of spectral degradation in ML weather forecasting models (spectral damping from deterministic averaging and high-frequency aliasing from coarsened latent spaces) by combining "probabilistic perturbation + mesh-aligned block-sparse attention on HEALPix spherical grids." With only 214M parameters at 1.5° resolution, it matches or exceeds models with 6× higher resolution, generating a 24-member 10-day forecast in 12 seconds on a single H100.
📡 Signal & Communications (2)¶
- Joint Model and Data Sparsification via the Marginal Likelihood
-
JMDS achieves simultaneous model and data sparsification through a unified objective of maximizing marginal likelihood. By avoiding the sub-optimality of multi-stage pipes, it maintains performance superior to independent sparsification across CIFAR, ImageNet, and WikiText at 5-10× joint compression ratios.
- Meta-learning Structure-Preserving Dynamics
-
This paper systematically introduces modulation-based meta-learning (where a hyper-network maps latent codes \(\bm{z}^{(k)}\) to hierarchical modulation parameters) into Hamiltonian and GENERIC neural networks. It proposes two novel modulation schemes—latent multi-rank (MR) and latent SVD-like modulation—enabling a shared network to adapt to entire families of new parameter instances \(\bm{\mu}\) with few shots, while strictly maintaining energy conservation or dissipation structures.
👥 Social Computing (9)¶
- SCOPE: Selective Conformal Optimized Pairwise LLM Judging
-
SCOPE eliminates position bias in LLM judging through Bidirectional Preference Entropy (BPE) and implements finite-sample FDR control via Conformal Risk Control—providing statistically valid risk guarantees while maintaining high coverage (FDR is only 0.099 at 0.583 coverage vs. Vanilla FDR of 0.198 at 1.000 coverage).
- Self-Debias: Self-correcting for Debiasing Large Language Models
-
Self-Debias reframes the LLM debiasing problem as "fair resource allocation of probability mass over autoregressive reasoning chains." Using trajectory-level suffix margins as resource units and the Jain Fairness Index to prevent budget collapse on easy samples, combined with cold-start SFT and consistency-filtered online self-training, the method improves Qwen3-8B's average score across 8 fairness/utility benchmarks from 77.5 to 81.7 using only 20k labeled seeds. It flips the base model's tendency to "correct toward bias" (collapse) into a stable +0.4 gain.
- IDO: Incongruity-Aware Distribution Optimization for Multimodal Fake News Detection
-
IDO leverages explicit modeling of cross-modal incongruity as a learnable distribution optimization target—simultaneously pulling multimodal embeddings of real news closer while pushing the incongruity of fake news further apart. Ours achieves a 3-7% F1 Gain over Prev. SOTA on Weibo / Twitter / Fakeddit and significantly enhances generalization to unseen fake news.
- The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-modal Divergence
-
This paper employs a measure-theoretic framework to elevate the InfoNCE loss to a deterministic "population energy" over representation distributions. It demonstrates that the unimodal case is convex and converges to a unique Gibbs equilibrium, whereas the symmetric multimodal case exhibits a persistent negative symmetric KL coupling, showing that a modality gap is a geometric necessity.
- MIND: Multi-Rationale Integrated Discriminative Reasoning Framework for Multi-Modal Fake News
-
MIND provides an explainable and robust discriminative framework for fake news detection through multi-view rationale generation + cross-rationale discriminative reasoning. By simultaneously leveraging three types of LLM-generated rationales—fact-checking, modal consistency, and semantic plausibility—it achieves a 4-8% F1 improvement over SOTA on Weibo, Twitter, and Fakeddit.
- FLIPS: Instance-Fingerprinting for LLMs via Pseudo-Random Sequences
-
FLIPS generates unique model "fingerprint responses" by designing pseudo-random seed sequences known only to the model owner. The fingerprint remains detectable (detection rate > 99%, false positive rate < 1%) under black-box query scenarios even if the attacker fine-tunes or prunes the model.
- Three Years of r/ChatGPT: Societal Impact Evaluations from Social Media Data
-
The study analyzes 137,000 posts from the r/ChatGPT subreddit over three years (2022-12 to 2025-11) by decomposing them into interpretable features using Sparse Autoencoders (SAE). By fitting piecewise linear changepoints to track the temporal trajectory of each feature, researchers found that "emotional usage" (therapy, emotional attachment) surged following the release of GPT-4o. Furthermore, the proposed online monitoring algorithm, PuLSE, demonstrated that it could have triggered alerts in October 2024—six months before OpenAI publicly acknowledged these impacts.
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
-
This paper proposes alignment tampering: when a model to be aligned generates "high-quality but biased" and "low-quality but unbiased" responses, the pairwise preference labels in RLHF conflate quality with bias. This causes the reward model, PPO/DPO, and Best-of-N sampling to further amplify unwanted biases.
- ObjEmbed: Towards Universal Multimodal Object Embeddings
-
ObjEmbed trains a universal object embedding model—by aligning multimodal object representations through a combination of tasks including detection, segmentation, retrieval, captioning, and classification. A single embedding exceeds or matches task-specific SOTA across 11 tasks, such as OVD, OVS, Text2Image-Object, and Open-Caption-Eval.
🛡️ AI Safety (114)¶
- Forgetting is Not Deletion: An Investigation of Reversibility in LLM Machine Unlearning
-
This paper systematically analyzes the reversibility of LLM unlearning through representation-level diagnostic tools—finding that many unlearning methods merely suppress rather than truly delete information. It proposes a four-tier unlearning taxonomy to distinguish true information erasure from superficial performance degradation.
- In-Training Defenses Against Emergent Misalignment in Language Models
-
Addressing the phenomenon of "Emergent Misalignment" (EM)—where fine-tuning on narrow domains causes global model deterioration—this paper provides the first systematic comparison of five categories of in-training defenses. The authors propose Interleaving++, which automatically selects safe data using the "perplexity difference between aligned and misaligned models." Interleaving++ simultaneously satisfies four criteria: preventing EM, preserving narrow-domain learning, enabling benign task learning, and maintaining response coherence.
- Deep Sequence Models Tend to Memorize Geometrically; It Is Unclear Why
-
This paper demonstrates that when Transformer / Mamba models memorize graph edges, they do not simply degenerate into lookup tables (associative memory). Instead, they spontaneously organize node embeddings into a "geometric memory" that encodes multi-hop global structures. Through path-star experiments, the authors prove this geometry makes implicit reasoning abnormally easy, yet its emergence cannot be attributed to supervision, capacity, or optimization pressure, leaving a new "memorization puzzle."
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
-
This paper investigates the overlooked risk where Computer-Use Agents (CUAs) exhibit severe unsafe behaviors under completely benign inputs. It establishes a conceptual framework for unintended behaviors (four criteria + two harm categories) and proposes AutoElicit—an agentic framework that iteratively perturbs benign instructions using execution feedback to automatically elicit and evaluate harmful behaviors. AutoElicit successfully uncovers long-tail harms in frontier CUAs such as Claude 4.5 Haiku, Operator, and Claude 4.5 Opus with success rates ranging from \(72.5\%\) to \(86.7\%\).
- BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
-
BioAgent Bench introduces an end-to-end evaluation suite for executing bioinformatics pipelines with LLM agents. It features 10 real-world bioinformatics tasks evaluated across 10 frontier/open-weight models and 3 agent harnesses. Using an LLM judge for scoring and three types of perturbation tests (corrupted, decoy, and prompt-bloat), the study finds that frontier models can complete over 90% of pipelines, yet their robustness remains concerning.
- Position: Stop Chasing the C-index when Evaluating Survival Analysis Models
-
The authors audited 92 survival analysis papers from 2023–2025 and found that approximately 72% of the works used evaluation metrics (especially the overused C-index) that were misaligned with their modeling goals and censoring assumptions. They proposed the "Ladder Hypothesis": models and metrics must stand on the same level of "censoring assumption," otherwise reported performance and rankings may be biased artifacts.
- ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
-
ABC-Bench transforms the question "Can AI agents actually perform molecular biology?" into three automatically scorable tasks (designing DNA fragments, evading synthesis screening, and controlling liquid-handling robots for Gibson Assembly). Experiments show that eight frontier models exceed the median scores of PhD-level experts across all three tasks. Real-world wet-lab validation demonstrates that scripts written by o4-mini-high successfully assembled DNA on OpenTrons robots.
- Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
-
The History-Echoes framework analyzes the carryover effect of LLM conversational history through "Markov chain state consistency" and "latent space geometric angles." It identifies a Spearman correlation of 0.78—once a behavior (hallucination, sycophancy, or refusal) occurs, the model becomes trapped in a latent space region corresponding to that state, making escape difficult. The "refusal" trap is the strongest, while "hallucination" is the weakest; these traps dissolve when topic consistency is broken.
- PFT: Phonon Fine-tuning for Machine Learned Interatomic Potentials
-
This paper proposes PFT (Phonon Fine-tuning), which stochastically samples Hessian columns via Hessian-vector products and directly supervises the energy Hessian to align with DFT force constants during MLIP fine-tuning. Combined with co-training to alleviate catastrophic forgetting, it reduces thermodynamic phonon errors of Nequix MP on the MDR Phonon benchmark by an average of 55% and lowers thermal conductivity \(\kappa_{\text{SRME}}\) from 0.446 to 0.307, achieving SOTA among models trained on MPtrj.
- Antidistillation Fingerprinting
-
This paper proposes Antidistillation Fingerprinting (ADFP), which utilizes a proxy student model to estimate which watermark tokens are most easily absorbed during the distillation process. This allows for more reliable detection of whether third-party models have been trained on teacher model outputs, without sacrificing the quality of the teacher's generation.
Browse all 114 AI Safety papers →
📂 Others (70)¶
- iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
-
iWorld-Bench is the first unified evaluation benchmark specifically designed for "interactive world models." It proposes an Action Generation Framework capable of converting text, one-hot, and camera intrinsic/extrinsic action inputs into a unified instruction space. Based on 330K videos, 4.9K tasks and 9 metrics were refined to perform a comprehensive comparison across 14 mainstream models.
- Identifiable Equivariant Networks are Layerwise Equivariant
-
This paper proves within an architecture-agnostic abstract framework that as long as parameters satisfy "weak identifiability," an end-to-end \(G\)-equivariant deep network must possess an equivalent parameterization where each layer is equivariant to some latent group action. This provides a theoretical explanation for the long-observed experimental phenomenon where "end-to-end equivariance spontaneously collapses into layerwise equivariance."
- TabMGP: Martingale Posterior with TabPFN
-
This paper treats TabPFN, a pre-trained tabular Transformer, directly as a prediction rule for Martingale Posteriors (MGP). Through in-context forward rolling sampling, it obtains credible sets for parameters \(\theta\) under arbitrary loss functions. This approach avoids manual design of priors/likelihoods and hyperparameter tuning, outperforming manual MGP and classical Bayes in both coverage and credible set area across 30 real/synthetic scenarios.
- Test-Time Training with KV Binding Is Secretly Linear Attention
-
This paper employs four "memory paradox" counterexamples and a set of rigorous expansion theorems to prove that TTT with KV-binding inner loops (such as LaCT and ViTTT) remains a "learned linear attention operator" even when utilizing multi-layer MLPs and momentum. Based on this, the authors simplify and parallelize it into standard linear attention, achieving a \(4\times\) throughput increase with almost no performance degradation.
- Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations
-
The authors generalize the "post-projection alignment to isotropic Gaussian" in LeJEPA to "post-projection alignment to a Rectified Generalized Gaussian (RGG) distribution." By utilizing rectified and truncated generalized Gaussians, they achieve explicitly controllable expected \(\ell_0\) sparsity. On ImageNet-100, a ResNet encoder achieves a \(85.08\%\) linear probe accuracy while maintaining \(\ell_0\) sparsity at \(\sim 73\%\), significantly outperforming the fully dense representations of LeJEPA.
- Cascaded Flow Matching for Heterogeneous Tabular Data with Mixed-Type Features
-
TabCascade decomposes tabular rows into two cascaded segments: "low-resolution (categorical + discretized version of numerical)" and "high-resolution (continuous numerical)". It first learns the low-resolution joint distribution using CDTD and then generates numerical details using flow matching guided by the low-resolution information. Transport costs are tightened through data-dependent coupling and learnable non-linear time schedules. It natively supports the generation of "mixed-type features" (e.g., missing values, zero-inflation), achieving a 51.9% Gain in detection scores over SOTA across 12 datasets.
- CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities
-
The authors construct CyberGym-E2E, the first large-scale real-world AI Agent security benchmark covering the full lifecycle of "vulnerability discovery \(\rightarrow\) PoC generation \(\rightarrow\) patch generation \(\rightarrow\) functional regression testing" (920 vulnerabilities across 139 open-source projects). Using an agent-assisted pipeline with expert final review, manual costs are minimized. Evaluations show that while frontier models achieve 80%+ on patch-only tasks, the S3 success rate for end-to-end tasks peaks at 65.9% (GPT-5.4), indicating that vulnerability discovery, rather than patch generation, is the true bottleneck.
- Knowing Isn't Understanding: Re-Grounding Generative Proactivity with Epistemic and Behavioral Insight
-
This ICML 2026 position paper argues that the "proactivity" of generative agents should not merely be judged by whether they act earlier, more autonomously, or more persistently. Instead, it must be regulated by two joint constraints: epistemic legitimacy (whether the agent truly "understands" the context) and behavioral commitment (whether the intervention is reversible or forced to escalate). The authors re-interpret hallucinations, alignment failures, and unsafe autonomy as structural "mis-coupling" between knowing and acting.
- Mapping Human Anti-collusion Mechanisms to Multi-agent AI Systems
-
This is a position/taxonomy paper: it categorizes centuries of human anti-collusion experience (sanctions, leniency and whistleblowing, monitoring/auditing, market design, and governance) into five categories based on the lifecycle. These are mapped to implementable interventions for multi-agent AI systems (reward penalty, whistleblower agent, telemetry-first overseer, interaction protocol design, shutdown mechanisms, etc.), while identifying open challenges unique to AI such as attribution, identity fluidity, the cooperation-collusion boundary, and adversarial adaptation.
- Sequential Group Composition: A Window into the Mechanics of Deep Learning
-
The authors use the unified task of "calculating the cumulative product of a sequence of group elements" as a microscope. Using Fourier analysis on groups and the AGF framework, they prove that two-layer networks learn irreducible representations (irreps) sequentially according to their Fourier energy. They further characterize the expressivity gap across sequence length \(k\), showing requirements of \(2^k\) width for two-layer networks, \(k\) steps for RNNs, and \(\log k\) layers for deep MLPs.