Skip to content

🧠 NeurIPS2026 Accepted Papers

233 NeurIPS2026 paper notes covering Optimization & Theory (16), Computational Biology (11), Time Series (11), AI Safety (10), Interpretability (10), Reinforcement Learning (10), Learning Theory (9), Multimodal VLM (9) and other 50 areas. Each note has TL;DR, motivation, method, experiments, highlights, and limitations — 5-minute reads of core ideas.


💡 LLM Reasoning (8)

Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization

PEPO weights token updates in successful reasoning trajectories by uncertainty relative to their neighborhoods while preserving each trajectory's total weight, outperforming the compared GRPO and global-entropy baselines in the main mathematical reasoning experiments, with explicit limits on trend robustness and cross-task generalization.

Diversity Combining for Multi-Path LLM Reasoning

The paper explains diminishing returns in multi-path reasoning through correctness correlation and the design effect, then estimates a fixed deployment budget from a labeled offline four-path pilot; across five model–task configurations, 4–10 paths retain 96%–103% of the binary majority-vote accuracy at 32 paths, although that evaluation metric is not the actual plurality-vote accuracy for open answers.

Externalized CPDAG Summaries Improve LLM Causal Deduction

Structured Thinking asks the same large language model to produce a format-constrained CPDAG summary before checking whether a causal hypothesis holds in every compatible DAG, raising Qwen3.5-27B primary-seed F1(YES) on Corr2Cause from 73.01 to 86.36 without formally guaranteeing graph validity or reasoning over the entire equivalence class.

Provable Test-Time Scaling for Beam Search in LLM Reasoning

With sampling access to next tokens but no full logits, this paper filters low-frequency tokens before beam pruning and, under alignment between prefix likelihood and correctness, reduces the worst-step coverage dependence of search samples from worst-case quadratic to nearly linear; accuracy on an LLM matrix-multiplication task increases from 37.2% to 38.8%.

Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning

By evaluating answer correctness separately from the continued production of complete explicit reasoning, this paper shows that ordinary fine-tuning without reasoning traces can improve Chemistry accuracy while reducing valid reasoning to approximately zero, and that reasoning-region loss masking mitigates this structural degradation, with model- and task-dependent effects.

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

SAGE uses two soft potentials—whether an operation addresses unresolved constraints and whether its resulting state approaches a target structure—to guide both candidate sampling and rewards during post-training, improving long-horizon reasoning while retaining ordinary inference-time decoding; Qwen3.5-35B mathematical average accuracy rises from 61.92% with GRPO to 64.86%.

Structured Sparse Memory for Recurrent Reasoning

CoSE replaces a large task table with a factorized FiLM conditioning branch and rank-32 instance residuals, reducing task-memory parameters to roughly 1/15 while improving pass@2 in controlled ARC experiments; CHARM combines this memory with synthetic data, recurrent computation, and extended training to reach 84.0% / 46.7% on ARC-AGI-1 / 2 public evaluation, which cannot be attributed entirely to CoSE.

Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning

This diagnostic study charges diffusion generation and reward scoring to a shared inference budget, finds that deterministic top-1 PRM guidance on Dream-7B trails independent sampling with a task-matched ORM, and locates the failures through candidate-pool, terminal-scoring, and readout controls.


🦾 LLM Agent (4)

AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

AgentHop places 1,011 scientific four-option questions in a seven-tool sandbox with three resource caps and diagnoses failures through paper recall, conditional conversion, tool-call patterns, and resource failures; Gemini-3 Pro achieves 89.1% accuracy in the paper's evaluation, while similar overall scores can conceal different retrieval and synthesis bottlenecks.

ContractBench: Can LLM Agents Preserve Observation Contracts?

ContractBench uses a virtual clock, byte checks, and HTTP-trace validation to evaluate whether agents preserve cross-step contracts attached to tool artifacts; in the paper's experimental snapshot of 38 model variants, the highest success rate is 77.8%, and greater scale or newer versions do not guarantee improved reliability.

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

Without updating model weights, MoMHa uses a proposer to rewrite Python harnesses from execution traces, jointly optimizing accuracy, behavioral safety, and token cost; the authors report leading joint means on synthetic and real benchmarks, but source conflicts concerning metric scaling, the two-phase definition, and cost accounting require clarification.

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

ToolSearcher trains a multi-turn tool selector using a category-constrained curriculum, event-level advantages for first discoveries of target tools, and progress-dependent credit assignment, improving Qwen2.5-7B-Instruct's overall StableToolBench F1 from GDPO's 0.496 to 0.513 and improving AppWorld task outcomes with a separate execution agent.


👥 Multi-Agent (4)

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

AgentGrad identifies repairable prompts through single-agent interventions, extracts textual gradients from differences between original and intervention-adjusted intermediate outputs, and clusters and abstracts shared corrective patterns; with GPT-5-mini, it improves the five-task average over unoptimized prompts by 11.76 percentage points, although some performance and cost summaries require the exceptions in the tables.

DAGent: Evaluate-then-Grow Planning for Deep Research Agents

DAGent evaluates completed sub-tasks before incrementally growing a research DAG, combining hierarchical evidence propagation with DAGRPO structural rewards: training-free Qwen3-32B achieves 47.3 / 55.3 / 65.0 Pass@1 across three benchmarks, outperforming its same-architecture Plan-then-Patch variant by 5.3 / 4.8 / 4.0 percentage points.

Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

The paper decomposes debate among three copies of one model into preservation, collapse, correction, and unrepaired transitions, using pre-debate screening and round-level traces to locate risk; however, an offline freeze replay prevents 29 collapses while losing 108 corrections, showing that a risk signal is not an effective controller.

XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication

XBridge transfers the full context token sequence through deterministic lexical anchor mapping and reads sender hidden states through a pairwise-trained cross-attention bridge, outperforming 128-token text-summary communication on all seven tasks for three heterogeneous model pairs and reducing per-sample latency from 1.70 to 0.15 seconds in the specified H200 test.


⚖️ Alignment & RLHF (2)

Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

Holding the wrong answer fixed and changing its endorser, the paper uses direction removal and an independently fitted attribution patch to show selectively intervenable differences between wrong-source deference and user agreement in three open-weight families, so one sycophancy evaluation cannot substitute for the other.

Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment

GAPO constructs an ephemeral geometric anchor near the current policy along the direction that decreases the batch-average preference margin, then reduces brittle pairs' update weights according to their margin degradation, improving multiple alignment evaluations and controlled noise experiments without winning every model–metric combination.


🔒 LLM Safety (5)

Contrastive Representation Shaping for LLM Unlearning

CLReg uses augmented views of the same forget example as positives and retain examples as negatives to shape hidden representations alongside a base unlearning objective, improving aggregate scores in multiple settings without equating representation separation with knowledge removal or providing privacy guarantees.

LLM Alignment–Utility Asymmetry under Semantic-Preserving Transformations

Through paired evaluation of original and semantic-preserving representations, this paper finds that task capability transfer does not guarantee corresponding transfer of safety behavior, and qualifies this empirical finding through benign-data conditions, protocol controls, and defensive supervision analysis.

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

Using 213 selected cases in SkillCascade-Bench, the paper shows that passing individual skill reviews does not imply safe joint execution; sandbox stress tests across three agent systems and eight models yield an author-reported average success rate of 89.4%, while a behavior-composition defense achieves only partial improvement.

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

The paper measures inconsistent safety judgments between reasoning traces and final answers with DSAR and introduces SARA, an on-policy reinforcement learning method combining early safety awareness, full-trace safety, and answer safety; it improves consistency under perturbed reasoning on two DeepSeek-based models, but does not outperform answer-only rewards in every setting or safety metric.

UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?

UnlearningSoup replaces some repeated tuning in training-based large language model unlearning with evaluation-guided weight interpolation: EfficientSoup searches using the original model and two trained unlearned models, while PerformanceSoup merges existing candidates, improving empirical forgetting–retention quality and selection costs across benchmarks without eliminating base unlearning training or providing certified forgetting.


👻 Hallucination Detection (1)

The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining

The paper localises a sparse Commit-Abstain Circuit (CAC) using the pre-generation commit–abstain logit margin, identifies early commitment accumulation with insufficient late abstention correction, and trains a lightweight policy on component contributions that raises reported mean decision accuracy from 0.692 to 0.814, rather than demonstrating improved answer-content accuracy.


📊 LLM Evaluation (5)

LLM Judge Validation Under Sparse Overlap: From Inference to Design

This paper treats the quantity and allocation of overlapping annotations for LLM-judge validation as a statistical design problem: variance analysis guides overlap planning, while stratified allocation improves representativeness; low overlap substantially changes deployment decisions and judge rankings on real benchmarks, but stratification can increase false approval even as it reduces false rejection.

OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

OTROPE uses optimal transport in semantic embedding space to reweight human-labeled residuals from an older model and correct a newer model's mean proxy evaluation, reducing cross-model preference-win-rate estimation error without response likelihoods, but its effectiveness depends on representation coverage and residual transferability, and it does not outperform PPI under every shift.

SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

SciR generates verifiable deduction, induction, and causal-discovery problems before rewriting their premises as multi-document scientific discourse, separately varying inference complexity and premise obfuscation to diagnose LLMs; document formalisation remains a substantial bottleneck even with symbolic solvers.

Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning

OracleLadder fixes teacher-written mathematical sub-goals and crosses isolated milestone tests with increasing roadmap and answer assistance on the parent problem: across six models, 33–48% of a screened 354-problem NuminaMath set falls into a composition gap, where all milestones are solvable but the parent remains unrecovered; the share is 24–37% after excluding grading issues flagged by automated review, but this is not a pure causal measurement of an internal composition ability.

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

SpanUQ distills claim consistency across offline sampled responses into a probe over frozen LLM hidden states, using set prediction to locate semantic spans and assign continuous uncertainty; five separately trained backbone probes achieve 0.908–0.944 AUROC and support finer-grained risk filtering than rejecting entire responses.


⚡ LLM Efficiency (6)

Adaptive Mass-Segmented KV Compression for Long-Form Reasoning

AMS leaves token-importance scorers unchanged and instead uses attention-derived quality mass to form adaptive segments and allocate retention quotas before local selection, mitigating contiguous context loss in long-form reasoning; on Math500 with a 256-token cache budget, AMS-Expected improves over AdaKV-ExpE2 by 16.0 percentage points.

Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

Without training the model or changing the base sampler's stopping rule, ADAS dynamically discounts confidence using a candidate's attention to positions already selected in the current step and their uncertainty, improving parallel decoding at low numbers of denoiser evaluations; task-and-method-average gains are 9.11 and 10.46 percentage points for the two models.

Block Sparse Flash Attention

BSFA first computes all causally visible QK scores exactly inside FlashAttention-2, then uses offline-calibrated block-maximum thresholds to skip V loads, PV, and softmax-state updates for low-scoring blocks, achieving a LongBench score of 39.78% versus the dense baseline's 40.24% on Llama-3.1-8B and a 1.13× end-to-end prefill speedup on the longest 10 samples.

Fractional State Space Transition for Long Sequence Modeling

Frac approximates fractional long memory with finite exponential modes on a shared geometric timescale bank, then adds token-wise controls and independent read/write routing, preserving bounded recurrent state while raising the 1.3B model's average LongBench score from GDN's 16.0 to 17.9, without leading on every task or throughput measure.

It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

Synth expands sourced encyclopedic seeds into heterogeneous instruction and reasoning data, enabling the final Baguettotron student to acquire instruction-following and factual capabilities from random initialization in one training process; the 600M variant reaches 79.3% precision on verifiable facts about seed entities, but this does not establish broad knowledge coverage or assume that teachers and data generation are free.

SEED: Self-Speculative Decoding via Implicit Encoder–Decoder

SEED trains the last two layers of an existing decoder-only model as a drafter that reuses deep KV caches from the previous full-model verification, achieving average decoding speedups of 2.6×/2.7× over AR on Qwen3-1.7B/4B while maintaining or improving task quality under task-specific fine-tuning and adaptive tree verification.


📚 Pretraining (1)

Spectral-Sphere-Constrained Hyper-Connections

s²HC replaces nonnegative doubly stochastic constraints on multi-stream residual matrices with a mean-preserving spectral norm sphere, using dynamic rotations and bounded scaling in the zero-sum subspace to enable non-degenerate mixing and achieving average accuracy of 47.1, 50.2, and 50.7 across eight benchmarks on three language models pretrained from scratch.


✏️ Knowledge Editing (1)

Open-Vocabulary Domain Unlearning

The paper extends visual-domain unlearning from failure on training classes to failure on classes absent from unlearning training, combining Fisher-masked visual-parameter updates with Targeted Manifold Scattering (TMS) to improve the forgetting–retention trade-off in few-shot CLIP classification, without establishing certified data deletion or guaranteed semantic preservation.


🗣️ Dialogue Systems (1)

Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State

The paper lets users rewrite the state of an entire conversation through natural language rather than append more corrective messages; a controlled study with 17 participants supports usability, and three representative cases show shorter retained transcripts, without establishing improvements in long-term task performance or total API cost.


🌐 Multilingual & Translation (1)

One Threshold Does Not Fit All Languages: Language-Conditional Deferral for Reliable and Efficient Low-Resource Text Classification

The paper assigns per-language conformal thresholds to a single inexpensive multilingual classifier, addressing language-specific coverage gaps hidden by pooled calibration while exposing human-review costs; the guarantee concerns wrong automatic outputs over all traffic, not accuracy conditional on automatic output.


🔍 Information Retrieval & RAG (1)

LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting

LogicTree-RAG organizes a research paper into a retrievable, recursively expandable logic tree with evidence associations, then uses section-specific hybrid traversal to generate a long patent-description draft, achieving 20.22k output tokens, 43.79 coverage, and 66.34 source-combined factuality on the Pap2Pat test set; these results do not certify patentability or legal validity.


💻 Code Intelligence (5)

CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

CodeScaler distills high-quality execution feedback into a code reward model, uses strict code extraction and reward shaping to support reinforcement learning and candidate reranking without online execution, and exceeds RLVR by 1.55/4.23 percentage points in average Avg@8 across four benchmarks for 8B/14B policies trained on DeepCoder.

Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents

CRR restores a small number of decision states during software-agent training and scores realized actions using the terminal-return difference between the original trajectory and an alternative-action continuation, raising SWE-bench Verified pass@1 from extended GRPO's 36.7% to 41.7% within the same 40.125-hour training window on an eight-H100 node, without eliminating fork execution costs.

CTE-Bench: Counterfactual Trace Evaluation for Stateful Software Simulators

CTE-Bench executes software with and without an intervention to evaluate sustained response prediction on fixed future calls: effect-step accuracy reaches 54.3%–61.5% with correct earlier answers, falls to 24.8%–33.2% in free rollout, and whole-trace exact match peaks at just 1.2%.

Execution Guided Line-by-Line Code Generation

At inference time, EG-CFG previews and executes several short continuations of a code prefix, injects runtime traces into the prompt, and generates code token by token with dual-distribution classifier-free guidance while refreshing feedback at line boundaries; with DeepSeek-V3-0324, the paper reports 96.6% / 99.4% / 69.9% accuracy on MBPP / HumanEval / DS-1000, respectively, but requires executable provided tests and additional search computation.

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation

MemDocAgent combines a traversal order constrained by dependencies and directory hierarchy with a readable, writable, and verifiable RepoMemory to let one agent continuously produce component-, module-, and repository-level documentation; across 20 Python repositories, its Qwen3-Coder configuration achieves 0.979 completeness, 0.916 truthfulness, and 0.690 helpfulness, but its coverage-preserving fallback does not guarantee that every document is trustworthy.


🎨 Image Generation (6)

All Roads Lead to Rome: Flow-driven Multi-Anchor Exploration for Open-Environment Active 3D Mapping

The method generates multiple long-horizon exploration anchors with conditional flow matching, then combines obstacle-aware planning, exploration-mode clustering, and hierarchical path selection to improve mapping coverage in same-difficulty AiMDoom evaluation, difficulty transfer, and AiMDoom-to-MP3D transfer, without guaranteeing generalization to arbitrary open environments.

One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation

OVIE uses a frozen depth model to turn 30M single images into partially valid pseudo-view training pairs, learns observed regions through masked reconstruction and perceptual losses, and completes unknown regions through a full-image adversarial prior, ultimately generating a novel view from only a source image and target pose in one forward pass at 116 FPS on an H100, with multi-view fine-tuning improving in-domain performance.

Panoptic Scene Program Diffusion Transformer

PSP-DiT turns a panoptic scene program containing instances, attributes, relations, and counts into a latent variable jointly denoised with the image, uses ownership and cycle-consistency supervision to constrain its visual realization, and improves GenEval 2 from 32.8 to 38.2 over a matched internal baseline while reporting approximately 1.12 times the inference latency.

PixelDiT2: Representation-Grounded Pixel Diffusion Transformers

PixelDiT2 re-encodes the current noisy image at each pixel-denoising evaluation, using frozen DINO patch representations and a timestep adapter as spatial conditioning without an autoencoder, achieving FID 1.46 on ImageNet-256 after 600 epochs and 1.48 on ImageNet-512 after 680 epochs.

Rethinking Cross-Layer Information Routing in Diffusion Transformers

DAR replaces incremental residual addition in diffusion Transformers with timestep-aware attention over historical sublayer outputs, achieving an unguided ODE FID of 7.56 after 600K steps on ImageNet 256×256, a 2.11 improvement over the SiT baseline trained for 1.75M steps, while remaining compatible with REPA's representation-alignment loss.

Revisiting Diffusion Fine-Tuning for Unsupervised Domain Adaptation

MUSE separates diffusion-generator adaptation into a shared semantic branch trained only with source-label supervision and a target-conditioned style branch, allowing one source-involved fine-tuning process to generate bridge data for multiple targets; it improves average downstream UDA accuracy on three closed-set classification benchmarks while reducing generator fine-tuning time to roughly one-half to two-fifths of Terra's cost.


🎬 Video Generation (8)

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

SplitMoE divides the sparse experts of a video diffusion Transformer into semantic and generic groups, guides semantic routing with clean VAE features, and improves video quality over same-source baselines with 27B total and 14B activated parameters, rather than forcing every expert to process an equal number of tokens.

GLARE: Generating Listening Heads with Appropriate Reactions

GLARE models listener feedback such as nodding and smiling as typed temporal events, improving reaction consistency on RealTalk and Seamless through prosody-conditioned flow matching and training-time reaction supervision, while its metrics measure single-reference consistency rather than uniquely correct social responses.

Motion Forcing: Decoupling Ego and Object Motion via Sparse Inputs for Structured Video Generation

Motion Forcing first converts sparse object trajectories and camera motion into dynamic depth, then renders RGB video with a shared diffusion backbone, achieving FVD 157.8, FVMD 205.2, and Physics-IQ 33.2 on Waymo while trading some distributional visual similarity for better motion consistency and physical plausibility.

PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos

PhyProbe trains a lightweight scoring head on a frozen video encoder, combining ranking, severity regression, and endpoint anchoring from heterogeneous sources to score individual videos; it achieves the highest pairwise accuracy on three of four physical benchmarks, but remains a proxy for perceived plausibility rather than a test of physical laws.

PISCO: Precise Video Instance Insertion with Sparse Control

PISCO propagates a few instance keyframes into existing footage through variable-density conditioning, pre-encoding frame completion and post-encoding masking, and depth and appearance augmentation; on PISCO-Bench, first-and-last-frame control with its 14B model reduces whole-video FVD from VACE's 371 to 204, although the compared methods receive different inputs.

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models

DyMoS subtracts an adjustable bias from non-reference-query-to-reference-key self-attention logits in the text-conditioned branch during early image-to-video denoising, raising Wan 2.2's VBench Dynamic Degree from 51.7 to 64.8, although stronger dynamics do not automatically imply more realistic motion or higher reference fidelity.

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

ReMind trains video generators to recover states across interruptions using event frame graphs, reliable-anchor training, and KV caches that retain original spatiotemporal addresses, reaching the highest STEVO-Bench Total of 300.9, but only 14.5% State Progress; the evidence primarily supports short-horizon recovery rather than long-term unobserved physical evolution.

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

ViTeX-Bench uses paired training data from real videos, a frozen evaluation split, and 13 metrics across three axes to distinguish text correctness, motion stability, and background preservation, while providing a motion-aligned glyph-conditioned reference editor, ViTeX-Edit-14B, with the highest mean character accuracy among the evaluated video-native editors rather than the best performance on every metric.


🧩 Multimodal VLM (9)

Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

RWM makes VLM agents infer the previous action from textual belief states derived from consecutive observations and converts action-transition consistency into training feedback; the full configuration achieves the paper's reported 0.81 Overall success rate across four task families, but its gains jointly involve retrospective reasoning, Bi-Level GAE, external judging, and SCR rather than SCR alone.

Binding Multiple Modalities via Multimodal Wasserstein Barycenter

BaryBind replaces a fixed modality anchor with a learnable Wasserstein barycenter map and applies volumetric contrastive learning to modality gap vectors around that anchor, improving zero-shot retrieval and classification on a VAST backbone, although its published dual derivation, potential cancellation, and some experimental numbers require clarification.

Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

Exemplar2VQA compiles spatial QA exemplars into reusable simulator capture and geometric annotation programs, using four-role collaboration and execution feedback to reduce generation errors; approximately 10K synthetic examples raise the author-reported average score of Qwen2.5-VL-3B on the multiple-choice subset of VSI-Bench from 35.3 to 42.9.

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

MulTaBench constructs a benchmark of 20 image-tabular and 20 text-tabular tasks by requiring both complementary multimodal signal and gains from target-aware representations over frozen embeddings, finding that adaptation gains generalize to new learners but are not consistently significant on every selected dataset.

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

SemMSA feeds incomplete language, visual, and acoustic evidence into a frozen LLM, recursively appends continuous hidden states as auxiliary semantics, and learns robust representations through within-instance kernel spectral alignment and cross-instance separation, achieving average Acc-2 scores of 74.36/73.91 on MOSI, 79.61/79.38 on MOSEI, and 75.46 on SIMS across ten intra-modal missing rates.

Spherical Interpolation for Backward-Compatible Multimodal Representations

Without rebuilding the old gallery index, Procrustes first aligns new-model queries to the old space, then spherical interpolation combines them with old-model queries for the same input; endpoint complementarity accounts for most average retrieval gains, while support-set weight selection further improves compatibility success rates.

The Alignment Illusion in Multimodal Large Language Models

Using norm-matched noise and irrelevant images in 13 multimodal large language models, this paper shows that shared MLP weights can produce high similarity and introduces the principal-angle gap to distinguish one-directional collapse from multi-directional structure, while emphasizing that geometry alone cannot establish the use of question-relevant visual content.

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

Argus compares uncertainty methods using unified single-step GUI click records under different observable interfaces, finding stronger cross-dataset method-ranking transfer at a fixed model than across models or from open weights to API-only systems, while error discrimination cannot replace probability calibration or spatial coverage checks.

Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs

Through representational analysis and activation interventions on controlled successful trajectories, this paper explains how audio-visual LLMs bind “who says what” through temporal and position IDs, then uses an existing active speaker detector for visual prompting and optional lightweight fine-tuning to improve conversation understanding, with training-free gains depending on visual-marker grounding capabilities.


🧠 VLM Reasoning (2)

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

LIFT extracts the hidden-state difference at identical answer tokens with and without a reasoning trace from a base large language model, then injects it into the corresponding vision-language model's language layers without updating the backbone, raising six-benchmark average accuracy from 66.6/67.8 to 68.3/69.0 with static intervention and 69.1/69.4 with vector adaptation.

OneCanvas: 3D Scene Understanding via Panoramic Reprojection

OneCanvas lifts multi-view image patch features into 3D and places them as separate continuous tokens in a shared panoramic coordinate system, using native RoPE for angles and frame order and an additive embedding for metric position, then combines synthetic spatial pretraining with real-scene QA adaptation to score 65.3, 71.3, and 72.1 on SQA3D, VSI-Bench, and SPBench, respectively.


⚡ VLM Efficiency (2)

Beyond Selection: Token Parameterization for Extreme Visual Token Compression

Rather than selecting a few original patches, Braco constructs a compact visual interface using a DCT low-frequency backbone, basis-coordinate embeddings, budget-dependent coordinate organization, and sparse-pooled spatial residuals; at 576→16 tokens, it retains a Vanilla-normalized aggregate score of 94.0 while reducing full-pipeline prefill latency to 40.59 ms.

G2TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

G2TR uses VAE latent anchors to select and merge understanding-side visual tokens before LLM prefill in separate-encoder unified multimodal models; retaining 50% of the tokens yields 99.0% relative-average understanding performance and 98.0% editing performance on BAGEL, with approximately 1.94× lower prefill FLOPs, but not lossless performance on every task.


🎵 Audio & Speech (5)

Audible World Models: Spatially Aware Sound Generation for 3D Worlds

Audible World Models connects panorama generation, semantic source parsing, dry-audio synthesis, and geometric acoustics to store sound as persistent 3D world state; it produces more accurate motion-dependent directional cues while retaining strong semantic alignment, but currently remains an offline construction system for static worlds.

OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations

OpenWhistle organizes years of underwater recordings from one dolphin group into a pretraining corpus containing approximately 180,000 whistles and an expert-annotated subset of 8,354 whistles, establishing session-split classification and detection benchmarks with frozen-representation linear probes; corpus-pretrained Wav2Vec2.0 achieves 81.1% classification accuracy and 75.6% detection mAP, without decoding the meaning of dolphin communication.

Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound

Passing resamples a monorail-window recording into a nonlinear journey with branching playback, lets viewer presence indirectly change the visual path, and uses an installation-adapted SpecMaskFoley to generate sound in real time, presenting an exhibition case study and creative observations rather than a controlled experiment in acoustic reconstruction accuracy.

Reconstructing the Vocal Tract with Differentiable Acoustic Simulation

The paper models the vocal tract as a differentiable dynamic acoustic tube and combines frequency-domain solving, smooth turbulence gating, and neural geometry parameterization to infer area functions and MRI visualizations from speech; its default forward synthesis runs at 71.5 times real time, but the reconstructed geometry is an acoustically compatible explanation, not unique anatomical ground truth.

SENSE: Semantic Neural Speech Synthesis from Brain Dynamics via Spatial Graph Encoding

SENSE combines electrode-geometry graph encoding with EEG semantic conditioning learned only from congruent trials to improve waveform reconstruction on the passive-listening N400 dataset; unseen-subject WER falls from FE-Phoneme's 1.1231 to 0.9948, but remains far from reliable language recovery.


🔎 AIGC Detection (1)

Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance

The paper exactly characterizes the optimal robust separation available to pixel-only passive provenance verification through total variation distance, separates this statistical ceiling from interface information exposure in a deployed verifier, and validates the identity in a finite domain without establishing a robustness certificate for realistic images.


🧊 3D Vision (8)

DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding

DirectUV generates mesh textures directly in the latent UV space of a frozen Flux VAE, embeds 3D surface coordinates into each attention head's rotary positional encoding, and uses the reference image and coarse UV only as conditions, achieving 25.13 PSNR and 0.9527 SSIM on GSO while improving consistency across UV islands and texture detail.

M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals

M-plicits reconstructs oriented point clouds with a full-domain coarse SIREN and progressively band-supervised residual SIRENs, reusing this hierarchy for tracing, mesh extraction, and attribute mapping to deliver real-time rendering and strong synthetic measurement-noise robustness with compact models, without winning every accuracy or speed metric.

NRF-GS: Neural Residual Fields for Expressive and Compact Gaussian Splatting

NRF-GS replaces per-Gaussian spherical-harmonic appearance with shared low-/high-frequency neural residual branches, achieving higher average PSNR with fewer Gaussians under unchanged densification and pruning rules, but not uniformly better perceptual quality or rendering speed.

Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction

AFG uses changes in adjacent frames' internal features to generate a positive frame-level write weight in frozen CUT3R/TTT3R inference, multiplying it with the token gate to reduce redundant overwriting and lower mean KITTI ATE from 68.54 m to 43.96 m without discarding frames, retraining, or growing memory, although not every sequence or error metric improves.

SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding

SceneScaffold allocates a fixed budget of 100 visual tokens to details, entities, spatial references, relations, and a global summary, organizing scene evidence before language reasoning and improving ScanRefer / Multi3DRefer mIoU from 43.3 / 42.7 to 47.0 / 47.9 over 3D-LLaVA.

Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

Seeing Speech learns facial motion along spreading, opening, and protrusion coordinate directions, retrieves local motion features through a speech–articulatory memory, and composes them into 3D mesh animation through topology-aware residuals, reducing facial and lip reconstruction errors on VOCASET and TFHP without recovering the physiological mechanisms of actual speech organs.

Seg3DParts: Segmentation-Grounded Controllable Part-Level 3D Generation

Seg3DParts injects externally supplied part segmentations into a two-stage 3D generator and exchanges structural information across part latents, directly producing separate meshes in a space that preserves training-time relative coordinates; its strength is controllable decomposition with less interpenetration, not automatic discovery of unknown parts.

Towards Generalizable 3D Anomaly Detection via Relational Inconsistency Modeling

GRIM learns a defect criterion from controlled local relational violations in normal point clouds through edge-aware graph refinement and within-sample cluster-deviation modeling, achieving 97.4/94.5 object-/point-level AUROC on Anomaly-ShapeNet and 83.6/89.6 when transferred directly to Real3D-AD.


🎯 Object Detection (1)

Beyond Normal References: Discriminative Few-Shot Anomaly Detection

IDEAL uses a fixed set of few normal and masked anomalous references to suppress local normal variations, extract diverse intrinsic deviation directions, and score their projections, improving seen and unseen anomaly detection without target-domain retraining; under N1A1, it achieves 96.3 / 96.8 image-level / pixel-level AUROC on MVTecAD.


✂️ Segmentation (3)

Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement

The paper converts localized segmentation disagreement between prompt templates into binary supervision and adapts open-vocabulary segmentation with region-localized preference optimization and outside-region consistency; with a ground-truth-based preference oracle, CAT-Seg-L improves its mean MESS mIoU from 35.26 to 45.88.

Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation

AdvPCS examines vulnerabilities in promptable concept segmentation through prompt variation, concept perception, and temporal memory, reducing three-model average video mIoU on SA-CO with text prompts from 70.24% to 4.61% in controlled digital-input experiments, without establishing failure guarantees for arbitrary prompts, models, or real-world inputs.

When Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-Labeling

FTC-Seg combines prototype-based residual feature reconstruction with pseudo-label thresholds driven by labeled accuracy and class-distribution bias in teacher–student segmentation, improving DINOv2-S over UniMatch V2 by 2.09–8.54 mIoU percentage points across four benchmarks with 2% labels, without winning every configuration.


🖼️ Image Restoration (4)

Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion

Rather than treating source images as ideal fused images, RCS-Fusion learns a relational supervision space from frozen DINO/CLIP features and trains different fusion backbones with shared, complementary, and conflict-coordination constraints; its Transformer variant achieves 0.730 mAP in detection using fused M3FD images.

i-DEQ: A stable inertial deep equilibrium model for image restoration

i-DEQ integrates restarted inertial optimization into training and inference for an explicit-energy deep equilibrium model, achieving reconstruction quality close to ELDER with substantially shorter equilibrium-solving time in small image restoration experiments, but its theory establishes conditional stationarity complexity and practical training is not always stable.

Learning Where and What to Restore for Composite Image Restoration

CART allocates restoration compute through window-level spatial routing inside a U-Net, then controls image-level channel routing with task features shaped by multi-hot degradation supervision, achieving 29.95 dB average PSNR on CDD-11 while trading its speed advantage for more parameters and higher peak memory than MoCE-IR.

Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment

The paper represents images as Gaussian distributions in a low-dimensional perceptual space and fits human choices through noisy candidate–reference distance comparisons; the CORnet-S Multi observer reaches 95% of its own above-chance performance at 3, 9, and 96 dimensions on BAPPS, PieAPP, and NIGHTS, respectively, suggesting task-dependent perceptual structures.


🛰️ Remote Sensing (2)

PolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing Imagery

Using two aerial-image datasets, four object categories, and 11 public baselines, PolyTopoBench separates region, ring-boundary, and complete ring-structure evaluation, showing that complex instances still have an 84%–100% topology error rate even when their IoU reaches 0.9.

SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery

SatNav automatically constructs 118,494 long-horizon navigation episodes from satellite imagery and OpenStreetMap and uses SwiftVLN to isolate memory designs: retaining only short-term memory reduces success rate by 21.3 percentage points on both test splits, while real-flight transfer remains supported by just 18 preliminary episodes.


🧑 Human Understanding (7)

BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion

BiMoGen combines a unified motion-text vocabulary with a bidirectional masked diffusion model for motion generation and captioning, using correspondence learning before conditional generation and generation-aware self-correction to achieve T2M R@1 of 0.555 and M2T CIDEr of 60.2 on HumanML3D, without outperforming existing unified models on every metric.

FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery

FactorizedHMR first deterministically estimates the torso, shape, and camera-space trajectory, then fixes that structure while masked flow matching completes limbs, the head, and world motion; its clearest benefits concern severe occlusion and trajectory drift rather than uniformly superior conventional metrics.

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

GazeFlow generates whole gaze trajectories through flow matching conditioned on local video features and global task tokens, reaching F1 0.491 and average angular error 9.01 on EGTEA Gaze+ while improving displacement statistics, but its default model uses future frames and is not directly an online eye tracker.

PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition

Instead of adding an RGB action-recognition branch, PoseBridge preserves visual semantics inside pose estimation before they are compressed into joint coordinates, then transfers them through skeleton-conditioned bridging and semantic prototype adaptation, improving over the strongest compared baseline by 13.3–17.4 percentage points across eight Kinetics-200/400 splits.

Re:Cognize: Open-Set Comic Character Re-Identification

Re:Cognize separates identity emergence from identity maintenance through four gallery protocols over the same reading-order stream, finds that growth under predicted labels harms recognition with random seeds, and uses Re:Cast's character averages, page evidence, and cross-page binding to recover part of the benefit of correctly labelled updates under explicit conditions.

STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts

STRIDE precompiles textual crowd contexts into behavioral questions, deterministic measurement functions, and expected ranges, then checks generated trajectories against them; its benchmark contains 936 scenarios, Text-Crowd achieves an overall score of 0.645, and humans agree with benchmark answers on 80% of pairs, although the latter result comes from a limited annotation subset.

Towards Unified Dynamic Face Landmark Detection

The paper describes landmarks using a face part and a normalized sequence position, then uses image-conditioned queries and iterative decoding to train one model across annotation formats and predict landmarks on demand; the unified ViT-B achieves full-set NME of 4.05, 2.80, and 1.02 on WFLW, 300W, and AFLW-19, respectively, offering unified training and interfaces rather than superiority on every metric.


📹 Video Understanding (1)

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

TT-VidT reconstructs later frames using a first-frame anchor that retains full spatial information and compact learned motion tokens, substantially improving motion-sensitive tasks under a shared recipe of roughly 1.7 million videos and 8 epochs while reducing TT3D encoder compute to 456.1 GF, without establishing universal video-understanding superiority or strict appearance–motion disentanglement.


🚗 Autonomous Driving (5)

AutoExpert: Automating 3D LiDAR Annotation from Expert-Crafted Guidelines

AutoExpert turns textual expert rules and a few 2D examples into an adapted detector, then uses a vision-language model to supply instance-specific size and orientation priors for fixed-size multi-hypothesis search, achieving 25.4 mAP3D on AutoExpert-nuScenes without target-task 3D training annotations.

DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution

DriveHierarchy organizes vision-language driving capabilities into perceptual grounding, contextual memory, mental reasoning, and closed-loop execution, diagnoses 15 models through unified open-loop tasks and 100 interactive simulation scenarios, and shows on Qwen3-VL-8B that repairing selected open-loop weaknesses can raise the closed-loop composite score from 0.902 to 8.021, without establishing real-world driving safety.

Driving Video Retrieval for Complex Queries with Structured Grounding

STRIVE-D calibrates motion rules against a weak-label ranking objective on separate driving videos, reuses or adapts those rules per query, and fuses them with visual and lexical retrieval, raising DrivingDojo Acc@1 from the main table's strongest dense baseline of 14.6% to 26.8%; it does not simply ask an LLM to inspect videos and determine true geometry.

Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds

MaDE learns inverse dynamics and physical residuals from state transitions without control labels, then corrects predicted trajectories in control space so that outputs satisfy its learned discrete dynamics map by construction while reducing inequality violations on a best-effort basis; known-model residuals fall substantially on inD, but position errors increase, without establishing true-physics correctness or safety.

Progressive Risk Estimation for Accident Anticipation

PRE-ACT uses continuous distance-to-event progress supervision and within-video clip ranking to improve traffic risk prediction from past-only sliding windows, raising CAP mAUC from TOP's 0.429 to 0.481 under FPR ≤ 0.1 and introducing a Separation Score to examine false-alarm tendencies across complete risk curves.


🤖 Robotics & Embodied AI (6)

Action Chunking Proximal Policy Optimization with Feedback Correction

ACPPO-Corr extends PPO with a low-frequency action-chunk planner and a stepwise feedback corrector while retaining a state-only value network, improving final normalized interquartile mean over PPO by 30.4% across 25 simulated robotics tasks—not by 30.4 percentage points of success rate.

Behavioral Foundation Models for Quality Diversity

BFM-QD freezes an offline-pretrained behavioral foundation model, builds a repertoire of high-quality, diverse behaviors by searching latent codes rather than policy weights, and uses closed-form backward inference for small directed mutations, substantially improving sparse navigation and contact-rich manipulation while still requiring environment rollouts and pretraining investment.

GeoWind2Plan: Mission-Time 3D Urban Wind Prediction for Energy-Efficient UAV Planning

GeoWind2Plan converts building geometry and background wind into a mission-corridor 3D mean wind field and jointly optimizes UAV path and speed; on block A at a background speed of 4 m/s, CFD-evaluated energy decreases from \(13.3\pm1.8\) for wind-agnostic planning to \(12.5\pm1.5\) Wh/km, while wind inference for a 20% corridor takes about 3 seconds.

Language-Conditioned World Modeling for Visual Navigation

The paper introduces a language-conditioned visual navigation dataset and compares two alternative approaches—“diffusion world model + latent-space actor–critic” and “unified autoregressive action/observation prediction”: the former has stronger image structural fidelity in seen environments, whereas the latter predicts offline trajectories better in unseen environments, without validating real closed-loop control.

Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

The paper encodes object motion histories as linked linguistic captions, sparse 3D anchors, and visual anchors, improving semantic trajectory retrieval in long egocentric videos while trading expensive one-time construction for sub-second online queries.

ProCompNav: Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries

ProCompNav collects same-category objects in an unknown environment, then recursively identifies distinguishing attributes and asks users yes/no questions, achieving success rates of 23.7%, 28.1%, and 17.0% across the three simulated CoIN-Bench splits while substantially shortening user-simulator responses.


🎮 Reinforcement Learning (10)

AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning

AlphaPareto encodes the current alpha pool with a frozen LLM, conditions token-by-token MaskPPO formula search on existing signals, and uses adaptive multi-objective rewards for predictive power, ranking stability, perturbation robustness, and diversity, achieving out-of-sample ICs of 3.92%, 5.70%, and 10.10% on CSI300, CSI800, and the full market.

From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning

QTPT replaces behavior-action prediction in a parameter-shared Transformer with context-conditioned Q-target regression, enabling a fixed-weight model to improve decisions using reward and transition history from the same task; gains are pronounced under weak behavior data, but the theory requires coverage and realizability rather than guaranteeing success on arbitrary low-coverage datasets.

HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning

HaM-World integrates selective history memory and a Soft-Hamiltonian prior on part of the latent coordinates into one planning model, ranking first on four and second on two state-based control tasks, but its imagined-prediction advantage is confined to planning-relevant horizons, and memory has a substantially larger observed ablation impact than geometry.

Learning Chance-Constrained MDPs with Bellman Distributional Certificates

The paper expresses cumulative-cost violation events through remaining-budget Bellman recursions, obtains nearly matching deterministic-policy sample complexity under bounded successor support and a certified planning oracle, and combines local approximate-KKT optimization with independent validation for stochastic policies; synthetic and power-system simulations distinguish statistical conservatism from conservatism induced by expected-cost surrogates.

Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching

LFIRL converts local action scores from a frozen diffusion policy into soft-Q gradient supervision, then sequentially fits values, calibrates state offsets, and regresses rewards, removing alternating reward–policy optimization and typically running about 2–3 times faster than the fastest baseline when diffusion pretraining is included, without leading every reward-quality metric.

Modeling Quantum Neural Network Gradient With Reinforcement Learning

RLQ-Grad uses a classical PPO policy with spectral normalization to generate surrogate update signals from a quantum neural network's training state, improving validation accuracy and reducing additional differentiation overhead in the reported classification simulations without proving recovery of true gradients or universal avoidance of barren plateaus.

Replay-buffer engineering for noise-aware quantum circuit optimization

The paper shifts quantum circuit optimization from changing the agent to reusing experience more effectively: annealed reliability-aware replay improves sample efficiency, blockwise evaluation reduces quantum–classical calls, and buffer-only transfer accelerates noise adaptation, although accuracy, gate-count, and convergence advantages differ across tasks.

Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning

Safe Score Matching (SSM) uses learned Hamilton–Jacobi safety values to switch a diffusion policy between reward and recovery training targets, combining multimodal action modeling with low violation costs in online safe reinforcement learning, although its hard-constrained formulation does not establish a strict safety guarantee for the learned policy.

Trust Guided Decision Transformer

TGDT first filters trustworthy history suffixes using rolling next-state prediction errors on realized transitions, then ranks their candidate actions with a frozen IQL critic, improving return and persistent-error behavior in long-horizon navigation without guaranteeing closed-loop coverage or safety.

Verifying Neural Networks with Reinforcement Learning

Rsb uses graph-structured observations and actor-critic learning to reweight Fsb neuron-branching scores without replacing verification logic, increasing solved problems from 148 to 165 among 600 challenging test instances, although inconsistencies in rewards, feature masks, and some statistics need to be distinguished from the performance gains.


🎁 Recommender Systems (2)

Goal-Conditioned Supervised Learning for Multi-Objective Recommendation

MOGCSL preserves multiple future session rewards as a goal vector, trains a goal-conditioned next-item predictor with ordinary cross-entropy, and selects inference goals using statistics or two CVAEs, improving purchase prediction and training cost without dominating every click metric or guaranteeing that specified goals are attainable.

Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models

The paper reframes personalized generation as selecting suitable candidates already produced by a generator, using an MLP with 2.8M parameters by default to rank frozen generator representations; it beats four generalist reward models across nine datasets but remains behind a dataset-finetuned 8B reward model, while single-trajectory decoding guidance recovers only part of the selection gain.


🔄 Self-Supervised Learning (6)

Beyond Missing Rates: Rethinking Incomplete Multi-View Clustering with Protocol Divergence

The paper exposes evaluation-protocol divergence in incomplete multi-view clustering through effective missing rate and complete-sample proportion, and introduces CRAFT, which fuses only each sample's observed views: it wins 12 of 13 information-matched incomplete-training conditions, while a separate evaluation tests cross-protocol deployment of checkpoints trained with all views available.

Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

SFM approximately interprets linear gradient matching on frozen pretrained vision backbones as matching relative vectors between class centers, then supervises synthetic images with global statistics computed once, improving classification with one image per class; its approximately 10-fold memory reduction and fourfold speedup compare single-augmentation SFM against ten-augmentation LGM, rather than equal-budget configurations.

Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

Using pixel-PCA projections and frozen-encoder probes, this paper studies the distribution of handwritten text recognition signals: high-variance pixel directions support transcription better on six Latin-script benchmarks, while real-data-pretrained MAE achieves 5.5% mean probe CER and 4.5% mean CER after adding a language decoder and full fine-tuning.

I Have a Stream: Making Self-Supervised Learning Work on Continuous Video

StreamMAE preserves MAE's pixel reconstruction objective while adapting temporally ordered video training through stronger regularization, example-level DataDrop, and motion-biased two-stage cropping, improving classification, segmentation, and depth transfer on WT++12h and approaching same-data offline i.i.d. MAE on most metrics.

Latent Video Prediction for World Modeling: An Evaluation Uncovering Intriguing Favorable Evidence

Under a unified readout protocol, this study evaluates four frozen video encoders, finds that V-JEPA variants better preserve usable action semantics under degraded inputs and support more reliable planning in limited simulated manipulation experiments, but neither evaluates generation quality nor causally isolates the pretraining objective.

Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding

SGMA combines content-adaptive quadtree tokenization with cross-scale structure-conditioned masking for MAE pretraining, preserving scientific microstructure within a fixed sequence budget for standard ViTs; SGMA-SAM reaches 95.68% Dice on SpringXCT and SGMA-SAM 2 reaches 83.21% on PAIP, while the reported maximum 24.8× inference speedup comes from a separate, shorter-sequence configuration.


📐 Optimization & Theory (16)

Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients

The paper analyzes finite-horizon calibrated Adam without bias correction under generalized smoothness and second-moment ABC noise, shows that the confidence exponent \(\delta^{-1/2}\) cannot generally be removed, and separates expectation upper bounds for \(p<1\) from worst-case lower bounds for a specified algorithm family when \(1\leq p<2\).

Amortized Optimal Transport from Sliced Potentials

The paper uses inexpensive one-dimensional Kantorovich potentials as features, predicts original-space potentials with shared linear coefficients, and reconstructs approximate transport plans; RA-OT and OA-OT reduce training costs and support variable numbers of atoms, but are not universally fastest at inference or best in generation quality.

Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods

The paper separates surrogate geometry from utility sensitivity through the pullback Fisher tensor of the posterior map, then replaces TuRBO's kernel-lengthscale weights with regularized local Fisher diagonal weights, achieving competitive SE-kernel benchmark performance without a regret or global convergence guarantee for outer Bayesian optimization.

Bidirectional Information Flow (BIF) - A Sample Efficient Hierarchical Gaussian Process for Bayesian Optimization

BIF constructs a soft parent prior from child Gaussian processes' acquisition maps and allocates real parent responses to children using uncertainty-aware weights for continual learning, improving reconstruction and learning trajectories in low-budget composite tasks; this feedback comprises biased pseudo-responses, not true subtask decomposition.

Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models

The paper decomposes the noncommutative residue of two training updates into a signed token readout, localizing and intervening on training-order differences under local SGD conditions; targeted interventions close the held-out loss gap by a median 32.0% on Qwen-3-4B, while paired-endpoint order assignment reaches 66/72 = 91.7%.

Cumulative-Goodness Free-Riding in Forward-Forward Networks: Real, Repairable, but Not Accuracy-Dominant

The paper proves that cumulative goodness exactly attenuates deep blocks' local discrimination gradients on examples already separated upstream, and repairs block health through history removal, hardness gating, and gradient compensation, but finds no accuracy-dominant benefit from these repairs; MGC even reduces CIFAR-100 Stage-1 single-crop accuracy by 1.05 percentage points.

Dynamic Regret in Online Convex Optimization with Indicator Switching Costs

The paper aggregates distributions from randomized lazy FTRL learners restarted at dyadic scales through a movement-aware meta-learner, uses maximal coupling to control actual action changes, and obtains near-optimal expected dynamic regret for piecewise-constant comparators alongside a separate guarantee for comparators with small path length.

Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models

EoupCT uses a frozen pretrained model and learnable soft prompts to construct differentiable knowledge proxies vulnerable to new-task updates, then combines distillation with gradient projection to protect historical tasks and general capabilities; experiments on SuperNI/MMLU subsets across six models improve retention, but its “first-order-only” and “absolute zero forgetting” claims require qualification.

Finite-Sample Performance of Gradient Descent in Logistic Regression with Gaussian Design

Under well-specified Gaussian logistic regression, the paper combines population curvature analysis with approximate invertibility of the empirical gradient to prove linear convergence of gradient descent to a statistical-error neighborhood, and obtains a sharper high-dimensional error bound by estimating direction and norm separately; large-step acceleration is only local, and the stated near-optimal regime contains a condition discrepancy that must be retained.

FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models

FluxLite jointly adjusts sparse jump rates and compensating weights without retraining a discrete diffusion model, preserving the target marginal-distribution path through graph divergence while mitigating particle-weight degeneracy with the local HEU rule or a small nonnegative quadratic program, D-VCG.

Browse all 16 Optimization & Theory papers →


📐 Learning Theory (9)

Adaptive inference for functionals of M-estimands

This paper extends inference for smooth functionals of nonparametric M-estimands to adaptive data with known sampling policies, combining one-step bias correction with per-round conditional variance stabilization, or self-normalization when variance converges to a random nonzero limit, to construct asymptotically valid confidence intervals; dynamic pricing simulations show undercoverage for ordinary GLMs and GLMs using inverse propensity weighting alone.

Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS

The paper turns significance testing for partial least squares (PLS) supervised subspaces into tests of held-out predictive correlation, supplies an NB-corrected fast approximation with an explicit empirical scope and a finite-sample valid permutation test under the omnibus independence null, and separates ordered inference on original components from interpretation of rotated coordinates.

Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models

COCOCO separately calibrates the concepts and task labels of neuro-symbolic concept-based models, then removes unsupported candidates through one deduction–abduction intersection step, producing smaller, mutually consistent sets with explicit coverage-loss bounds and supporting adaptive size budgets through COCOCO*.

Cost-Aware Best-LLM Identification using Dueling Feedback

The paper formulates fixed-confidence identification of the best model under pairwise preferences as a heterogeneous-cost dueling bandit, allocates comparisons using rejection evidence per unit cost, tracks sampling targets, and stops through best-arm hypothesis testing, with an almost-sure asymptotic cost guarantee; however, the confidence-interval variant DCTAC costs less than the main algorithm DCTAS in the finite experiments.

Estimation of the Label-Noise Transition Matrix with Performance Guarantees via Selective Classification

The paper learns high-purity subsets through one-sided selective classification and estimates the transition matrix from the empirical frequencies of all noisy labels within those subsets, decomposing error into subset contamination, learning suboptimality, and accepted-sample fluctuations without requiring accurate pointwise class-posterior recovery.

Even Sharper Bounds for Transductive Learning and Its Applications

The paper proves Bernstein-type supremum concentration through a two-parameter entropy closure for the swap walk, then uses local complexity to analyze transductive empirical risk minimization, removing both earlier sample-imbalance restrictions and an extra logarithmic confidence factor under bounded-loss and appropriate localization conditions, with applications to realizable VC classification and the empirical kernel spectrum.

Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers

The paper establishes the minimax rate of nonparametric in-context learning on heterogeneous manifold mixtures whose complexity grows with sample size, and proves that a structure-informed transformer approximates a tangent local-polynomial estimator attaining that rate; optimality of a trained predictor still requires sufficient training tasks and a sufficiently small empirical-risk optimization gap.

The Price of Locality: Why Forward-Forward Underperforms Backpropagation?

This mechanism-analysis paper combines conditional local convergence theory, spectral analysis of representation kernels and error signals, and interventions on training groups and update spectra to explain why the Forward-Forward Algorithm (FFA) often trails back-propagation (BP); relaxing gradient isolation between CNN12 layers raises accuracy from 53.02% to 78.74%, still below BP's 85.67%.

Two-Fidelity Best-Action Identification for Stochastic Minimax Tree

Under a known fast-evaluation bias envelope and slow evaluations unbiased for true node minimax values, 2FFS uses endpoint certificates and recursive budgets to choose adaptively between expansion and sampling, achieving high synthetic-tree accuracy with substantially fewer node visits, while stopping and cost guarantees require additional regularity conditions.


🔗 Causal Inference (1)

Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition

The paper separates root cause ranking accuracy into whether the cause enters the candidate set and whether it ranks highly once retrieved, audits six suites from four datasets, and uses multi-signal retrieval plus single-call LLM reranking to show why the two failures require different remedies rather than simply more causal structure or reasoning capacity.


🔬 Interpretability (10)

Can Circuit Alignment Predict OOD Generalization?

The paper defines Circuit Alignment Score (CAS) as same-class cross-domain circuit similarity minus cross-class circuit similarity, achieving a mean Spearman correlation of 0.88 on PACS for source-domain model selection without target data, but its OOD ranking consistency requires an additional monotonicity assumption and is not a fully data-free weight diagnostic.

Deep Minds and Shallow Probes

The paper derives a polynomial hierarchy of shallow probes from coordinate symmetries at the final readout, uses low-rank CP probes to read interaction concepts and probe-visible quotients to transfer concept readouts, and improves cross-token agreement AUROC by 16.8–20.0 percentage points while explicitly separating transfer accuracy from concept coverage.

Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations

Using a shared generator to separate latent geometric changes from generator-concept changes, the paper compares prediction stability and smoothed-classifier certificates for standard and concept bottleneck models, finding that interpretability offers no uniform robustness advantage but changes where sensitivity appears; the conclusions remain conditional on generator-native inputs and specific task settings.

Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

After freezing an LLM, the method calibrates an attention head and selects answers using pre-RoPE dot products between the final prompt query and option-ending keys, exposing a gap between internal selection and final output; particular zero-shot configurations gain 27.4 and 49.8 percentage points on HellaSwag and HaluDialogue, respectively, but not every model benefits.

Parameter symmetries determine representational geometry in overparameterized nonlinear networks

For one-hidden-layer nonlinear networks, this paper reduces parameter symmetries preserving the global function to feature addition, duplication, and scaling, proves that they can substantially reshape representational geometry, and gives sufficient conditions for minimum-norm selection to restore geometric identifiability.

PersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona Representations

PersonaManifold models LLM persona activations as a curved, anisotropic low-dimensional manifold, using graph shortest paths weighted by local metrics to measure similarity and interpolate personas; combined BST triplet consistency improves over Euclidean distance by 5.1–6.1 percentage points across three 7–8B models, primarily benefiting high-deviation persona pairs.

PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders

PULSE identifies SAE features associated with demonstration utility from a small labeled discovery set, then uses them for complete-set ranking and cacheable per-example retrieval, improving selection across six text tasks while providing associational rather than causal-intervention evidence.

SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery

SMILE identifies decomposable structure in data, fits a network with fixed symbolic activations, and recovers compact expressions through pruning, constant refitting, and gradient-based rounding, achieving strong symbolic recovery under high noise in SRBench without universally leading noise-free recovery or prediction accuracy.

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

The paper uses standard knowledge distillation as an analytical tool to transfer an event teacher's predictions into an RGB student, finding joint improvements in tolerance to color changes, shape preference, and high-frequency-band noise, but with lower clean accuracy and sensitivity to disrupted geometric continuity rather than comprehensive robustness.

Witness Overlap: Directional Provenance Inside Open-Weight Model Families

Within known same-family open-weight models with aligned parameters, Witness Overlap introduces a third checkpoint, compares each candidate endpoint's update overlap toward the target and witness, and predicts the lower-overlap endpoint as the parent; Frobenius cosine achieves 95.3% orientation accuracy over 1,542 single-witness triplets from 16 LLM families.


📦 Model Compression (5)

CASS: Contribution-Aware Structured Sparsity for Model Merging

CASS uses unlabeled task samples to identify attention heads and FFN neurons with prominent contributions, then filters task vectors or constrains fine-tuning gradients to reduce merging interference, improving Task Arithmetic's average accuracy across 20 tasks on ViT-B/16 from 65.02% to 69.99%.

Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models

CaRE-KD selects Forward or Reverse KL per token using relative teacher–student confidence and uses MC-dropout BALD to reject supervision when the teacher is uncertain but the student is comparatively certain, improving distillation quality in the tested settings and reaching 62.4 on MBPP and 73.9 on GSM8K, although gains are not universal and training becomes substantially more expensive.

Importance-Aware OBS Pruning for Diffusion Models

The paper injects prompt-related spatial importance into the layer-wise OBS reconstruction objective to improve subject fidelity in highly sparse diffusion models without fine-tuning, but its global mask-coefficient ablation has an unexplained inconsistency with the stated exact equations.

Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer

FGQ weights channel reconstruction errors in block-wise post-training quantization of VGGT using normalized squared gradients from its geometric tasks, prioritizing sensitive features during calibration; under W4A4, mean completeness error on 7-Scenes falls from QuantVGGT's 0.085 to 0.059, although not every metric recovers FP16 performance.

Stabilizing the Dynamic Low-Rank Training

SDLRT retains the leading singular directions discarded in the previous iteration and combines compensated basis augmentation with negative feedback on truncation tolerance to mitigate rank collapse under aggressive compression, achieving 85.15% average accuracy on six SuperGLUE validation tasks.


🕸️ Graph Learning (6)

Efficient Dynamic Algorithms for Graph Neural Networks with Non-Linear Propagation

The paper maintains nonlinear propagation fixed points on dynamic undirected graphs using residual certificates, endpoint rescaling, and selective pushes, proves deterministic amortized update bounds under degree-normalized error, and shows that nonlinearity can improve static classification accuracy without substantially increasing dynamic propagation cost.

HoTS: Homophily-Aware Temperature Scaling for Graph Neural Network Calibration

HoTS generates positive node-wise temperatures from predictive entropy and local homophily estimated by an auxiliary GCN, using an idealized CSBM inverse-homophily temperature law as a structural prior; it preserves predicted classes and achieves the best mean ECE of 4.79% across 18 datasets, but does not win on every dataset or coverage level.

Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge

NSR binds entities, relations, and relation chains to readable network units and connections, performing link prediction through relation-association retrieval and explicit graph traversal; it achieves an MRR of 0.8142 on Nations, but its advantages are dataset-dependent, and the efficiency of its accelerated static implementation should not be equated with that of its online brain-inspired dynamics.

Relation-Aware Graph Foundation Model

REEF treats transferable relation semantics as the unit of graph foundation modeling and uses semantic hypernetworks to generate aggregators, classifiers, and dataset projectors, reaching 48.66%, 48.65%, and 47.72% accuracy on three target graphs under 1-shot classification, while retaining clear limitations for unseen relations, zero-shot tasks, and computational cost.

Spectral Reversal: Counteracting Singular Value Bias for Graph Prompting

SRP preferentially reads weak directions through singular-value soft masks at each frozen GNN layer, supplements features discarded by the original weights through null-space PCA, and adds low-rank prompts after aggregation; it achieves 68.11% accuracy on 5-shot Cora with GraphCL pretraining, but does not outperform the strongest baseline on every dataset.

Understanding and Mitigating Under-Confidence in GNNs from the Final Layer

SCAR explains GNN under-confidence through the final classifier, reduces only final-layer weight decay during training, and moves representations toward predicted-class prototypes at inference time using training-set adjacency groups, reducing ECE across multiple node classification settings without guaranteeing unchanged labels or accuracy.


📈 Time Series (11)

AdaST: Adaptive Coupling for Spatial-Temporal Forecasting

AdaST decomposes spatial-temporal inputs into temporal-specific, spatial-specific, and jointly coupled representations, models them separately, and adaptively recomposes them with correlation-modulated gates, achieving the best reported results on four short-term sensor forecasting benchmarks and reducing PurpleAir MAE from the strongest comparator's 0.511 to 0.489.

Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport

SBOG uses Sinkhorn support energy to identify latent boundary anchors and selects locally perturbed candidates with weak support but controlled semantics, raising CARLA's average pointwise F1 across five datasets from 0.362 to 0.432 and improving image out-of-distribution detection; these results measure downstream utility, not recovery of the true outlier distribution.

DiffPTS: Rethinking Diffusion ELBO for Probabilistic Time Series Forecasting

DiffPTS jointly trains history-conditioned Gaussian mean and variance estimators with a denoising network, replacing hand-crafted variance supervision with a Gaussian negative log-likelihood derived from the diffusion ELBO and improving CRPS/MSE across nine benchmarks, including a reduction from 0.378/0.637 to 0.270/0.456 against NsDiff on Traffic, without uniformly dominating calibration metrics.

Kairos: Toward Adaptive and Parameter-Efficient Time Series Foundation Models

Kairos handles local time-series complexity and instance-level spectral differences through mixture-of-size encoding and dynamic rotary positional encoding, then predicts multiple future patches in parallel, reaching a normalized MASE of 0.738 on GIFT-Eval with 53M parameters, without uniformly leading in probabilistic or task-level performance.

Multivariate Time Series Forecasting needs Cross Variable Loss

CvLoss leaves the forecasting backbone unchanged and supplements point-wise MSE with constraints on residual differences between forecast patches from different variables, achieving 111 wins, 2 ties, and 1 loss across the paper's 114 backbone-controlled comparison cells, with no loss computation at inference time.

Progressive Memory Transformer: Memory-Aware Attention for Time-Series

PMT exposes writable memory updated along sliding windows as an explicit mid-range representation and separately supervises tokens, memory, and sequence summaries, achieving 84.4% average accuracy across seven low-label classification datasets while a cue-retention probe verifies that memory carries information across windows.

Raw-Routed Mixture of Adapters: A Causal Intervention for Routing Collapse in Time Series Foundation Models

RR-MoA gives a lightweight gate the input before instance normalization while adapter experts retain frozen-backbone hidden states, repairing an input-side information bottleneck and winning all 54 paired comparisons reported against fixed adapters, although the basic version still trails a linear predictor trained from scratch.

Revitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language Perspective

ViRe renders the same numerical EEG/ECG segment as a waveform and uses frozen CLIP visual features as a query to retrieve evidence from temporal and channel numerical tokens; the mean six-metric score across six benchmarks rises from Medformer's 77.90 to 82.90, but ViRe does not lead on every dataset, and this is not patient-level diagnostic accuracy.

TimeES: Probabilistic and Deterministic Time Series Forecasting via Evolutionary Spectra

TimeES predicts time-varying complex spectral amplitudes with a lightweight network and synthesizes trajectories using fixed Fourier phases and frequency-level random variables shared across the forecast horizon, unifying deterministic and probabilistic forecasting; its average performance on nine probabilistic benchmarks is strong, but ILI clearly deteriorates, and spectral recovery and uncertainty training require careful interpretation.

TimeTok: Granularity-Controllable Time-Series Generation via Hierarchical Tokenization

TimeTok encodes time series into discrete token prefixes with coarse-to-fine semantics, adds detail through granularity-block autoregression, and controls output granularity through conditional flow matching, achieving strong conditional refinement and distributional metrics without leading on every downstream predictive metric.

Browse all 11 Time Series papers →


🏥 Medical Imaging (8)

Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning

Aegis adds a task-relevant synthetic-image gradient update after each client-local epoch to increase sample mixing in linear leakage; it reduces the measured reconstruction rate to 9.38%–11.63% across three MedMNIST modalities, without providing zero-leakage or differential privacy guarantees.

DeepArrhythmia: Segment-Contextualized ECG Arrhythmia Classification via Selective Evidence Acquisition

DeepArrhythmia produces R-peak-aligned beat labels within 10-second ECG segments and uses initial segment confidence to decide whether to acquire numerical and morphology evidence, demonstrating the value of context and physiological grounding across four datasets without outperforming always-rich inference on every metric.

FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs

FOCUS connects harmonized disease labels from ten fundus datasets and the prediction interfaces of three foundation-model families to a shared evaluation layer, finding that ranking, calibration, and subgroup behavior do not align, while MLLM supervised fine-tuning primarily improves binary decisions and calibration with little average AUROC gain.

Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis

IBAHGT combines adaptive message passing on a fixed nearest-neighbor hypergraph with a parallel global Transformer, then uses node-level fusion and information-bottleneck auxiliary training to improve fMRI brain-network classification, achieving 77.6% / 86.3% accuracy on randomly split ABIDE / ADNI, although its mutual-information derivations contain important conditional and sign issues.

Modeling Whole-Slide Images as Dynamic Tumor Microenvironment Fields

TMEvolve models pathology slides as repeatedly updated latent region fields, learning slide representations through intra-region diffusion and concept-guided directed boundary interactions; it achieves the highest mean C-index on three of four survival cohorts, but its “evolution” denotes representation refinement rather than actual tumor chronology.

Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity

Using three-image similarity choices across six histopathology cohorts, MOSAIC directly contrasts same-class cross-domain candidates with different-class same-domain candidates and finds that zero-shot multimodal large language models usually resist acquisition-context shortcuts better than Euclidean distances between pathology foundation model embeddings, without establishing clinical diagnostic competence or reliable cross-domain performance.

SheafStain: Sheaf-Theoretic Schrödinger Bridge for Spatially and Biologically Coherent Virtual Staining

SheafStain adds neighborhood VFM spatial conditioning, overlap-consistency losses, and pathology supervision to an unpaired Schrödinger bridge, then uses adaptive extra patches and weighted stitching to generate IHC from H&E, substantially reducing seams on 1024×1024 assembled regions from two breast pathology datasets without guaranteeing diagnostic correctness.

Tokenizer-Generator Coupling in Medical Image Generation

This controlled factorial study treats the tokenizer–generator–sampler triple as its experimental unit: across 70 generation cells on ChestMNIST-64, reconstruction quality does not reliably predict generation quality, quantizer rankings depend on the generator, and validation-selected reduced-step sampling lowers LFQ-1024 + D3PM FID-192 from 0.44 to 0.09.


🩺 Medical LLM (1)

MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models

MedKIT converts 6,196 timestamped oncology evidence updates into fact-centered, multi-task probes and compares 12 knowledge integration methods across 5 models, showing that recalling an updated fact usually does not imply using it in open-ended reasoning or generation.


🧬 Computational Biology (11)

At FullTilt: Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections

FullTilt feeds aligned 2D tilt-series directly into a visually prompted multiclass 3D detector, replacing volumetric sliding-window detection with cross-tilt row attention, near-zero-tilt query initialization, and training-time geometric augmentation to achieve subsecond zero-shot detection on three real-world cryo-ET datasets, while still requiring simulated-data pretraining and input alignment.

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

NMO tests molecular optimization beyond pharmaceutical priors through three nanophysics tasks constrained by electrode binding, and combines explicit-anchor GGS representations, random-molecule pretraining, and genetic-guided GFNs to discover scientifically promising candidates, although the full model does not maximize AUC on every task.

CellMSA: Context Modeling for Single-Cell Representation Learning

CellMSA aligns cross-batch and related-type cells by gene identity, extracts gene-pair dependencies from low-dimensional context, and uses them to guide target-cell encoding, achieving strong results in label-informed integration, classification without test-label retrieval, and perturbation prediction combined with STATE-ST.

Data-Driven Soft Labeling Scales DNA Read Classification to Whole-Body Cell-Type Deconvolution

Syto replaces single-origin hard labels with the cell-type distribution associated with methylation patterns in reference data, then separately models read classification, sample deconvolution, and proportion calibration; its best reported configuration reduces pseudobulk MSE across 39 cell types from CelFiE's \(3.28\times10^{-4}\) to \(0.88\times10^{-4}\), without establishing clinical cancer-detection performance.

Evolutionary foraging in grids: Intermittent search dynamics emerge in finite, depletable landscapes

On finite two-dimensional grids with non-renewable resources, the empirical distributions of step length, velocity, and turning angle evolve through a genetic algorithm rather than a prescribed power law; the resulting second and fourth displacement moments favor an intermittent-search description, without proving explicit state switching or globally optimal foraging performance.

FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases

FlyAOC evaluates fruit-fly knowledge base curation end to end, from large-scale full-text retrieval to ontology-grounded candidate outputs, showing that multi-agent context partitioning improves recall of known recoverable annotations while semantic scores do not establish exact biological validity.

LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling

LEMON-ZEST clusters evolutionarily conserved fragments into shared tokens and combines stochastic segmentation with a dual-head encoder for global retrieval and residue information, allowing a 200M-parameter model to outperform large baselines on several fold-level retrieval metrics without being best at every classification level or downstream task.

PHOEBI: An Open-World Benchmark for Multi-Label Bacterial Identification in Phase-Contrast Microscopy

PHOEBI tests multi-label identification of known species in unseen mixtures using 40 combinations of six bacteria and approximately 120,000 phase-contrast images, revealing severe failures of per-image trained classifiers and greater stability of geometric prototype decoders on frozen features, although their gains must be interpreted against the all-present baseline and combination-level confidence intervals.

PocketVE: Stable and Property-Guided Structure-Based Drug Design with Variance-Exploding Diffusion

PocketVE combines VE/EDM denoising that preserves the shared protein–ligand coordinate scale, training-time pocket perturbation, and inference-time multi-property CFG, raising 3D validity from 58.6% to 80.6% and reducing strain energy from 457.4 to 127.9 relative to guided TAGMol on CrossDocked2020, without attributing these gains to VE alone or claiming superiority on every docking metric.

Preserving DEG Rankings for Gene Discovery in Histology-Based Spatial Gene Expression Prediction

The paper shifts histology-based spatial expression prediction from matching each gene's spatial profile to preserving which genes should be prioritized for a tissue contrast, using morphology-derived proxy groups, differentiable U statistics, and an across-gene correlation loss to improve DEG ranking and pathway overlap without guaranteeing better conventional spatial PCC.

Browse all 11 Computational Biology papers →


⚛️ Physics & Scientific Computing (5)

Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events

The paper compresses a reference-nudged stochastic coarse ensemble into dynamical-sensitivity context and uses FiLM to modulate existing temporal correctors, improving QG extreme-event area and frequency statistics with limited high-resolution training data without guaranteeing better global distributions or every tail diagnostic than a long-data baseline.

Neural Harmonic Measure Operator

NHMO learns a geometry-only harmonic-measure density, integrates boundary data against its normalized kernel, and uses a learned lift for source contributions and approximation error, outperforming four compared baselines across all five 3D MCB-B Poisson categories while reducing repeated solves on one geometry to a cached matrix-vector product and a lift forward pass.

Phaedra: Learning High-Fidelity Discrete Tokenization for the Physical Sciences

Phaedra separates physical-field latents into multidimensional morphology tokens and high-precision one-dimensional amplitude tokens, then learns to recombine and decode them, reducing in-distribution PDE reconstruction nMAE from FSQ's 2.603 to 1.522 and compressing 65.0 GB of scientific data to 3.44 GB.

PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation

PR-Smoother learns a non-Gaussian initial-state posterior and future-conditioned stepwise corrections in physical space while retaining the prescribed simulator, using an observation-only variational objective to jointly estimate states, physical parameters, and sensor biases, with validation on nonlinear Lorenz–96 and a 16,384-dimensional fluid system.

Soft Geometric Inductive Bias for Object Centric Dynamics

The paper encodes provided object states as Clifford multivectors and learns next-timestep states with geometric-product layers and a block-causal Transformer that do not enforce exact equivariance, improving prediction and autoregressive rollouts in symmetry-breaking settings such as wall collisions, anisotropic confinement, and real driving trajectories.


🌍 Earth Science (1)

ThousandWorlds: A benchmark for climate emulation of potentially habitable exoplanets

ThousandWorlds standardizes 1,689 simulations of tidally locked waterworlds from five global climate models into a low-data parameter-to-mean-climate-field benchmark, measures practical emulator accuracy against same-planet GCM disagreement, and reports a normalized geometric-mean RMSE of 0.462 for GPLFR on the full benchmark.


📡 Signal & Communications (1)

Geometric Inductive Biases for Semi-Supervised Equalization: The Constellation-Aware Transformer

CAT jointly processes the known constellation and received signals from the first layer, with a bidirectional FIR signal branch for inter-symbol interference, reducing pilot requirements through per-block semi-supervised adaptation; the benefit still requires sufficient pilots and receiver-side training, rather than enabling unsupervised or zero-shot decoding across channels.


🛡️ AI Safety (10)

A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion

DSS-GNN expands node representations along graph-frequency and stochastic chaos-order axes and propagates uncertainty through deterministic quadrature; its standalone mode obtains the lowest mean Brier among compared methods on 14 calibration benchmarks, while its hybrid mode outperforms the listed baselines on most node-OOD settings and all 7 GOOD concept-shift settings.

Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework

MABT combines empirical manifold anchoring with bilevel initialization learning under local surrogate uncertainty, increasing error transfer in controlled cross-model ImageNet evaluation, while its geometric interpretation, unknown-model distribution approximation, and convergence claims have explicit limits.

ASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction

ASIRF moves sensitive-attribute definitions from fixed model labels into an inference-time knowledge base and adapts through context classification, definition retrieval, and value extraction; the authors report that at least one architecture exceeds OPF recall in 68 of 80 model–domain combinations, but the multi-agent architecture does not consistently outperform the single agent with the same retrieval tool.

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

GT-HarmBench organizes 1,535 strategic-interaction instances into six two-player, two-action game families, reports average utilitarian accuracy of 0.62 across 15 models, and observes aggregate improvements of +0.13 to +0.18 from institutional descriptions in a separate eight-model experiment, but measures textual advice rather than institutional execution or deployment safety.

How Much Must a Private Mempool Hide? Exact Leakage Thresholds for Sandwich Attacks

For a single-order model of a fee-free constant-product automated market maker, the paper proves that interval-leakage security boundaries are governed primarily by the lower endpoint of order size rather than interval width, while separating all-type guarantees, posterior expected incentives, and residual post-trade risks.

Less Supervision, Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Images

ReGFLoW trains a dual-encoder, VAE-reconstruction-guided regional scoring network with image-level real/fake labels, improving average cross-generator F1 in mixed partial-edit/full-synthesis evaluation at a substantial in-domain localization cost; it does not beat the strongest mask-supervised baseline's OOD average on partial edits alone.

Rank-Constrained Adaptation for Reliable Real-World Performance

MARLA uses error probabilities from a frozen ERM model on a held-out adaptation set with task labels to construct a weighted feature subspace, learns a low-rank logit correction only within that subspace, and improves worst-group accuracy without subgroup labels while separating model-selection conditions with no, partial, and complete subgroup knowledge.

Social Choice Foundations for Simulation-Augmented Generation

This paper formalizes representative routing in simulation-augmented generation as proportional clustering: predict simulation-response viewpoint embeddings, then select a small set of simulators with SEAR; it establishes approximate representation guarantees under explicit reward-factorization and error assumptions and outperforms clustering and random-routing baselines on two simulation pools, without studying final-answer synthesis.

Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

TRACE treats gradient leakage in embodied reinforcement learning as a temporally correlated privacy problem, evaluating observation and action leakage through temporal learning under a strict ordered per-step gradient visibility assumption, reaching approximately 18.8 dB PSNR in the main experiment and showing that lightweight compression cannot replace sequence-level privacy protection.

Watermarking Should Be Treated as a Monitoring Primitive

Through an observer-based threat model and controlled text and image experiments, this paper argues that persistent entity bindings and reliable inference can make watermarking support repeated attribution, motivating governance of both internal attribution access and design-dependent external source identification beyond per-sample robustness.


📂 Others (3)

Building Transformation Layers for Riemannian Neural Networks

The paper reinterprets fully connected layers as signed point-to-hyperplane responses, constructs manifold-valued layers using multiple tangent spaces and a tractable pseudo-distance, and extends them to convolution through product manifolds, validating representational flexibility across ten geometric instantiations while performance and cost remain geometry- and data-dependent.

D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Amplitude and Pixel Spaces

D-GAP uses task-loss gradients with respect to Fourier amplitudes to determine frequency-wise cross-domain mixing strengths, then fuses the result with pixel mixing, improving over the respective best generic methods by an average of 5.3 percentage points on four real-world datasets; however, not using target labels does not mean not accessing target data during training.

Position: Let’s Strengthen Verifiability If We Can’t Enforce Reproducibility

Drawing on a code-availability survey of five leading ML/CV conferences from 2021–2025, this paper proposes integrating experiment-log checks and metric recomputation from prediction files into submission, review, and publication when full reproduction cannot be enforced; it is a proposal for partial consistency verification, not a demonstrated fraud-detection system.


🧠 Mixture of Experts (1)

SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization

SpecDrop shapes modular specialization through a fixed category-to-branch assignment, nonzero cross-category soft weights, and fixed-denominator merging, outperforming label-free matched controls on vision tasks with trusted inference-time category labels, while label-aware dense logit masking is more accurate and routing gains on fuzzy text partitions are near zero.


📄 world_models (1)

StarWM: Self-Supervised Trained Attention Routing for Robust World Models

StarWM uses self-supervised dynamics signals to determine spatial attention, then constrains reconstruction through stop-gradient barriers and a dual-stream decoder so that world models retain predictable entities under randomized video distractions; coherent video backgrounds still confuse default routing, while the reward-guided StarWM-R variant improves cross-condition performance.