Skip to content

๐ŸŽจ Image Generation

๐Ÿง  NeurIPS2026 ยท 6 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (328) ยท ๐Ÿ“ท CVPR2026 (490) ยท ๐Ÿ”ฌ ICLR2026 (352) ยท ๐Ÿ’ฌ ACL2026 (5) ยท ๐Ÿงช ICML2026 (141) ยท ๐Ÿค– AAAI2026 (79)

๐Ÿ”ฅ Top topics: Diffusion Models ร—4

All Roads Lead to Rome: Flow-driven Multi-Anchor Exploration for Open-Environment Active 3D Mapping

The method generates multiple long-horizon exploration anchors with conditional flow matching, then combines obstacle-aware planning, exploration-mode clustering, and hierarchical path selection to improve mapping coverage in same-difficulty AiMDoom evaluation, difficulty transfer, and AiMDoom-to-MP3D transfer, without guaranteeing generalization to arbitrary open environments.

One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation

OVIE uses a frozen depth model to turn 30M single images into partially valid pseudo-view training pairs, learns observed regions through masked reconstruction and perceptual losses, and completes unknown regions through a full-image adversarial prior, ultimately generating a novel view from only a source image and target pose in one forward pass at 116 FPS on an H100, with multi-view fine-tuning improving in-domain performance.

Panoptic Scene Program Diffusion Transformer

PSP-DiT turns a panoptic scene program containing instances, attributes, relations, and counts into a latent variable jointly denoised with the image, uses ownership and cycle-consistency supervision to constrain its visual realization, and improves GenEval 2 from 32.8 to 38.2 over a matched internal baseline while reporting approximately 1.12 times the inference latency.

PixelDiT2: Representation-Grounded Pixel Diffusion Transformers

PixelDiT2 re-encodes the current noisy image at each pixel-denoising evaluation, using frozen DINO patch representations and a timestep adapter as spatial conditioning without an autoencoder, achieving FID 1.46 on ImageNet-256 after 600 epochs and 1.48 on ImageNet-512 after 680 epochs.

Rethinking Cross-Layer Information Routing in Diffusion Transformers

DAR replaces incremental residual addition in diffusion Transformers with timestep-aware attention over historical sublayer outputs, achieving an unguided ODE FID of 7.56 after 600K steps on ImageNet 256ร—256, a 2.11 improvement over the SiT baseline trained for 1.75M steps, while remaining compatible with REPA's representation-alignment loss.

Revisiting Diffusion Fine-Tuning for Unsupervised Domain Adaptation

MUSE separates diffusion-generator adaptation into a shared semantic branch trained only with source-label supervision and a target-conditioned style branch, allowing one source-involved fine-tuning process to generate bridge data for multiple targets; it improves average downstream UDA accuracy on three closed-set classification benchmarks while reducing generator fine-tuning time to roughly one-half to two-fifths of Terra's cost.