Improved Immiscible Diffusion: Accelerating Diffusion Training by Reducing Miscibility¶
Conference: ECCV 2026
Paper: ECCV Official Link
Code: https://github.com/yilun-li/improved-immiscible-diffusion
Area: Image Generation
Keywords: Diffusion Models, Immiscible Diffusion, Training Acceleration, Optimal Transport, KNN Noise Sampling
TL;DR¶
Reveals the theoretical bottleneck of crossed diffusion trajectories causing denoising supervision collapse in noisy layers, proves that vanilla denoising holds an intrinsic local bijection between generated images and noise origins, and introduces \(O(n)\) KNN noise selection and pairing-free image scaling to accelerate diverse diffusion models by 2.5ร to 4.5ร without sacrificing diversity.
Background & Motivation¶
Diffusion-based models have become the cornerstone of high-fidelity generative modeling and multimodal reasoning, yet their exorbitant training computational cost severely limits deployment and iterative research. For instance, Stable Diffusion V1.1 required over 24 days of training across 256 GPUs, and V2's training time expanded to 32 days on identical hardware. Recent advances such as Flow Matching and minibatch Optimal Transport (OT) seek to straighten or shorten diffusion trajectories to lower denoising variance. However, empirical audits reveal that approximate OT only reduces average image-noise distance by ~2% and the standard deviation of denoising by ~4%, leaving the fundamental acceleration mechanism incompletely understood. Immiscible Diffusion previously introduced linear assignment between images and noise points to reduce miscibility in noise space, achieving noticeable speedups; yet its reliance on Hungarian-style assignment incurs an \(O(n^3)\) computational complexity that fails to scale with large batches. More critically, restricting images from freely diffusing across the entire noise space provoked pervasive skepticism regarding potential degradation in generation diversity and conditional alignment.
The underlying tension arises from an unexamined assumption in conventional diffusion theory versus its empirical generative dynamics. The orthodox perspective assumes that each training image must diffuse isotropically across the entirety of Gaussian space to preserve open-ended diversity. However, this isotropic mixing produces severe high-dimensional "trajectory miscibility": trajectories originating from drastically different images inevitably collide at identical noisy states at higher noise levels (\(t \to T\)). Under such overlap, the single-valued denoising function \(\epsilon_\theta(x_t, t)\) is forced to simultaneously predict conflicting noise targets. As a result, its training objective mathematically collapses to predicting the average of the whole training dataset, eliminating any meaningful sample-specific guidance; in the reverse process, these tangled paths trigger substantial trajectory confusion and oscillation.
This paper cuts into the problem by investigating whether vanilla diffusion actually preserves an isotropic map or secretly maintains strong local correlations. Through systematic perturbation experiments on reverse trajectories, the authors demonstrate that reverse denoising naturally forms a robust, local bijective correspondence: even under 20% independent Gaussian perturbation, generated images preserve identical semantic categories and object outlines. This critical empirical finding dispels the fear that reducing miscibility hurts generative diversity. Core idea: generalize immiscible diffusion from restrictive batch-wise linear assignments into an implementation-agnostic framework of trajectory miscibility reduction, introducing linear-complexity \(O(n)\) KNN noise selection and pairing-free image scaling to eliminate trajectory confusion and deliver up to 4.5x training speedups across diverse architectures and tasks.
Method¶
Overall Architecture¶
In vanilla (miscible) diffusion models, intermediate noisy latents are formed via linear combinations of clean data \(x_0\), noise schedules \(\omega_t\), and sampled Gaussian noise \(\epsilon \sim \mathcal{N}(0, I)\): \(x_t = \sqrt{\omega_t} x_0 + \sqrt{1 - \omega_t} \epsilon\). At maximum noise timesteps \(t \to T\), \(\omega_t \to 0\) and \(x_t \to \epsilon\). When two distinct images \(x_{0,1}\) and \(x_{0,2}\) intersect at the same latent state \(x_t\), the denoising network can only output the arithmetic average \(\frac{1}{2}(\epsilon_1 + \epsilon_2)\). As dataset size \(N \to \infty\), the supervision signal inevitably collapses toward \(\frac{1}{\sqrt{1-\omega_t}}(x_t - \sqrt{\omega_t} \bar{x}_0)\), an uninformative constant independent of specific sample semantics.
Improved Immiscible Diffusion resolves this bottleneck by enforcing geometric separation between the diffusion paths of different images. The pipeline first confirms the stable bijection between noise points and generation outcomes, then breaks trajectory entanglement using either efficient nearest-neighbor noise selection or coordinate space scaling, and finally feeds clear, well-separated trajectories into standard diffusion training objectives.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Data Batch x_0<br/>and Gaussian Noise Prior"] --> B["Noise Space Bijection Analysis<br/>Validate Stable Local Attractor"]
B --> C["KNN Noise Selection<br/>O(n) Nearest-Neighbor Assignment"]
B --> D["Image Geometric Scaling<br/>Enlarge Inter-Sample Distance"]
C --> E["Immiscible Trajectory Formation<br/>Eliminate High-Noise Ambiguity"]
D --> E
E --> F["Efficient Denoising Training<br/>Activate Noisy Layer Supervision"]
Key Designs¶
1. Noise Space Bijection Analysis: Resolving Diversity Concerns and Measuring Miscibility
To rigorously test whether restricting noise assignments compromises generation diversity, this design assesses the sensitivity boundary of the reverse process under controlled perturbations on trained vanilla DDIM checkpoints. Sampling a reference noise point \(N_{\text{orig}}\) and adding 10 independent Gaussian perturbation vectors \(N_{\text{pert}}\), the perturbed initial noise is formulated as:
Empirical inspection demonstrates that at perturbation weight \(W = 10\%\), all 10 generated images retain identical cat contours and foreground structures; at \(W = 20\%\), only minor background variations appear with zero category shift; even at \(W = 30\%\), structural layouts remain remarkably consistent despite occasional class changes. This proves that denoising dynamics naturally construct stable local basins of attraction, confirming that reducing trajectory overlap does not restrict the accessible generative manifold. Furthermore, the paper measures miscibility by computing the average \(L_2\) distance between noise clusters assigned to each image: vanilla DDIM exhibits an average distance of merely \(0.92 \pm 0.06\) (heavily miscible), whereas immiscible diffusion expands this distance to \(4.11 \pm 0.37\) (highly immiscible), demonstrating clear geometric separation between sample diffusion corridors.
2. KNN Noise Selection: Efficient \(O(n)\) Immiscibility Implementation
While batch-wise linear assignment maintains an exact global Gaussian distribution, its \(O(n^3)\) Hungarian complexity creates severe latency bottlenecks when scaling batch sizes to 512 or 1024. This design replaces global bipartite matching with an \(O(n)\) KNN noise selection operator. For each image \(x\) in a batch, the system independently draws \(k\) Gaussian noise candidates \(\{n_1, n_2, \dots, n_k\}\) and selects the candidate with minimum Euclidean distance:
To ensure that selectively dropping unchosen noise candidates does not corrupt Gaussian properties, the authors collected 50,000 noise samples selected via KNN (\(k=8\)) on CIFAR-10 dimensions. The resulting KL divergence against theoretical Gaussian distribution is 48.60, matching standard Gaussian random sampling (48.25) almost perfectly. On an A5000 GPU at batch size 256, KNN executes in just 0.2ms compared to 6.7ms for linear assignmentโa 33.5x speedupโwhile preserving continuous noise trajectory distributions across diffusion steps.
3. Image Geometric Scaling: Pairing-Free Trajectory Dispersal
Grounded in the insight that immiscibility requires reducing the spatial overlap of intermediate diffusion corridors rather than explicitly enforcing sample-noise pairings, this design introduces Image Scaling as a completely training-free geometric mechanism. The noise distribution and variance schedule remain standard, while all pixel values of the input images are multiplied by a constant factor greater than 1 (scaling pixel standard deviation from the default 0.5 to 1.0 or 2.0). Mathematically, this linear expansion magnifies the pairwise Euclidean distances between clean data points in image space. Because noise variance at intermediate timesteps \(t < t_{\max}\) remains unchanged, the enlarged distances between trajectory origins drastically reduce the geometric volume of intersection between diffusion paths. Feature-level diagnostics confirm that scaling reduces t-SNE trajectory tangling, suppresses denoising confusion, and accelerates convergence without incurring any distance computation overhead.
Loss & Training¶
Immiscible Diffusion integrates seamlessly into standard diffusion objectives across architectures including DDIM, Flow Matching, and Consistency Models without modifying backbone architectures or adding auxiliary regularizers:
The optimal choice of candidate pool size \(k\) scales with data dimensionality: for CIFAR-10 (3072 dimensions), \(k \in [4, 8]\) achieves the optimal trade-off; for ImageNet latent representations (4096 dimensions), \(k = 64\) provides sufficient geometric separation. Furthermore, because Immiscible Diffusion specifically reshapes the geometric training trajectories, it operates completely orthogonal to complementary training acceleration techniques like Perception Prioritized Training (P2 weighting) and Curriculum Learning, enabling additive compound efficiency gains.
Key Experimental Results¶
Main Results¶
The authors evaluate Immiscible Diffusion across unconditional generation (CIFAR-10), conditional synthesis and fine-tuning (ImageNet-1k, MS-COCO), and extended downstream applications including image in-painting, out-painting, and visuomotor robotics policy learning.
| Architecture / Task | Dataset / Setting | Method Configuration | Steps Speedup to Best Baseline FID | Converged Metric (FID / Coverage) | Baseline Metric (FID / Coverage) | Metric Gain (Delta) |
|---|---|---|---|---|---|---|
| DDIM | CIFAR-10 Unconditional | Immiscible (KNN, k=8) | 3.5ร faster | 4.45 (FID) | 6.63 (FID) | -2.18 FID (Better) |
| DDIM | CIFAR-10 Unconditional | Immiscible (Assignment) | 2.5ร faster | 5.14 (FID) | 6.63 (FID) | -1.49 FID (Better) |
| Flow Matching | CIFAR-10 Unconditional | Immiscible (KNN, k=4) | 4.5ร faster | 3.56 (FID) | 4.92 (FID) | -1.36 FID (Better) |
| Flow Matching | CIFAR-10 Unconditional | Immiscible (Assignment) | 3.0ร faster | 5.02 (FID) | 4.92 (FID) | +0.10 FID (Parity) |
| Consistency Models | CIFAR-10 Unconditional | Immiscible (KNN, k=4) | 2.8ร faster | 9.16 (FID) | 10.22 (FID) | -1.06 FID (Better) |
| Consistency Models | CIFAR-10 Unconditional | Immiscible (Assignment) | 3.2ร faster | 8.54 (FID) | 10.22 (FID) | -1.68 FID (Better) |
| Stable Diffusion | ImageNet-1k (In-painting) | Immiscible SD (KNN) | - | 17.32 (FID) | 18.35 (FID) | -1.03 FID (Better) |
| Stable Diffusion | ImageNet-1k (Out-painting) | Immiscible SD (KNN) | - | 27.57 (FID) | 29.34 (FID) | -1.77 FID (Better) |
| Diffusion Policy | PushT Robotic Planning | Immiscible (KNN, k=2) | - | 82.83% (Last 10 ckpt avg) | 79.56% (Last 10 ckpt avg) | +3.27% Coverage |
Ablation Study¶
1. Execution Latency and Complexity Comparison (Single A5000 GPU)
| Batch Size | Linear Assignment Latency (ms) | KNN Noise Selection Latency (ms) | Speedup Ratio (\(t_{\text{assign}} / t_{\text{knn}}\)) |
|---|---|---|---|
| 128 | 5.4 | 0.2 | 27.0ร |
| 256 | 6.7 | 0.2 | 33.5ร |
| 512 | 8.8 | 0.3 | 29.3ร |
| 1024 | 22.8 | 0.7 | 32.6ร |
2. Candidate Count \(k\) Sensitivity Ablation (DDIM on CIFAR-10)
| Configuration (\(k\) Value) | Noise-Image \(L_2\) Distance Reduction | Best Converged FID | Mechanism Observation |
|---|---|---|---|
| Baseline (\(k=1\)) | 0.00% | 6.63 | Severe trajectory miscibility in noisy layers |
| \(k=2\) | -0.62% | 4.94 | Initial trajectory uncrossing yields sharp FID gain |
| \(k=4\) | -1.10% | 4.77 | Intermediate trajectory overlap suppressed |
| \(k=8\) (Optimal) | -1.58% | 4.45 | Optimal balance between miscibility reduction and Gaussian prior |
| \(k=16\) | -1.95% | 4.85 | Minor drift from standard isotropic Gaussian prior |
| \(k=32\) | -2.31% | 5.19 | Tail noise contraction begins to degrade generation diversity |
Key Findings¶
- Unlocking High-Noise Layer Supervision: Layer-wise t-SNE and step-wise FID analyses show that while vanilla DDIM at step \(\tau = S\) (pure noise) exhibits chaotic feature collapses, Immiscible Diffusion provides distinct denoising targets from the very first step, accelerating early-stage reconstruction FID substantially.
- KNN Outperforms Linear Assignment: Beyond executing orders of magnitude faster (0.2ms vs 6.7ms at batch size 256), KNN forms continuous, smooth local noise neighborhoods that enhance performance in near-clean timesteps, achieving superior final FIDs compared to hard bipartite assignments.
- Orthogonal Synergy with Existing Acceleration Techniques: Under Flow Matching, combining Immiscible Diffusion with Curriculum Learning drops the 100k-step FID from 6.43 to 4.75; adding it to Perception Prioritized Training reduces FID from 5.58 to 4.26, validating complete orthogonality.
Highlights & Insights¶
- Debunking the High-Noise Inefficiency Dogma: The community previously assumed that high-noise stages are inherently intractable due to vanishing SNR; this work reveals that the true culprit is multi-path trajectory collisions causing target mean-collapse.
- Reconciling Diversity with Optimization Efficiency: Demonstrating that vanilla diffusion inherently maintains a local bijection dispels the long-standing belief that trajectory uncrossing inherently damages sample diversity.
- Practically Zero-Cost Drop-In Acceleration: KNN noise selection and image scaling introduce negligible runtime overhead and zero architectural modifications, allowing seamless integration into large-scale diffusion pipelines.
Limitations & Future Work¶
- High-Dimensional \(k\) Scaling: In extreme latent dimensions (e.g., billion-parameter latent video diffusion), the candidate size \(k\) required to maintain adequate geometric separation may increase, motivating hierarchical or partitioned nearest-neighbor schemes.
- Theoretical Bounds: The framework relies primarily on empirical feature diagnostics and geometric metrics; deriving rigorous statistical-physics or Wasserstein error convergence bounds remains an open challenge.
- Scale-Induced SNR Shift: Extreme pixel scaling alters the effective signal-to-noise ratio at boundary timesteps, which may require careful normalization tuning in dynamic-range-sensitive modalities.
Related Work & Insights¶
- vs Immiscible Diffusion (Li et al., NeurIPS 2024): The predecessor constrained immiscibility to \(O(n^3)\) Hungarian assignment on image-noise pairs and left diversity concerns unaddressed; this work redefines immiscibility as a general trajectory phenomenon, proposes \(O(n)\) KNN and pairing-free scaling, and validates transferability to downstream editing and robotics.
- vs Minibatch Optimal Transport Flow Matching (Pooladian et al., ICML 2023 / Tong et al., TMLR 2024): While OT literature attributes speedups to trajectory straightening, this paper reveals that approximate OT's modest 2% distance reduction accelerates learning primarily by functioning as an implicit realization of trajectory de-miscibility.
Rating¶
- Novelty: โญโญโญโญโญ Decouples diffusion trajectory miscibility as the root cause of noisy-layer degradation and provides an elegant geometric rethinking of optimal transport.
- Experimental Thoroughness: โญโญโญโญโญ Rigorously tested across CIFAR-10, ImageNet, MS-COCO, spanning DDIM, Flow Matching, Consistency Models, editing, and robotics.
- Writing Quality: โญโญโญโญโญ Cohesive narrative with compelling perturbation diagnostics and clear step-by-step layer analyses.
- Value: โญโญโญโญโญ Highly practical, nearly zero-cost plug-and-play acceleration for pretraining and fine-tuning.