Skip to content

Which Layer Causes Distribution Deviation? Entropy-Guided Adaptive Pruning for Diffusion and Flow Models

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/changlin31/EntPruner
Area: Model Compression
Keywords: diffusion model pruning, flow matching models, conditional entropy deviation, adaptive progressive pruning, zero-shot neural architecture search

TL;DR

To tackle distribution drift and mode collapse during downstream transfer and pruning of Transformer-based diffusion and flow models, this paper proposes Conditional Entropy Deviation (CED) to quantify module-wise damage to the output generative distribution, coupling it with NTK condition numbers and ZiCo gradient proxies into an adaptive progressive pruning framework, EntPruner, which delivers up to 2.22x inference speedup on DiT and SiT with near-lossless generation quality.

Background & Motivation

In recent years, large-scale visual generative models—predominantly diffusion models and flow matching models—have demonstrated extraordinary synthesis quality. As architectural backbones have systematically transitioned from legacy U-Nets to Transformer-based topologies (such as DiT, SiT, and PixArt-α), these models leverage superior scalability to generate photorealistic imagery. However, this scalability comes at the expense of massive parameter footprints and intense memory demands, severely hindering practical deployment on edge devices and in low-latency interactive scenarios.

Nevertheless, prevailing pruning strategies (e.g., BK-SDM, Diff-Pruning, LD-Pruner) predominantly treat generative backbones as discriminative classification networks, relying heavily on parameter magnitudes or gradient norms to gauge module importance. This approach fundamentally ignores the central requirement of generative models: preserving both the diversity and condition-fidelity of continuous output distributions. When transferring pretrained foundation models to downstream domains, different layers play vastly disparate roles in maintaining distributional balance. Crucially, aggressive one-shot pruning violently fractures the pretrained generation manifold, inducing irreversible distribution divergence or catastrophic mode collapse, which renders subsequent post-pruning fine-tuning remarkably inefficient.

The authors thoroughly investigate how removing different layers impacts generative distributions and reveal two contrasting degradation patterns: dropping certain blocks leads to positive signed entropy divergence, drifting toward pure stochastic noise, whereas removing other critical blocks causes sharp negative entropy divergence, collapsing the model into oversimplified trivial modes. Core idea: quantify the structural distribution disruption caused by dropping each block using absolute Conditional Entropy Deviation (CED) to establish task-aware redundancy rankings, and dynamically govern pruning schedules via NTK and ZiCo zero-shot proxies during training to achieve fidelity-preserving progressive model compression.

Method

Overall Architecture

The execution pipeline of EntPruner consists of two coordinated stages: Stage 1 performs block-level importance estimation and redundancy ranking based on Conditional Entropy Deviation (CED); Stage 2 carries out zero-shot proxy-guided adaptive progressive pruning across multiple training stages. The system autonomously determines which blocks to drop and what pruning ratio to apply at each training step while directly inheriting fine-tuned weights from previous stages, thereby eliminating the trauma of one-shot structural destruction.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Pretrained Generative Backbone<br/>(DiT / SiT-XL/2)"] --> B["Conditional Entropy Deviation (CED)<br/>Drop blocks individually to evaluate entropy deviation"]
    B --> C["Entropy-Guided Redundancy Ranking<br/>Prioritize low absolute CED blocks for pruning"]
    C --> D["Zero-Shot Adaptive Candidate Search<br/>NTK condition number + ZiCo gradient score + size penalty"]
    D --> E["Multi-Stage Adaptive Progressive Pruning<br/>Inherit weights stage-by-stage to prevent mode collapse"]
    E --> F["Compressed Lightweight Generator<br/>30%~50% parameters pruned, 1.33x~2.22x speedup"]

Key Designs

1. Conditional Entropy Deviation Estimation: Quantifying Generative Output Distribution Deviation

Unlike discriminative networks that focus purely on decision boundaries, generative models must preserve the fidelity and coverage of the learned continuous probability distribution \(p(x)\). Assuming the denoising network output feature distribution follows a Gaussian distribution \(X \sim \mathcal{N}(\mu, \sigma^2)\), its continuous differential entropy is expressed analytically as: $$ \mathcal{H}(X) = -\int p(x) \log p(x) dx = \log(\sigma) + \frac{1}{2}\log(2\pi) + \frac{1}{2} $$ When temporarily omitting the \(i\)-th block, the sign of the entropy difference directly indicates the physical mode of distributional degradation: \(\Delta \mathcal{H} > 0\) denotes uncontrolled variance drifting toward isotropic Gaussian noise, whereas \(\Delta \mathcal{H} < 0\) signifies feature collapse into degenerate low-entropy modes. Regardless of direction, any significant deviation degrades generative quality. Thus, Conditional Entropy Deviation measures the absolute deviation magnitude: $$ \text{CED}i = \left| \mathcal{H}_i}) \right| $$ Blocks with large CED values are structurally essential for preserving generative distribution equilibrium and must be protected, whereas low-CED blocks introduce minimal distributional perturbation upon removal, identifying them as redundant components on the target task that receive top pruning priority.}}(X) - \mathcal{H}_{\text{pruned}}(X \mid \text{Drop}{\textit{block

2. Zero-Shot Adaptive Search via NTK and Gradient Statistics: Validation-Free Dynamic Scheduling

Conventional progressive pruning relies on rigid handcrafted layer-decay schedules or incurs prohibitive validation overhead to evaluate candidate subnetworks at each step. EntPruner reformulates subnetwork selection at pruning stage \(k\) into a lightweight zero-shot NAS optimization problem: $$ \psi_k^* = \arg\min_{\psi_k \in \Lambda_k} \mathcal{K}(\omega(\psi_k), \Omega) $$ To assess candidate trainability and generalization capability without intermediate fine-tuning, the framework couples the flow matching Neural Tangent Kernel (NTK) condition number with the signed ZiCo gradient proxy. In flow matching velocity prediction \(v_\theta(x_t, t)\), the eigenvalues of the empirical NTK matrix \(\hat{\Theta}\) dictate the convergence rate of the slowest learning mode, characterized by the NTK condition number: $$ \mathcal{K}\kappa(\psi) = \frac{\lambda_0}{\lambda_m} $$ A smaller condition number corresponds to a smoother optimization landscape and faster, more stable gradient descent. Concurrently, the modified signed ZiCo proxy evaluates the ratio between the expected absolute gradient and its standard deviation across layers: $$ \mathcal{K} \right) $$ Because higher ZiCo values correlate positively with superior trainability, the authors apply a negative sign to cast it as a minimization objective, steering the search toward architectures featuring both smooth loss surfaces and high gradient signal-to-noise ratios.}}(\psi) = \sum_{l=1}^N \log \left( \sum_{\omega \in \omega_l} \frac{\mathbb{E}[|\nabla_\omega \mathcal{L}|]}{\sigma(\nabla_\omega \mathcal{L})

3. Rank-Based Voting with Parameter Regularization: Robust Multi-Criteria Selection

Because the NTK condition number and ZiCo gradient score differ vastly in numerical scale, direct linear aggregation could allow a single metric to overpower the decision process. EntPruner incorporates a parameter-free rank-based voting mechanism augmented with an explicit model size penalty \(\Omega\): $$ \psi_k^* = \arg\min_{\psi_k \in \Lambda_k} R(\psi_k), \quad \text{s.t.} \quad R(\psi_k) = R(\mathcal{K}\kappa(\psi_k)) + R(\mathcal{K}(\psi_k)) + \gamma R(\Omega) $$ where }\(R(\cdot)\) denotes the rank index within the candidate pool (1st, 2nd, ...), and \(\gamma\) is fixed at 0.5 as an efficiency regularization weight. Candidates scoring lower aggregate ranks represent optimal trade-offs and are selected as \(\psi_k^*\) for subsequent fine-tuning. The selected subnetwork directly inherits parameters optimized during the preceding stage, ensuring smooth re-parameterization and avoiding the disruption inherent to one-shot pruning.

Key Experimental Results

Main Results

The framework was comprehensively evaluated on ImageNet 256×256 alongside three fine-grained downstream transfer benchmarks: CUB-200-2011, Oxford Flowers, and ArtBench-10, utilizing both DiT-XL/2 and SiT-XL/2 backbones across ODE and SDE numerical solvers (evaluated at 50 sampling steps and a classifier-free guidance scale of 4.0).

The table below summarizes generation quality (FID, IS) and computational efficiency comparisons across different pruning techniques on the SiT-XL/2 backbone:

Model & Solver Setting Method Sparsity CUB (FID↓ / IS↑) Flowers (FID↓ / IS↑) ArtBench (FID↓ / IS↑) Params (M) Speedup
SiT w/ ODE Full Fine-tuning 0% 5.32 / 6.02 11.78 / 3.71 8.80 / 7.32 675.12 1.00×
SiT w/ ODE LD-Pruner 35% 5.70 / 6.03 12.02 / 3.75 10.78 / 6.63 435.78 1.82×
SiT w/ ODE EntPruner (Ours) 35% 5.48 / 6.07 11.75 / 3.82 10.03 / 6.88 435.78 1.82×
SiT w/ ODE LD-Pruner 50% 6.86 / 6.16 12.09 / 3.79 12.81 / 6.35 334.67 2.22×
SiT w/ ODE EntPruner (Ours) 50% 6.68 / 6.18 11.86 / 3.82 12.65 / 6.41 334.67 2.22×
SiT w/ SDE Full Fine-tuning 0% 5.17 / 5.87 12.47 / 3.77 13.33 / 6.56 675.12 1.00×
SiT w/ SDE LD-Pruner 35% 5.24 / 6.10 12.32 / 3.74 16.16 / 6.20 435.78 1.49×
SiT w/ SDE EntPruner (Ours) 35% 5.22 / 6.11 12.10 / 3.74 15.25 / 6.31 435.78 1.49×
SiT w/ SDE LD-Pruner 50% 5.98 / 6.19 12.92 / 3.75 18.74 / 5.90 334.67 1.85×
SiT w/ SDE EntPruner (Ours) 50% 5.83 / 6.15 12.77 / 3.77 18.69 / 5.90 334.67 1.85×

On ImageNet 256×256 under direct pretrained compression (SiT at 30% pruning ratio with an ODE solver), EntPruner achieves an FID of 2.69 (compared to 2.15 for full uncompressed SiT-XL/2), significantly surpassing BK-SDM (3.48), Diff-Pruning (2.79), and LD-Pruner (6.81). On DiT-XL/2 pruned by 30%, EntPruner achieves 11.99 FID on Flowers (versus 21.05 for full fine-tuning), marking a 43.04% relative error reduction and outperforming Parameter-Efficient Fine-Tuning (PEFT) baselines including LoRA, DiffFit, and BitFit.

Ablation Study

To isolate the empirical contributions of the CED entropy importance criterion and the adaptive progressive search mechanism, ablation experiments were conducted on Oxford Flowers with the SiT backbone:

Config FID ↓ IS ↑ Params (M) MACs (G) Latency (s) Speedup Note
Full Fine-tuning 11.78 3.71 675.12 228.85 0.20 1.00× Unpruned baseline
w/o CED (pruning high-CED blocks) 12.06 3.81 435.78 147.13 0.09 1.82× Pruning distribution-critical blocks degrades FID by 0.31
w/o Ada. Pruning (one-shot pruning) 11.84 3.80 435.78 147.13 0.09 1.82× Lacks training dynamics adaptation; leads to suboptimal convergence
EntPruner (full model) 11.75 3.82 435.78 147.13 0.09 1.82× Fully recovers fidelity and even surpasses full fine-tuning

Key Findings

  • Layer sensitivity exhibits stark architectural heterogeneity: In SiT-XL/2, intermediate blocks (e.g., Block 6) display strong positive entropy deviation (omission causes drift toward random noise), whereas in DiT-XL/2, deep output blocks (Blocks 27-28) display severe negative entropy deviation (omission triggers mode collapse). This demonstrates that generative pruning cannot rely on discriminative heuristics and must be tailored to generative distributional behavior.
  • Progressive scheduling mitigates mode collapse: Hard one-shot structural pruning at initialization causes irreversible damage to the probability manifold that downstream fine-tuning struggles to mend. Progressive multi-stage pruning guided by NTK and ZiCo provides an essential buffer for continuous re-parameterization.
  • Structured pruning provides beneficial regularization: On specialized downstream benchmarks like Flowers, 30%~35% pruned models actually outperform full fine-tuning (FID 11.75 vs 11.78), indicating that shedding over-parameterized generative layers suppresses overfitting on limited-scale target datasets.

Highlights & Insights

  • CED metric tailored for generative distribution dynamics: Measures continuous differential entropy deviations to simultaneously detect noise drift and mode collapse, departing from conventional weight-norm heuristics.
  • Diffusion pruning reframed as zero-shot NAS: Combines theoretical NTK convergence metrics with empirical ZiCo gradient statistics through rank-based voting, eliminating costly search and validation passes.
  • Universal compatibility across generative paradigms: Operates seamlessly across discrete DDPM-style diffusion and continuous flow matching ODE/SDE solvers, unlocking practical on-device deployment for large Diffusion Transformers.

Limitations & Future Work

  • Gaussian simplification in entropy calculation: Assumes high-dimensional feature representations follow a local univariate Gaussian distribution (\(\mathcal{N}(\mu, \sigma^2)\)), which may overlook higher-order moments on complex multi-modal manifolds.
  • Coarse block-level granularity: Restricts pruning to complete Transformer blocks, leaving finer-grained structural dimensions (such as attention head pruning or FFN channel pruning) unaddressed.
  • Validation on large-scale multimodal and video models: While theoretically generalizable, empirical evaluations focus on class-conditional image generation; validating the approach on text-to-image foundation models (e.g., FLUX, SD3) and video diffusion models remains for future work.
  • vs LD-Pruner: LD-Pruner employs task-agnostic operator metrics focused on discriminative features. EntPruner introduces data-dependent CED to protect generative distribution fidelity, consistently outperforming LD-Pruner at 35% sparsity (e.g., CUB FID 5.48 vs 5.70).
  • vs BK-SDM: BK-SDM relies on handcrafted layer-deletion rules and expensive distillation retraining. EntPruner automates pruning via CED and zero-shot NAS without manual tuning.
  • vs DiffFit / LoRA: PEFT methods only reduce training-time memory without improving inference speed. EntPruner physically reduces parameters and MACs, delivering a 1.82×~2.22× speedup while matching or surpassing fine-tuning fidelity.

Rating

  • Novelty: ⭐⭐⭐⭐ [Pioneering entropy-based distribution deviation perspective for generative pruning combined with NTK-ZiCo zero-shot scheduling]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Spans DiT/SiT across ODE/SDE solvers on ImageNet and three downstream benchmarks with detailed ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, rigorous mathematical formulation of entropy and zero-shot proxies, and structured presentation]
  • Value: ⭐⭐⭐⭐⭐ [Provides a practical, plug-and-play solution for accelerating large Diffusion Transformers on edge devices]