Skip to content

Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Model Compression
Keywords: Diffusion Model, Token Reduction, Output-Similarity-Awareness, Matching Reuse, Frequency Penalty

TL;DR

Addressing the severe recovery error caused by conventional input-similarity token matching in diffusion models, DiTo introduces an output-similarity-aware token reduction framework that leverages temporal consistency across adjacent timesteps to reuse prior-output matching correspondences, paired with PMR-guided interval scheduling and a frequency penalty, yielding 1.6โ€“3.9 dB PSNR improvements and up to 18.6% latency reduction on Flux and SD3.

Background & Motivation

Diffusion Transformers (DiTs) have established a prominent standard for high-fidelity image generation owing to their scalable transformer backbone architectures. Nonetheless, the core self-attention operator scales quadratically (\(O(N^2)\)) with the sequence token length, presenting a severe computational bottleneck during high-resolution text-to-image synthesis. While kernel-level techniques like FlashAttention alleviate low-level memory bandwidth constraints, they do not reduce the underlying algorithmic complexity. Consequently, Token Reduction (TR) approaches that prune or merge redundant tokens have attracted significant interest for accelerating diffusion inference without retraining.

Existing token reduction methods for diffusion modelsโ€”such as ToMeSD, ToFu, SiTo, and ToMAโ€”directly inherit design principles from discriminative Vision Transformers (ViTs), relying exclusively on current-layer input token similarities to determine which tokens to reduce. In discriminative classification, discarding redundant tokens is sufficient. In generative diffusion models, however, tokens correspond strictly to physical spatial locations; reduced tokens must subsequently undergo an explicit recovery stage where their activations are reconstructed by copying the outputs of their matched destination tokens. Hence, the overarching objective in diffusion TR is to minimize the recovery error between the restored output and the unreduced dense output. Because reconstruction copies output activations, the recovery error is fundamentally determined by token similarity in the output space rather than the input space. Empirical profiling reveals that conventional input-driven matching severely deviates from this error minimization objective, causing conspicuous structural distortion and texture blurring.

Directly computing output similarities within the current timestep is impossible because the dense output is unavailable until the full forward pass finishes. To overcome this obstacle, this paper exploits the intrinsic temporal consistency along the diffusion denoising trajectory: while intermediate activations are well-known to remain correlated across adjacent timesteps, the relative token similarities in the output space also exhibit strong temporal preservation. Core idea: propose DiTo, an output-similarity-aware token reduction paradigm that decouples the diffusion trajectory into matching steps and reduction steps, using prior-step output similarities as an accurate proxy for current-step golden correspondences and reusing them across subsequent reduction steps with PMR-guided interval scheduling and frequency-aware penalty regularization.

Method

Overall Architecture

DiTo restructures the token reduction pipeline by decoupling token matching from reduction and recovery, operating via an interleaved schedule of Matching steps and Reduction steps across the diffusion timesteps. During a Matching step, the model executes a full dense computation and utilizes its output feature similarities to establish high-fidelity destination-source token correspondence indices. In subsequent Reduction steps, the model bypasses expensive pair-matching calculations altogether, directly reusing the stored compact index metadata to execute token pruning/merging, accelerated sparse attention, and spatial recovery reconstruction.

To simultaneously maximize computational efficiency and perceptual fidelity, DiTo incorporates two foundational mechanisms: offline Pair Match Ratio (PMR)-guided scheduling to dynamically calibrate the valid reuse interval bounds without per-prompt re-profiling, and frequency-aware token matching that penalizes repeatedly chosen source tokens to prevent localized error accumulation and eliminate blocking artifacts.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input noisy latent x_t"] --> B["Output-Similarity-Aware Token Matching<br/>Compute proxy similarity via prior output"]
    B --> C["Frequency-aware Penalty Correction<br/>Penalize cumulative selection counts"]
    C --> D["PMR-guided Interval Scheduling<br/>Designate current step as Matching or Reduction"]
    D -->|Matching Step: refresh mapping| E["Execute dense computation and store index metadata"]
    D -->|Reduction Step: reuse prior mapping| F["Execute token reduction and sparse attention"]
    F --> G["Token Recovery Stage<br/>Reconstruct tokens by copying matched destination outputs"]
    E --> H["Denoised latent x_{t-1}"]
    G --> H

Key Designs

1. Output-Similarity-Aware Token Reduction: Prior-step outputs as golden matching proxies

Prior TR techniques establish the similarity matrix \(\mathbf{A} \in \mathbb{R}^{|D| \times |S|}\) using input token embeddings at timestep \(t\). When the recovery stage subsequently reconstructs pruned source tokens \(s \in S\) by copying output representations from matched destination tokens \(d^\star(s)\), the recovery error \(\mathcal{L}_{\text{out}} = \|Y - \tilde{Y}\|_F^2\) is directly dictated by the distance between output representations \(Y_s\) and \(Y_{d^\star(s)}\). Given that the current dense output \(Y^{(t)}\) is inaccessible during matching, DiTo exploits adjacent-step temporal consistency, using output features from the previous step as a proxy. The authors quantitatively verify that candidate pairs derived from prior outputs (\(\Delta \ge 1\)) achieve markedly higher alignment with the ground-truth golden output matching than those derived from current inputs (\(\Delta = 0\)), strictly aligning the token matching process with recovery error minimization.

2. PMR-guided Interval Scheduling: Balancing reuse efficiency against temporal drift

Reusing correspondence mappings across consecutive timesteps eliminates repeated matching overhead, but increasing the interval \(\Delta\) inevitably degrades match quality due to cumulative temporal drift. To formalize and monitor this degradation, DiTo introduces the Top-\(k\) Pair Match Ratio (PMR), measuring the fraction of source tokens whose predicted Top-\(k\) destination candidate set contains the single golden destination token under current dense output \(Y\): $\(\mathrm{PMR}_{\text{Top-}k}(t, b, \Delta) = \frac{1}{|S|} \sum_{s \in S} \mathbb{I}\left[ \left| T^{(1)}_{t, b}(s) \cap T^{(k)}_{t-\Delta, b}(s) \right| > 0 \right]\)$ Averaging across all transformer blocks produces \(\overline{\mathrm{PMR}}(t, \Delta)\). During offline calibration under quality threshold \(\tau\), DiTo identifies the maximum permissible reuse interval \(\Delta_t^{\max}\) where \(\overline{\mathrm{PMR}}(t, \Delta) \ge \tau\). At inference time, DiTo adopts a one-step look-ahead rule: it computes the prospective interval \(\Delta_{t+1} = (t+1) - m\) relative to the latest Matching step \(m\). If \(\Delta_{t+1} > \Delta_{t+1}^{\max}\), step \(t\) is designated as a new Matching step to refresh \(m\); otherwise, it operates as a Reduction step. Redundant adjacent matching steps are collapsed to retain only the final one, guaranteeing minimal overhead.

3. Frequency-aware Token Matching: Suppressing localized error accumulation and blocking artifacts

When a single matching correspondence is repeatedly reused across multiple Reduction steps, specific tokens exhibiting slightly higher candidate similarity scores are repeatedly selected for reduction. Consequently, their approximation errors accumulate linearly over the reuse window, forming pronounced, visible blocking artifacts at high-intensity selection coordinates. To resolve this imbalance, DiTo maintains a global token selection history vector \(\mathbf{C} \in \mathbb{R}^N\) tracking how many times token \(i\) has been selected as a reduction target. During matching, DiTo normalizes the candidate score dynamic range \(\Delta s = \max(\hat{\mathbf{s}}) - \min(\hat{\mathbf{s}})\) and reduction ratio \(r\) to apply a scale-invariant penalty: $\(\mathbf{p} = \lambda \cdot r \cdot \Delta s \cdot \mathbf{C}[\mathcal{S}], \qquad \hat{\mathbf{s}}^{\mathrm{pen}} = \hat{\mathbf{s}} - \mathbf{p}\)$ where \(\lambda\) denotes the penalty strength hyperparameter. By penalizing frequently reduced spatial tokens, DiTo forces the model to rotate reduction targets dynamically over time, dispersing reconstruction residuals evenly and eliminating localized blocky artifacts while maintaining overall inference speedups.

Loss & Training

DiTo is completely training-free and operates as an orthogonal, plug-and-play acceleration framework without model parameter fine-tuning or backpropagation. The scheduling policy is executed using pre-calibrated interval lookup tables \(\Delta_t^{\max}\). In terms of memory footprint, DiTo stores only compact 1D integer index metadata during Matching steps rather than high-dimensional activation tensors, resulting in negligible runtime memory overhead compared to feature-caching baselines.

Key Experimental Results

Main Results

Evaluations are conducted on 1024ร—1024 text-to-image synthesis using FLUX.1-dev (35 steps) and Stable Diffusion 3 (SD3, 50 steps) across ImageNet-1k prompts (3,000 images over 3 seeds), measured against vanilla dense outputs on an NVIDIA RTX 6000 Ada GPU.

Model Ratio Method FIDโ†“ PSNRโ†‘ SSIMโ†‘ LPIPSโ†“ CLIPโ†‘ Latency (s)โ†“ Speedup (ฮ”)โ†“
Flux Baseline Vanilla Dense 31.62 โ€“ โ€“ โ€“ 23.71 26.64 0%
Flux 0.25 DiTo (Ours) 31.95 25.21 0.8552 0.2385 23.85 24.25 -9.0%
Flux 0.25 ToMeSD 33.76 21.59 0.7817 0.3395 23.93 24.53 -7.9%
Flux 0.25 ToFu 33.75 21.70 0.7851 0.3354 23.94 24.47 -8.1%
Flux 0.25 SiTo 32.30 22.39 0.7965 0.3462 23.93 25.25 -5.2%
Flux 0.25 ToMA 32.03 22.97 0.8112 0.2716 23.78 24.28 -8.9%
Flux 0.50 DiTo (Ours) 33.47 22.04 0.7757 0.3329 24.01 21.69 -18.6%
Flux 0.50 ToMeSD 90.00 19.73 0.6172 0.5384 23.04 22.00 -17.4%
Flux 0.50 ToFu 83.79 19.81 0.6328 0.5302 23.05 22.02 -17.3%
Flux 0.50 SiTo 61.05 20.17 0.6503 0.5245 23.45 22.54 -15.4%
Flux 0.50 ToMA 43.71 20.49 0.6936 0.4386 23.90 21.62 -18.8%
SD3 Baseline Vanilla Dense 27.67 โ€“ โ€“ โ€“ 24.67 9.32 0%
SD3 0.25 DiTo (Ours) 27.83 26.83 0.9160 0.1492 24.77 8.63 -7.4%
SD3 0.25 ToMeSD 28.42 23.32 0.8697 0.2045 24.74 8.82 -5.3%
SD3 0.25 ToFu 27.99 23.83 0.8835 0.1922 24.72 8.87 -4.9%
SD3 0.25 SiTo 27.92 23.01 0.8682 0.2211 24.79 9.30 -0.2%
SD3 0.25 ToMA 28.47 22.93 0.8624 0.2029 24.63 9.08 -2.5%
SD3 0.50 DiTo (Ours) 28.89 23.16 0.8510 0.2577 24.87 7.79 -16.4%
SD3 0.50 ToMeSD 31.45 21.22 0.8090 0.3139 24.80 8.00 -14.2%
SD3 0.50 ToFu 31.89 21.21 0.8112 0.3187 24.79 8.05 -13.6%
SD3 0.50 SiTo 33.13 21.23 0.8046 0.3170 24.82 8.42 -9.7%
SD3 0.50 ToMA 33.37 20.91 0.8049 0.3053 24.62 8.17 -12.3%

Ablation Study

Ablations dissect recovery error distributions and individual architectural components:

Configuration / Variant Observed Metric & Empirical Values Mechanistic Analysis & Conclusion
Output-based vs Input-based Matching Recovery error \(\mathcal{L}_{\text{out}}\): 100% of 500 samples fall strictly below the \(y=x\) line Confirms that output-space token similarity directly dictates reconstruction error, exposing input similarity as sub-optimal
DiTo w/o Frequency Penalty (\(\lambda=0\)) Peak token selection count exceeds 100, generating severe blocking artifacts Persistent target reuse concentrates approximation errors spatially, degrading local visual coherence
DiTo w/ Frequency Penalty (\(\lambda>0\)) Peak token selection count drops below 40, completely suppressing blocking artifacts Scale-invariant penalty dynamically distributes reduction across spatial tokens, restoring smooth textures
High Reduction Robustness (Flux 50%) Achieves 33.47 FID (vs ToMeSD 90.00, ToFu 83.79, SiTo 61.05, ToMA 43.71) Conventional input matching collapses structurally under aggressive reduction, whereas output matching retains global composition

Key Findings

  • Output similarity is the true governing factor of recovery error: 100% of tested samples achieve lower recovery errors with output-derived matching than with input-based matching, providing the foundational explanation for DiTo's superior Pareto trade-off.
  • Unrivaled resilience under aggressive reduction: At a 50% token reduction ratio, prior methods suffer catastrophic FID deterioration (e.g., ToMeSD degrading to 90.00), whereas DiTo retains a stable FID of 33.47 and a PSNR of 22.04 dB.
  • Negligible metadata storage overhead: In contrast to activation-caching paradigms (e.g., ToCa) that require multi-gigabyte feature buffers, DiTo retains only integer index arrays across steps, avoiding GPU memory expansion.

Highlights & Insights

  • Generative task alignment through output-centricity: Highlights that unlike discriminative ViTs, diffusion models require a recovery stage, shifting the optimal reduction criterion from input redundancy to output reconstruction fidelity.
  • Temporal proxy with scale-invariant frequency control: Elegantly leverages adjacent-step outputs as a non-invasive surrogate for inaccessible current outputs, combined with a dynamic range-normalized penalty that eliminates blocking artifacts with zero added FLOPs.
  • Cross-architecture generalizability: DiTo seamlessly scales across diverse architectures, extending effectively from MM-DiT backbones (Flux, SD3) to standard DiTs and classic U-Nets (SD1.5).

Limitations & Future Work

  • Offline calibration requirement: Although PMR distributions generalize across text prompts and seeds, altering the base sampling steps \(T\) or switching between disparate ODE/SDE solvers requires a one-time offline PMR profiling step.
  • Extreme transient dynamics: In diffusion regimes characterized by abrupt structural shifts between adjacent steps, prior-step output proxies may experience transient alignment drops, presenting an opportunity for online adaptive scheduling in future research.
  • vs ToMeSD / ToFu: ToMeSD and ToFu rely on input token similarity for bipartite soft matching or hybrid pruning; neither minimizes recovery error, causing structural distortion under aggressive pruning. DiTo maintains structural integrity and delivers 56+ FID point advantages at 50% reduction.
  • vs ToMA: ToMA optimizes matching latency via localized parallel attention matching on inputs; DiTo achieves comparable or superior speedups via multi-step index reuse while securing 1.5โ€“2.2 dB higher PSNR.
  • vs ToCa: ToCa achieves acceleration via caching high-dimensional layer activations at substantial memory cost; DiTo reuses lightweight index metadata, achieving similar latency reductions without GPU memory strain.

Rating

  • Novelty: โญโญโญโญโ˜† Pinpoints the fundamental misalignment of input-similarity TR in generative diffusion models and formulates an elegant output-aware proxy paradigm.
  • Experimental Thoroughness: โญโญโญโญโญ Rigorously tested on Flux and SD3 across multiple metrics, providing clear Pareto curves and comprehensive baselines.
  • Writing Quality: โญโญโญโญโญ Exceptionally articulate narrative with well-structured formulations, clear figures, and precise mathematical definitions.
  • Value: โญโญโญโญโญ Retraining-free, low-memory, and highly performant, providing a practical blueprint for edge-device deployment of large-scale diffusion models.