HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Model Compression
Keywords: diffusion model, model quantization, quantization-aware training, Single Hadamard Transformation, activation outliers
TL;DR¶
To resolve the catastrophic quality degradation in low-bit diffusion models caused by activation outliers, HQ-DM applies a Single Hadamard Transformation exclusively to activations, smoothing activation peaks while preventing weight outlier amplification and enabling hardware-friendly integer convolution and linear operations, boosting LDM-4 Inception Score by over 467% under W4A3.
Background & Motivation¶
Diffusion models have become the predominant paradigm for visual synthesis and multimodal content creation owing to their iterative denoising mechanism. However, this sequential, multi-step denoising process incurs massive memory consumption and high inference latency, posing severe obstacles for deployment on resource-constrained edge devices and real-time interactive platforms. Model quantization, which maps high-precision floating-point numbers to low-bit integer representations (such as INT4 or INT3), serves as a cornerstone technique to compress storage footprints and leverage integer arithmetic units for accelerated execution. Nevertheless, the noise prediction networks within diffusion models exhibit profound temporal non-stationarity across timesteps, where feature magnitudes and dynamic ranges fluctuate dramatically across sampling steps. Compounding this challenge, activation tensors feature heavy-tailed distributions with extreme outliers, leading to severe clipping and rounding errors under low-bit linear quantization that rapidly accumulate and amplify across recursive steps, ultimately collapsing generation fidelity.
Existing quantization paradigms fall short in extreme low-bit scenarios. Post-training quantization (PTQ) offers zero training overhead but experiences catastrophic degradation at or below 4-bit precision. Meanwhile, quantization-aware training (QAT) approaches utilizing LoRA distillation and learned step sizes (LSQ) improve training stability, yet their lack of explicit activation outlier suppression causes severe quality drops under aggressive regimes like W4A3. While orthogonal Hadamard transformations have recently emerged in LLM quantization to disperse outliers via spatial rotations, conventional Double Hadamard Transformations mandate simultaneous transformations on both activations and weights. This conventional design fundamentally clashes with convolutional operations featuring spatial sliding windows and strides; more critically, multiplying weight matrices (which typically have well-behaved, smooth distributions) by a Hadamard matrix inadvertently generates artificial weight outliers, worsening weight truncation error and degrading overall accuracy.
The critical insight of this work is to decouple activation smoothing from weight transformation: by rotating only the activation matrix and absorbing the normalization factor into the floating-point scaling scalars, the core linear and convolutional operations can run entirely using integer arithmetic without contaminating the weights. Core idea: propose a Single Hadamard Transformation-Based Quantization-Aware Training framework (HQ-DM) that performs block-diagonal orthogonal rotations exclusively on activations to eliminate outliers, seamlessly supports integer convolution and linear matrix multiplications, and combines LoRA distillation with timestep-adaptive step sizes to achieve high-fidelity low-bit diffusion synthesis.
Method¶
Overall Architecture¶
HQ-DM establishes an efficient quantization-aware distillation pipeline and a hardware-friendly inference flow tailored for low-bit diffusion models. Prior to each linear layer and convolutional layer throughout the denoising network, block-diagonal Hadamard orthogonal matrices are constructed according to the channel or spatial dimensions. The framework applies online orthogonal transformations exclusively to the incoming activation tensors, evenly diffusing extreme outlier peaks across the orthogonal subspace. Subsequently, the transformed activations and the intact weight tensors are fed into symmetric integer quantizers. Through algebraic identities, the Hadamard normalization constants are directly folded into the dequantization scale factors, allowing the underlying core compute to run as integer matrix multiplications and integer convolutions on hardware tensor cores. During the training phase, the original full-precision teacher weights are frozen, while low-rank LoRA adapters and timestep-dependent learnable quantization step sizes (LSQ) are jointly optimized via distillation loss to restore precision with minimal computational overhead.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Activations & Feature Maps<br/>Contains severe outlier channels"] --> SHT["Single Hadamard Activation Transform<br/>Block-diagonal orthogonal rotation diffuses peaks"]
SHT --> IntOp["Hardware-Friendly Integer Operator Formulation<br/>Absorbs normalization factor/supports INT linear & conv"]
IntOp --> QAT["LoRA Distillation & Timestep-Adaptive Optimization<br/>Freeze backbone weights/learn timestep-wise LSQ scales"]
QAT --> Out["Low-Bit Dequantized Outputs<br/>Proceed to subsequent denoising step"]
Key Designs¶
1. Single Hadamard Activation Transform: Smoothing Activation Peaks Without Amplifying Weight Outliers Traditional Double Hadamard Transformation (DHQ) attempts to insert \(H\) and \(H^\top\) symmetrically between activations \(X\) and weights \(W\). However, while neural network weights natively follow concentrated Gaussian-like distributions, applying an offline Hadamard transformation to them induces sharp, heavy-tailed weight outliers (for example, on LDM-4's output block weights, DHQ increases the maximum absolute value by 2484% from 0.037 to 0.964 and surges kurtosis by 65.81%), severely undermining low-bit weight representation. HQ-DM introduces a Single Hadamard Transformation that restricts the orthogonal rotation strictly to the activation side. When the input channel dimension \(C_i\) is not a power of two, HQ-DM constructs a block-diagonal Hadamard matrix composed of \(m\) base Hadamard submatrices \(H_k \in \mathbb{R}^{2^k \times 2^k}\):
When an activation channel carries a dominant outlier approximated by a scaled basis vector \(e_i\), the transformation yields \(e_i^\top H_k = 2^{-k/2} \cdot \mathbf{1}_{2^k}\), effectively spreading the peak into a uniform flat vector across the \(2^k\) coordinates. This substantially narrows the dynamic range of activations while leaving weights completely untouched in their original well-behaved distribution.
2. Hardware-Friendly Integer Operator Formulation: Direct Execution of INT Linear and Convolutional Layers To ensure high acceleration on edge AI accelerators, the Single Hadamard Transformation must execute without requiring high-precision floating-point inverse transformations during runtime. For linear layers \(Y = XW\), utilizing the orthogonality property \(H H^\top = I\), the operation is rewritten as \(Y = (X H) H^\top W\). By absorbing the orthonormal scaling coefficient \(2^{-k/2}\) (from \(H_k = 2^{-k/2} H_k^{\text{raw}}\)) into the floating-point quantization scales, the unnormalized Hadamard matrix \(H_k^{\text{raw}} \in \{-1, +1\}\) becomes a sparse integer operator, enabling the linear transformation to execute via pure integer arithmetic:
For convolutional layers where spatial neighborhood operations and non-unit strides prevent direct matrix multiplication distributivity, HQ-DM introduces a spatial reshaping strategy. Specifically, the 4D activation tensor \(X_c \in \mathbb{R}^{B \times C_{\text{in}} \times h \times w}\) merges batch, channel, and height dimensions into a 2D matrix \(\widetilde{X}_c \in \mathbb{R}^{(B \cdot C_{\text{in}} \cdot h) \times w}\), applying a square block-diagonal Hadamard matrix \(H_c \in \mathbb{R}^{w \times w}\) along the width dimension:
This formulation enforces orthogonal closure along the spatial width dimension, allowing integer convolutions with arbitrary strides and padding to run directly on INT tensor units without requiring intermediate layer-wise dequantization.
3. LoRA Distillation & Timestep-Adaptive Optimization: Parameter-Efficient Error Recovery To compensate for information loss under ultra-low bitwidths alongside outlier suppression, HQ-DM adopts a parameter-efficient distillation fine-tuning paradigm. The pre-trained backbone weights remain frozen, and trainable low-rank update matrices are attached in parallel: \(W' = W + B \cdot A\), where rank \(r \ll \min(C_i, C_o)\). Accounting for the significant activation distribution shifts across denoising steps, the framework allocates dedicated learnable quantization step sizes \(S(t)\) for each individual timestep \(t \in [1, T]\) for both weights and activations, employing the Straight-Through Estimator (STE) for gradient back-propagation:
The student model optimizes a data-free distillation loss against the full-precision teacher model outputs across time steps, converging rapidly with minimal training overhead.
Loss & Training¶
During the distillation phase, noisy latent representations \(x_t\) are concurrently fed into the frozen FP32 teacher \(\epsilon_{\mathrm{FP}}\) and the quantized student \(\epsilon_{\mathrm{quant}}\). The training objective minimizes the mean squared error across timesteps:
The AdamW optimizer is employed with differentiated initial learning rates for activation quantization scales, weight quantization scales, and LoRA parameters. Because the sparse Hadamard transformation adds negligible compute, the entire distillation finishes within 1.4 hours on a single A100 GPU, matching EfficientDM in training speed while surpassing it in convergence stability.
Key Experimental Results¶
Main Results¶
On conditional image synthesis using ImageNet \(256 \times 256\), HQ-DM was thoroughly evaluated on the U-Net-based LDM-4 model across multiple bitwidth configurations using 20 DDIM sampling steps. The results demonstrate that HQ-DM substantially outperforms prior PTQ and QAT baselines, with performance margins expanding under lower bitwidths.
| Method | Bit (W/A) | IS ↑ | FID ↓ | sFID ↓ | Precision (%) ↑ |
|---|---|---|---|---|---|
| FP32 Baseline | 32/32 | 365.35 | 11.22 | 7.78 | 93.68 |
| Q-Diffusion | 8/8 | 350.93 | 10.60 | 9.29 | 92.46 |
| PTQD | 8/8 | 359.78 | 10.05 | 9.01 | 93.00 |
| EfficientDM | 8/8 | 362.34 | 11.38 | 8.04 | 93.77 |
| HQ-DM (Ours) | 8/8 | 363.22 | 11.01 | 7.76 | 93.46 |
| Q-Diffusion | 4/8 | 336.80 | 9.29 | 9.29 | 91.06 |
| PTQD | 4/8 | 344.72 | 8.74 | 7.98 | 91.69 |
| EfficientDM | 4/8 | 351.97 | 9.96 | 7.80 | 92.48 |
| HQ-DM (Ours) | 4/8 | 355.59 | 10.07 | 7.14 | 93.07 |
| PTQD | 8/4 | 1.93 | 262.01 | 111.10 | 0.54 |
| EfficientDM | 8/4 | 252.00 | 6.75 | 8.48 | 84.90 |
| HQ-DM (Ours) | 8/4 | 292.76 | 7.68 | 9.93 | 88.68 |
| PTQD | 4/4 | 2.57 | 258.11 | 124.33 | 0.69 |
| QuEST | 4/4 | 202.45 | 5.98 | — | — |
| EfficientDM | 4/4 | 220.20 | 6.63 | 9.80 | 80.97 |
| HQ-DM (Ours) | 4/4 | 248.30 | 6.58 | 9.60 | 84.44 |
| PTQD | 4/3 | 1.62 | 297.88 | 150.46 | 0.31 |
| EfficientDM | 4/3 | 8.74 | 99.78 | 65.79 | 27.48 |
| HQ-DM (Ours) | 4/3 | 49.62 | 33.47 | 12.09 | 54.87 |
To verify architectural generalizability, HQ-DM was also evaluated on the Transformer-based DiT-XL model (50 steps). On ImageNet \(256 \times 256\) under W4A4, HQ-DM attains an Inception Score of 355.74 (a 60.92% relative increase over EfficientDM's 221.07) and reduces FID from 10.41 to 9.62, highlighting robust scalability across non-convolutional diffusion backbones.
Ablation Study¶
To isolate the contributions of timestep-aware step sizes (LSQ), LoRA fine-tuning, and the Hadamard transformation (Hada.), an ablation study was conducted on LDM-4 under W4A4 on ImageNet \(256 \times 256\).
| Config / Variant | LSQ | LoRA | Hada. | IS ↑ | FID ↓ | sFID ↓ | Precision (%) ↑ |
|---|---|---|---|---|---|---|---|
| FP32 Baseline | — | — | — | 365.35 | 11.22 | 7.78 | 93.68 |
| EfficientDM Baseline | ✓ | ✓ | ✗ | 220.20 | 6.63 | 9.80 | 80.97 |
| Naive PTQ + LSQ | ✓ | ✗ | ✗ | 5.09 | 193.59 | 67.24 | 3.38 |
| Naive PTQ + LSQ + Hada. | ✓ | ✗ | ✓ | 33.29 | 68.16 | 48.86 | 26.50 |
| Naive QAT + LSQ | ✓ | ✗ | ✗ | 159.24 | 7.69 | 8.38 | 76.85 |
| Naive QAT + LSQ + Hada. | ✓ | ✗ | ✓ | 203.69 | 6.47 | 10.98 | 82.53 |
| LoRA only | ✗ | ✓ | ✗ | 185.51 | 8.55 | 10.46 | 73.25 |
| LoRA + Hada. | ✗ | ✓ | ✓ | 229.75 | 6.52 | 10.72 | 79.55 |
| HQ-DM (Full Model) | ✓ | ✓ | ✓ | 248.30 | 6.58 | 9.60 | 84.44 |
Furthermore, when comparing the proposed Single Hadamard scheme against the conventional Double Hadamard Transformation (DHQ-DM), DHQ-DM suffered severe accuracy drops across all configurations (yielding an IS of only 4.37 and an FID of 179.40 under W4A4) while incurring over 10% latency overhead, validating the necessity of avoiding weight rotations.
Key Findings¶
- Independent orthogonal benefit across paradigms: Even in training-free PTQ setups, adding the Single Hadamard Transformation lifts Inception Score from 5.09 to 33.29 and slashes FID by 125.43 points, confirming that geometric outlier dispersal provides fundamental quantization advantages decoupled from distillation tuning.
- Amplified gains at ultra-low bitwidths: While HQ-DM improves IS by 12.76% over EfficientDM under W4A4, its margin surges dramatically in the W4A3 setting—elevating IS from 8.74 to 49.62 (+467.73%) and lowering FID from 99.78 to 33.47—successfully pushing generative diffusion models into practical 3-bit activation regimes.
- Hardware efficiency and low training overhead: Single Hadamard operations introduce virtually zero calibration latency (1.4 hours total on an A100 GPU versus 28 hours for PTQD), while delivering 10-14% faster per-sample inference than Double Hadamard approaches by eliminating redundant matrix multiplications.
Highlights & Insights¶
- From bilateral rotation to unilateral smoothing: Unlike LLM practices that routinely apply symmetric bilateral Hadamard rotations, HQ-DM identifies that weights inherently possess small dynamic ranges; applying Hadamard transforms to weights introduces artificial outliers and breaks convolution distributivity, whereas unilateral activation transformation achieves optimal outlier mitigation.
- Algebraic closure for full-integer acceleration: By absorbing normalization scales into subsequent layer dequantization constants, the transformed activation tensors can be represented as integer matrices multiplied by unnormalized \(\{-1, +1\}\) Hadamard bases, allowing direct implementation on standard integer MAC units.
- Data-free distillation utility: Because distillation is performed directly on Gaussian noise latents and intermediate representations, the entire compression process operates data-free without requiring original private training datasets.
Limitations & Future Work¶
- Reliance on weight stability: Because the framework strictly targets activation outliers, the overall accuracy remains sensitive to weight quantization noise; compressing weights to 2 bits remains challenging without mixed-precision support.
- Handling non-standard spatial resolutions: When spatial feature maps or channel dimensions fall below the minimal Hadamard base order \(2^k\), the transformation reverts to the identity matrix \(I\), leaving initial projection layers unprotected.
- Future directions: Integrating Single Hadamard transformations with dynamic layer-wise bit-allocation and exploring ultra-low-bit quantization for high-resolution video generation architectures (e.g., DiT-based video synthesis models).
Related Work & Insights¶
- vs EfficientDM: EfficientDM introduced LoRA distillation for diffusion QAT but relies on asymmetric quantization and suffers catastrophic degradation under 3-bit activations; HQ-DM adopts symmetric quantization with Single Hadamard outlier suppression, outperforming EfficientDM by 467.73% in IS under W4A3.
- vs QuaRot / Double Hadamard Transformation: QuaRot pioneered bilateral Hadamard rotations for post-training LLM quantization but fails on convolutional operators and induces weight outliers; HQ-DM resolves both bottlenecks via unilateral activation rotation and spatial tensor reshaping.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Breaches the conventional bilateral Hadamard assumption with an elegant unilateral activation rotation tailored for diffusion architectures.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation across ImageNet, LSUN-Bedrooms, and LSUN-Churches on both U-Net and DiT backbones, supported by rigorous ablation and latency benchmarks.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical derivation of orthogonal closures accompanied by clear empirical analyses.
- Value: ⭐⭐⭐⭐⭐ Solves the longstanding outlier bottleneck in W4A4 and W4A3 diffusion quantization, offering an actionable blueprint for high-efficiency edge deployment.