RotateAttention : RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation¶
Conference: ECCV 2026
Paper: ECCV Paper
Area: Video Generation
Keywords: Video Generation, Quantized Attention, 3D RoPE, Outlier Equalization, Range Rectification
TL;DR¶
Targeting the quadratic attention bottleneck in DiT-based video generation models equipped with 3D RoPE, RotateAttention introduces RoPE-aware orthogonal rotations to equalize Query-Key outliers with zero or negligible overhead, combines range-rectified affine quantization for probability matrix \(P\), and leverages selective FP16 fallback to achieve up to a 1.68× end-to-end speedup with virtually lossless visual quality.
Background & Motivation¶
Diffusion Transformer (DiT) models have achieved striking success in generative video synthesis. However, because video generation entails long token sequences often exceeding 10K, the self-attention mechanism scales quadratically with sequence length, standing as the primary latency and memory bottleneck. While 4-bit (INT4) FlashAttention quantization offers a compelling avenue for hardware acceleration, directly porting rotation-based quantization techniques established in large language models (LLMs) to 3D Rotary Position Embedding (3D RoPE) video DiTs encounters two prohibitive structural hurdles: first, online rotation transformations (such as Hadamard transforms) commonly utilized to suppress outlier activations in Query (Q) and Key (K) cannot be easily hidden or fused with sparse 3D RoPE kernels in compute-bound video DiT workloads; second, under causal or bidirectional attention, the non-negative attention probability matrix \(P = \exp(S - m) \in (0, 1]\) wastes the entire negative half \([-8, 0)\) of the INT4 dynamic range when symmetric quantization is applied, squandering half of the available quantization levels and resulting in severe truncation error.
The core tension underlying these issues is that low-bit numerical fidelity strictly requires mitigating extreme channel-wise outliers, yet the geometric partitioning of 3D RoPE creates strong, anisotropic outlier distributions across temporal, height, and width frequency sub-bands. Furthermore, while LLM quantization typically preserves FP16 precision for Queries and only quantizes Keys in the memory-bound KV cache, video DiTs are compute-bound and require simultaneous low-bit quantization of both Q and K to harness Tensor Core matrix multiplication speedups. Any asymmetric or non-orthogonal transformation intended to suppress outliers in Q unavoidably scales outliers up in K by the reciprocal singular values, rapidly destabilizing quantization precision.
The paper attacks this challenge by systematically characterizing the channel-wise incoherence profiles of Q and K after 3D RoPE, discovering that outliers are strictly partitioned along spatiotemporal segments, display cross-matrix symmetry between Q and K, and concentrate within one half-segment of each coordinate dimension. Core idea: design RoPE-aware block-diagonal orthogonal rotations (including mergeable Interleaved Rotation and negligible-overhead Half Rotation) to equalize Q/K outliers without breaking RoPE compatibility, paired with an affine range-rectified quantization for probability matrix \(P\) and selective FP16 fallback on sensitive steps/blocks to build a robust, high-performance INT4 FlashAttention pipeline.
Method¶
Overall Architecture¶
RotateAttention is tailored for DiT-based video generation architectures employing 3D RoPE. In the forward attention computation, input features are projected into Q, K, and V matrices, followed by channel smoothing and an offline-fused Hadamard transform applied to the Value matrix. Next, Query and Key tensors undergo RoPE-aware orthogonal rotation—either pre-fused with RoPE weights as Interleaved Rotation or executed via an ultra-sparse element-wise Half Rotation kernel—to equalize outlier magnitudes across channels with zero or negligible runtime overhead. Inside the tiled FlashAttention kernel, pre-softmax INT4 GEMMs compute the attention logits \(S\), after which the non-negative local attention matrix \(P\) is quantized via an affine range-rectified mapping that maps values onto the full \([-8, 7]\) INT4 range before performing the second INT4 GEMM with \(V\). Finally, a selective mixed-precision policy falls back to FP16 in accuracy-critical initial and final denoising steps as well as boundary DiT blocks, executing the remaining majority of attention operations fully in INT4.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Features & Linear Projections<br/>Q, K, V Tensors"] --> B["Offline Value Hadamard Transform<br/>Fused into Wv and Wo Weights"]
B --> C["RoPE-Aware Rotation Transform<br/>Interleaved Fused / Half-Segment Rotation"]
C --> D["INT4 Q-K Block Matmul<br/>Compute Attention Logits S = QK^T / sqrt(D)"]
D --> E["Range-Rectified P Quantization<br/>Affine Mapping with Scale 15 & Zero-Point -8"]
E --> F["INT4 P-V Block Matmul<br/>Accumulate Attention Outputs"]
F --> G["Selective FP16 Fallback Policy<br/>Preserve High Precision on Sensitive Steps/Blocks"]
Key Designs¶
1. RoPE-Aware Rotation: Orthogonal Outlier Equalization and Operator Fusion for Q and K
In video DiTs using 3D RoPE, the hidden dimension \(D\) is split into temporal, height, and width segments: \(D = D_f + D_h + D_w\). The authors quantify outlier severity via the channel-wise incoherence metric \(\Psi(X) = \max_i |X_{i,:} - \mathbb{E}[X]|\) for \(X \in \{Q, K\}\), uncovering that outliers strictly reflect the 3D RoPE frequency assignment: low-frequency pairs accumulate wide angular dispersion over long sequence lengths, producing heavy outliers concentrated largely within one half-segment of each coordinate band, with near-identical outlier magnitudes across corresponding Q and K channels. To ensure mathematically equivalent attention scores \(QK^\top = (QR)(KR^{-\top})^\top\), the SVD decomposition \(R = U \Sigma L^\top\) indicates that any singular value \(\sigma_i \ne 1\) scales the \(i\)-th principal direction of \(Q\) by \(\sigma_i\) while scaling \(K\) inversely by \(1/\sigma_i\). This asymmetric scaling inevitably inflates the dynamic range and quantization error on one of the matrices. Hence, strict orthogonality (\(\Sigma = I\), \(R \in \mathrm{O}(D)\)) is mandatory when quantizing both Q and K.
The authors propose two specialized orthogonal rotation configurations: - Interleaved Rotation: A block-diagonal orthogonal matrix \(R = \mathrm{diag}(\{R_i\}_{i=1}^{D/2})\) composed of \(2 \times 2\) orthogonal blocks \(R_i \in \mathrm{O}(2)\) (initialized via \(2 \times 2\) Hadamard matrices \(H_2\)), operating directly on adjacent channel pairs \((d_{2i}, d_{2i+1})\). Because 3D RoPE rotary matrix \(M\) is inherently structured from \(2 \times 2\) planar rotation blocks, the product \(M' = RM\) maintains identical block-diagonal sparsity. Consequently, Interleaved Rotation can be statically pre-multiplied into RoPE weights offline, executing within standard element-wise RoPE kernels at exactly zero additional runtime cost. - Half Rotation: Because outliers concentrate primarily in one half-segment of each coordinate \(j \in \{f, h, w\}\), Half Rotation defines a segment-wise permutation matrix \(\Pi = \mathrm{diag}(\Pi_f, \Pi_h, \Pi_w)\) that couples dimension \(d_i^j\) with \(d_{i+D_j/2}^j\) into adjacent pairs, followed by block-diagonal rotation \(R^j\) over \(D_j/2\) orthogonal \(2 \times 2\) blocks: $\(Q \Pi R = [Q_f \Pi^f R^f, \, Q_h \Pi^h R^h, \, Q_w \Pi^w R^w]\)$ Due to its sparse structured formulation, this permutation-rotation operation is compiled into a lightweight element-wise kernel that redistributes outlier energy across half-segments with an overhead equal to only \(\sim 3\%\) of a full-dimensional rotation.
2. Range-Optimized P Quantization: Doubling Effective Quantization Resolution
During the online softmax computation in FlashAttention, local attention probabilities for each tiled block are computed as \(P_{i,j} = \exp(S_{i,j} - m_{i,j}) \in (0, 1]\), where \(m_{i,j}\) tracks running row-wise maxima. Standard signed INT4 symmetric quantization \(\lfloor (P_{i,j} / \max P_{i,j}) \times 7 \rceil\) restricts integer representations to \([0, 7]\), entirely abandoning the negative half-range \([-8, 0)\) and using only 8 discrete levels out of the 16 available in 4 bits.
RotateAttention eliminates this dynamic range waste through an affine range rectification mapping configured with a fixed scale factor of \(15\) and a fixed zero-point of \(-8\): $\(\tilde{P}_{i,j} = \left\lfloor \frac{P_{i,j}}{\max(P_{i,j})} \times 15 \right\rceil - 8 \in [-8, 7]\)$ By expanding the discrete quantization levels from 8 to 16, this design doubles the numerical resolution for the attention probability distribution. Because the scale and zero-point are analytically fixed, the transformation incurs no dynamic calibration overhead while dramatically curbing rounding truncation errors, particularly on diffuse attention patterns in text-to-video diffusion tasks.
3. Zero-Overhead Value Transformation and Selective Precision Fallback
For the Value matrix \(V\), an orthogonal Hadamard transform \(H \in \mathbb{R}^{D \times D}\) is applied to disperse outlier activations. By exploiting the associative property of matrix multiplication: $\(\mathrm{Softmax}\left(\frac{QK^\top}{\sqrt{d_k}}\right) V W_o = \mathrm{Softmax}\left(\frac{QK^\top}{\sqrt{d_k}}\right) (V H) (H^\top W_o)\)$ the transform is absorbed directly into the projection weights during offline preprocessing via \(W_v' = W_v H\) and \(W_o' = H^\top W_o\), achieving complete outlier mitigation on \(V\) with zero inference penalty.
Furthermore, acknowledging that diffusion sampling trajectories exhibit non-uniform sensitivity—where initial noise layout steps and final high-frequency detail steps are most vulnerable to quantization noise, alongside boundary DiT layers—RotateAttention implements a selective FP16 fallback strategy. In Wan2.2-T2V, retaining FP16 across the first and last 2 sampling steps and DiT blocks allows \(81\%\) of FlashAttention operations to execute under high-speed INT4 kernels. For HunyuanVideo-T2V, \(76\%\) of operations run in INT4, and \(68\%\) run in INT4 for Wan2.2-I2V, securing maximum hardware speedup while preventing catastrophic error accumulation.
Key Experimental Results¶
Main Results¶
Evaluations are conducted on leading open-source 3D RoPE DiT models: Wan2.2 (I2V / T2V at 480P resolution, 81 frames) and HunyuanVideo (T2V at 480P resolution, 129 frames). Quantitative comparisons employ relative difference metrics against FP16 baselines (Cosine Similarity, MSE, SSIM, and PSNR) along with end-to-end inference speedups measured on NVIDIA A10 and H20 GPUs.
| Model / Dataset | Datatype | Optim-P | Rotation | Cosine ↑ | MSE ↓ | SSIM ↑ | PSNR ↑ | Speedup |
|---|---|---|---|---|---|---|---|---|
| Wan2.2-I2V (480P, 81f) | FP16 | - | - | 1.0000 | 0.0000 | 1.0000 | \(\infty\) | 1.00× |
| Wan2.2-I2V | INT4 (SageAttn baseline) | \(\times\) | W/O | 0.9809 | 519.30 | 0.7752 | 22.828 | - |
| Wan2.2-I2V | INT4 | \(\checkmark\) | W/O | 0.9789 | 575.17 | 0.7524 | 22.069 | - |
| Wan2.2-I2V (Ours) | INT4 | \(\checkmark\) | Half | 0.9824 | 493.35 | 0.7767 | 22.903 | 1.51× |
| Wan2.2-I2V | INT4 | \(\checkmark\) | Interleave | 0.9792 | 583.88 | 0.7502 | 22.057 | 1.51× |
| Wan2.2-I2V | INT4 | \(\checkmark\) | Full | 0.9807 | 515.34 | 0.7808 | 23.176 | - |
| Wan2.2-T2V (480P, 81f) | FP16 | - | - | 1.0000 | 0.0000 | 1.0000 | \(\infty\) | 1.00× |
| Wan2.2-T2V | INT4 (SageAttn baseline) | \(\times\) | W/O | 0.9458 | 1574.5 | 0.6348 | 17.829 | - |
| Wan2.2-T2V | INT4 | \(\checkmark\) | W/O | 0.9518 | 1410.1 | 0.6582 | 18.427 | - |
| Wan2.2-T2V | INT4 | \(\checkmark\) | Half | 0.9492 | 1466.0 | 0.6504 | 18.462 | 1.68× |
| Wan2.2-T2V (Ours) | INT4 | \(\checkmark\) | Interleave | 0.9562 | 1392.1 | 0.6628 | 18.544 | 1.68× |
| Wan2.2-T2V | INT4 | \(\checkmark\) | Full | 0.9527 | 1370.2 | 0.6737 | 19.211 | - |
| HunyuanVideo-T2V (480P, 129f) | FP16 | - | - | 1.0000 | 0.0000 | 1.0000 | \(\infty\) | 1.00× |
| HunyuanVideo-T2V | INT4 (SageAttn baseline) | \(\times\) | W/O | 0.9663 | 1095.3 | 0.6621 | 18.332 | - |
| HunyuanVideo-T2V | INT4 | \(\checkmark\) | W/O | 0.9712 | 1089.6 | 0.6620 | 18.945 | - |
| HunyuanVideo-T2V (Ours) | INT4 | \(\checkmark\) | Half | 0.9751 | 1041.9 | 0.6677 | 18.971 | 1.61× |
| HunyuanVideo-T2V | INT4 | \(\checkmark\) | Interleave | 0.9750 | 1045.8 | 0.6588 | 19.037 | 1.61× |
| HunyuanVideo-T2V | INT4 | \(\checkmark\) | Full | 0.9744 | 1076.4 | 0.6646 | 18.854 | - |
Ablation Study¶
The table below isolates the effects of rotation types, orthogonality constraints, and learned versus analytical rotation matrices on Wan2.2-I2V:
| Configuration Variant | Structure / Constraint | Cosine ↑ | MSE ↓ | SSIM ↑ | PSNR ↑ | Note |
|---|---|---|---|---|---|---|
| Baseline w/o Rotation | No rotation | 0.9809 | 519.30 | 0.7752 | 22.828 | Severe truncation error from raw outliers |
| Only Optimized P | No rotation, Optim-P enabled | 0.9789 | 575.17 | 0.7524 | 22.069 | Low-entropy I2V attention slightly sensitive to P-shift |
| Half Rotation (Ours) | \(2 \times 2\) orthogonal, cross-half paired | 0.9824 | 493.35 | 0.7767 | 22.903 | Best balance; eliminates half-segment outlier spikes |
| Interleaved Rotation | \(2 \times 2\) orthogonal, adjacent paired | 0.9792 | 583.88 | 0.7502 | 22.057 | Fused into RoPE weights at zero runtime cost |
| Full Rotation | \(D \times D\) global Hadamard | 0.9807 | 515.34 | 0.7808 | 23.176 | Strong global mixing but adds online latency |
| Learned Rotation (Unconstrained) | Gradient trained, unconstrained \(R\) | Degraded | Degraded | Degraded | Degraded | Violates \(\Sigma = I\); scales up K outliers |
| Learned Rotation (Orthogonal) | Gradient trained, \(R_i \in \mathrm{O}(2)\) | 0.9809 | 519.94 | 0.7746 | 22.713 | Matches static initialization on small calibration sets |
Key Findings¶
- Task-Dependent Strengths of Rotation Variants: Half Rotation delivers the strongest gains on Wan2.2-I2V and HunyuanVideo where outliers sharply isolate within half-segments. In contrast, Interleaved Rotation excels on Wan2.2-T2V where local adjacent channel correlations dominate, offering an attractive zero-cost pre-fused deployment path.
- Attention Entropy and P Rectification: Range-optimized P quantization yields substantial improvements across all metrics on T2V tasks (e.g. boosting Cosine Similarity on Wan2.2-T2V from 0.9458 to 0.9518 and PSNR from 17.829 to 18.427 dB). On I2V tasks, where strong image conditioning produces low-entropy, peaked attention distributions with probabilities crowded near 0 and 1, the benefit is marginal, suggesting it is best prioritized for moderate-to-high entropy diffusion generation.
- Kernel and Pipeline Speedups: Custom CUDA implementations on NVIDIA A10 GPUs achieve a 2.2× kernel-level acceleration over custom FP16 attention. In end-to-end execution, generation speeds up by 1.68× on Wan2.2-T2V, 1.61× on HunyuanVideo, and 1.51× on Wan2.2-I2V. For long sequences (>30K tokens), rotation overhead accounts for less than 1% of total attention latency, with Half Rotation consuming only \(\sim 3\%\) of full-dimensional rotation cost.
Highlights & Insights¶
- Geometric Origin of 3D RoPE Outliers: The paper identifies that channel-wise incoherence is not random; rather, it is strictly governed by 3D RoPE frequency assignments across temporal, height, and width coordinates, providing an intuitive geometric rationale for block-sparse rotations.
- Mathematical Invariance of Orthogonality: The proof based on SVD explains why unconstrained transformations from LLM KV-cache quantization fail in video DiTs: when both Q and K are quantized for compute acceleration, any non-unitary singular value inflates the opposite matrix's dynamic range.
- Sparsity-Preserving RoPE Fusion: Exploiting the shared \(2 \times 2\) block-diagonal structure between RoPE and adjacent-channel rotation allows seamless offline operator fusion, realizing outlier suppression with zero additional runtime FLOPs.
Limitations & Future Work¶
- Coupling to 3D RoPE Partitioning: The proposed Half and Interleaved rotations are explicitly structured around 3D RoPE coordinate bands; adapting to non-RoPE architectures (e.g., learnable absolute embeddings or NoPE) requires re-identifying coordinate groupings.
- Sample Efficiency in Learned Rotations: Preliminary explorations with Riemannian optimization on orthogonal manifolds did not surpass analytical Hadamard blocks under limited 1–2 calibration samples, leaving room for richer calibration datasets.
- Interaction with Sparse Attention on Short Sequences: While rotation overhead is negligible on long sequences (>30K tokens), combining online rotations with extreme sparse attention on short sequences (<10K tokens) can elevate rotation latency to \(\sim 10\%\) of attention time, warranting integrated fused kernels.
Related Work & Insights¶
- vs SageAttention Series (SageAttention / SageAttention3): SageAttention relies on channel smoothing (\(Q - \mathbb{E}[Q]\)) and NVFP4 formats, but lacks Q/K outlier rotations for 3D RoPE and discards negative INT4 ranges on matrix \(P\); RotateAttention supplements RoPE-aware rotations and range rectification, markedly enhancing structural fidelity and PSNR.
- vs LLM KV-Cache Quantization (QuaRot / SpinQuant / FlatQuant): LLM KV quantization focuses on memory-bound autoregressive decoding, keeping Q in FP16 and permitting non-orthogonal outlier transfer; RotateAttention addresses compute-bound video DiTs where Q and K must be quantized concurrently under strict orthogonal invariance.
- vs PAROAttention: PAROAttention reorganizes tokens to minimize tile-level variance in matrix \(P\), whereas RotateAttention operates in the orthogonal channel dimension via geometric rotations and affine scaling, creating an orthogonal, complementary combination.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Unveils the geometric relationship between 3D RoPE and outlier profiles, introducing elegant fused rotations and range rectification.
- Experimental Thoroughness: ⭐⭐⭐⭐☆ Thorough evaluations across Wan2.2 and HunyuanVideo with pixel-level difference metrics, ablations, and kernel-level latency profiles.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical justification, lucid structural narrative, and clean visual comparisons.
- Value: ⭐⭐⭐⭐⭐ Provides an immediately deployable, hardware-friendly recipe for mixed-precision INT4 DiT video generation.