TAQ: Static-Deployable Temporal-Aware Quantization for Real-World Video Super-Resolution¶
Conference: ECCV 2026
Paper: ECCV 2026 Official
Code: https://github.com/imaboybut/TAQ
Area: Model Compression
Keywords: Video Super-Resolution, Post-Training Quantization, Static Deployment, Temporal Consistency, Bound Ensembling
TL;DR¶
To tackle high runtime overhead from dynamic quantization and severe temporal drift in static baselines for edge-deployed video super-resolution, TAQ introduces an offline-only temporal-aware static post-training quantization framework that combines sequence-wise histogram initialization, consecutive frame-difference refinement, and quantizer bound ensembling, achieving a 3.56× practical speedup on Jetson Orin Nano while substantially eliminating temporal flickering.
Background & Motivation¶
Real-world video super-resolution (VSR) plays a critical role in mobile visual enhancement, high-definition streaming cost reduction, and historical film restoration. However, state-of-the-art VSR architectures—often featuring recurrent feature propagation and deformable transformer blocks—demand prohibitive compute and memory resources, severely hindering real-time edge execution. Post-training quantization (PTQ) offers an appealing compression pathway because it transforms a pre-trained floating-point model into a low-precision integer network without costly retraining. In resource-constrained edge deployments, hardware accelerators (e.g., compiled via NVIDIA TensorRT) strictly necessitate static execution pipelines: once the engine graph is constructed, all quantization scale and clipping parameters must remain fixed at runtime. Input-dependent runtime range/scale recalculations ("dynamic activation quantization") break operator fusion and induce severe kernel overhead; on an NVIDIA Jetson Orin Nano, dynamic INT8 VSR inference drops to 0.266 fps (slower than FP32 at 0.325 fps, yielding a 0.79× negative speedup).
When falling back to conventional static PTQ, video models face two fundamental dilemmas. First, videos display strong sequence-wise multimodality: activation distributions shift drastically across clips due to diverse scene contents, motion magnitudes, and codec degradations. Merging all calibration sequences into a single global distribution introduces severe statistical bias—either producing overly wide bounds that magnify rounding errors on standard clips, or overly tight bounds that clip high-frequency textures. Second, quantization perturbations cause temporal error drift: image-centric PTQ treats frames independently, allowing inter-frame quantization noise to accumulate along the time axis, which surfaces as glaring banding and flickering artifacts in the reconstructed video.
Addressing the tension between the strict edge requirement of a single static parameter set and the multi-sequence statistical heterogeneity accompanied by temporal drift, this work adopts a "specialize-then-aggregate" methodology. Video awareness and temporal alignment are pushed entirely into the offline calibration phase, keeping online inference purely static with zero runtime computational overhead. Core idea: calibrate sequence-specific activation bounds during offline calibration, refine them via a consecutive frame-difference alignment loss without weight retraining to suppress temporal error drift, and aggregate the refined bounds via arithmetic mean ensembling into a single static, highly robust parameter set ready for edge deployment.
Method¶
Overall Architecture¶
The TAQ pipeline produces a fully static, affine uniform quantized VSR model without changing original model weights or network topologies. Operating strictly during offline calibration, the framework proceeds in three sequential stages: first, sequence-wise histogram initialization estimates independent activation clipping bounds for each calibration clip, capturing multimodal clip statistics; second, temporal refinement executes parallel forward passes of the floating-point reference model and the quantized model, updating only the quantizer clipping bounds via a consecutive frame-difference alignment loss to eliminate temporal drift; third, sequence-wise quantizer bounds ensembling consolidates the clip-specific bounds through arithmetic averaging into a single static parameter set, which is directly baked into static inference engines such as TensorRT INT8.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Calibration Video Set<br/>Grouped Multi-Sequence Clips"] --> B["Sequence-wise Histogram Initialization<br/>Preserve Clip-Specific Multimodal Statistics"]
B --> C["Temporal Refinement<br/>Consecutive Frame-Difference Alignment Loss"]
C --> D["Sequence-wise Quantizer Bounds Ensembling<br/>Arithmetic Mean Aggregation into Static Parameters"]
D --> E["Static Engine Deployment<br/>TensorRT INT8 Zero-Overhead Inference"]
Key Designs¶
1. Sequence-wise Histogram Initialization: Eliminating Global Calibration Distribution Bias
Conventional PTQ methods (such as MinMax, PTQ4SR, and 2DQuant) concatenate activations across all calibration videos to construct a single global histogram and solve a global mean squared error (MSE) objective. However, heterogeneous video clips feature multimodal distribution spreads that skew global statistics. TAQ abandons global mixing and instead feeds each calibration sequence separately through the model to capture layer-wise activations. Before histogram accumulation, activations undergo percentile clipping (default lower bound 0.1%, upper bound 99.9%) to strip extreme outliers. The clipped tensor is subsequently partitioned into \(K\) bins (\(K=1024\)). Given the bin center \(c_k\) and sample frequency \(w_k\), TAQ solves an independent weighted reconstruction objective per sequence:
Here \(Q_b(\cdot; \ell, u)\) represents an affine uniform fake-quantization operator with step size \(\Delta = (u - \ell)/(2^b - 1)\). This sequence-specific initialization equips each clip with an optimal initial clipping range \((\ell^{(s)}, u^{(s)})\), preventing truncation and rounding distortions while laying a stable foundation for subsequent temporal optimization.
2. Temporal Refinement: Suppressing Temporal Error Drift via Frame-Difference Alignment
Even when per-frame reconstruction errors remain minor, unconstrained frame-to-frame fluctuations in low-bit quantized models manifest as noticeable temporal flickering. Furthermore, directly minimizing per-frame absolute distances \(\|F_t - Q_t\|_2\) frequently leads to oversmoothed spatial details. TAQ introduces a temporal consistency constraint: freezing all pre-trained weights and architectural configurations, only the clipping bounds \(\theta = \{(\ell_q, u_q)\}_{q \in \mathcal{Q}}\) across weight and activation quantizers are designated as learnable parameters (optimized via the Straight-Through Estimator). Given reference floating-point output \(F_t\) and quantized output \(Q_t\) at frame \(t\), the consecutive frame differences are defined as \(\Delta F_t = F_t - F_{t-1}\) and \(\Delta Q_t = Q_t - Q_{t-1}\). TAQ establishes the temporal consistency objective:
By forcing the inter-frame transformation rate of the quantized network to match that of the floating-point baseline, static reconstruction bias cancels out within the difference operator, guiding the bounds to specifically suppress transient inter-frame jitter. Running lightweight gradient descent for a few epochs yields per-clip bounds that simultaneously deliver spatial fidelity and temporal smoothness.
3. Sequence-wise Quantizer Bounds Ensembling: Reconciling Heterogeneous Statistics with Static Deployment
Hardware acceleration on edge platforms mandates a single immutable parameter pair \((\ell_q^{\mathrm{ens}}, u_q^{\mathrm{ens}})\) per quantizer during inference. To consolidate the sequence-specific optimal parameters \(\{\hat{\theta}_q^{(s)}\}_{s \in \mathcal{S}}\) obtained across all calibration sequences \(\mathcal{S}\) into an optimal static configuration, TAQ poses an ensembling optimization problem: finding a constant bound \(\theta_q\) that minimizes the weighted sum of squared Euclidean distances to all clip-specific optima. Under an uninformative prior where all sequence weights are uniform (\(\alpha_s = 1/|\mathcal{S}|\)), the closed-form global minimizer reduces to the arithmetic mean:
Because activations were already constrained by 99.9% percentile filtering during initialization, atypical sequence outliers cannot dominate the mean; the ensemble operation further dampens individual sequence variance by a factor of \(1/|\mathcal{S}|\). This formulation eliminates the need for dynamic runtime recomputation, allowing TensorRT to compile fully static INT8 kernels with maximum compute density.
Loss & Training¶
During the temporal refinement stage, all pre-trained weights, biases, and normalization layers are strictly frozen. Only the quantizer clipping bounds \((\ell, u)\) receive gradients. The optimization utilizes the Adam optimizer with a learning rate of \(5 \times 10^{-4}\). Calibration defaults to 9 REDS sequences with 20 frame differences each, trained over 5 epochs. The entire optimization is remarkably lightweight, requiring only 32 minutes on a single NVIDIA RTX A6000 GPU—slightly faster than the image-level baseline 2DQuant (36 minutes).
Key Experimental Results¶
Main Results¶
Evaluated on the representative VSR architecture RealViformer (\(\times 4\) upscaling) across benchmark datasets REDS, SPMCS, UDM10, and in-the-wild VideoLQ without paired ground truth. The benchmarks span 2/3/4/8-bit precision, highlighting perceptual quality (LPIPS) and temporal consistency metrics (TLPIPS, TOF).
| Dataset | Metric | 2-bit (2DQuant / Ours) | 3-bit (2DQuant / Ours) | 4-bit (2DQuant / Ours) | 8-bit (2DQuant / Ours) |
|---|---|---|---|---|---|
| REDS | PSNR (dB)↑ | 24.0987 / 24.5386 | 22.0758 / 22.2444 | 23.1155 / 23.8949 | 25.7180 / 25.7768 |
| SSIM↑ | 0.5209 / 0.5944 | 0.4243 / 0.4252 | 0.5306 / 0.5480 | 0.6767 / 0.6772 | |
| LPIPS↓ | 0.7298 / 0.7142 | 0.5848 / 0.5757 | 0.4828 / 0.4724 | 0.2660 / 0.2517 | |
| TLPIPS↓ | 10.0091 / 5.7869 | 20.0603 / 15.3259 | 12.9041 / 10.0249 | 1.4566 / 1.2161 | |
| TOF↓ | 43.8434 / 31.7238 | 52.9386 / 45.2512 | 29.0666 / 23.6122 | 6.9301 / 5.9729 | |
| SPMCS | PSNR (dB)↑ | 24.6055 / 24.7626 | 22.4048 / 22.3462 | 24.0131 / 24.0377 | 25.9200 / 25.9252 |
| SSIM↑ | 0.6054 / 0.6126 | 0.4067 / 0.4462 | 0.5512 / 0.5709 | 0.6932 / 0.6941 | |
| LPIPS↓ | 0.6931 / 0.6730 | 0.5675 / 0.5710 | 0.4647 / 0.4582 | 0.2979 / 0.2754 | |
| TLPIPS↓ | 17.6929 / 14.0521 | 28.3180 / 26.1131 | 24.1784 / 19.6660 | 4.2171 / 3.8087 | |
| TOF↓ | 14.0130 / 5.8894 | 55.8117 / 30.7197 | 34.5345 / 12.5978 | 1.9064 / 1.5042 | |
| UDM10 | PSNR (dB)↑ | 28.2226 / 28.3204 | 24.7232 / 24.7274 | 27.1506 / 27.2359 | 30.3029 / 30.3708 |
| SSIM↑ | 0.7683 / 0.7744 | 0.5553 / 0.5462 | 0.6964 / 0.7244 | 0.8601 / 0.8664 | |
| LPIPS↓ | 0.6009 / 0.5971 | 0.5867 / 0.6063 | 0.4763 / 0.4604 | 0.2495 / 0.2161 | |
| TLPIPS↓ | 14.6868 / 11.3868 | 23.8387 / 21.6668 | 19.7242 / 15.5704 | 3.3868 / 2.0095 | |
| TOF↓ | 85.8837 / 62.4279 | 126.6288 / 83.1930 | 89.3016 / 71.7848 | 28.7950 / 25.0208 | |
| VideoLQ | ILNIQE↓ | 40.7119 / 39.1392 | 29.6925 / 29.3606 | 27.8252 / 27.4141 | 26.3192 / 26.2863 |
| NRQM↑ | 2.6489 / 4.7491 | 4.0320 / 6.6185 | 4.7833 / 6.0281 | 4.5939 / 4.9738 |
Physical deployment performance measured on the NVIDIA Jetson Orin Nano edge board (BasicVSR backbone, 32 frames from Vid4): - FP32 Baseline: Throughput 0.3245 fps, Latency 2.987 s, Relative Speedup \(1.00\times\) - QBasicVSR Dynamic INT8: Throughput 0.2661 fps, Latency 3.738 s, Relative Speedup \(0.79\times\) (slower than FP32) - TAQ Static INT8 (Ours): Throughput 1.1459 fps, Latency 0.839 s, Relative Speedup \(3.56\times\)
Ablation Study¶
Ablation of Core Modules (4-bit RealViformer on REDS):
| Config | PSNR (dB)↑ | SSIM↑ | LPIPS↓ | TLPIPS↓ | TOF↓ | Note |
|---|---|---|---|---|---|---|
| Global single range baseline | 23.1268 | 0.4689 | 0.5555 | 15.3454 | 55.4030 | Conventional mixed-clip global histogram calibration |
| Sequence-wise ensembling only | 23.8104 | 0.5479 | 0.4915 | 11.3696 | 27.5246 | Resolves multimodality, +0.68 dB PSNR, TOF halved |
| Full model (Seq-wise + TempRef) | 23.8949 | 0.5480 | 0.4724 | 10.0249 | 23.6122 | Frame-difference alignment further curbs temporal drift |
Ablation of Bound Ensembling Strategies (4-bit RealViformer on REDS):
| Aggregation Method | PSNR (dB)↑ | SSIM↑ | LPIPS↓ | TLPIPS↓ | TOF↓ | Note |
|---|---|---|---|---|---|---|
| Median | 23.0800 | 0.5355 | 0.4821 | 10.1316 | 23.6815 | Discards distributional spread; lowest fidelity |
| Trimmed mean | 23.8851 | 0.5451 | 0.4786 | 10.0418 | 24.2053 | Trims extreme clips; marginal gain over percentile clip |
| Arithmetic Mean (Ours) | 23.8949 | 0.5480 | 0.4724 | 10.0249 | 23.6122 | Delivers optimal balance across fidelity and temporal metrics |
Key Findings¶
- Sequence-wise calibration provides the largest fidelity boost: Transitioning from a single global bound to sequence-wise ensembling increases PSNR by 0.68 dB and slashes TOF by 27.88, demonstrating that statistical multimodality is the primary source of quantization degradation in video tasks.
- Temporal refinement directly eliminates flicker: The inter-frame difference loss specifically suppresses high-frequency temporal perturbations, cutting TLPIPS by an extra 1.34 and removing banding patterns in y–t slices.
- Dynamic quantization introduces severe edge slowdowns: Although dynamic PTQ claims adaptive bit allocations, runtime range reductions break operator execution graph fusions in TensorRT, leading to a negative speedup (\(0.79\times\)). Static quantization is mandatory for genuine acceleration on edge devices.
Highlights & Insights¶
- Offloading temporal awareness strictly to offline calibration: While previous temporal quantization methods relied on dynamic flow-guided branches during inference, TAQ elegantly captures temporal dynamics into fixed static quantizer parameters via offline ensembling.
- Bias-cancellation via consecutive frame differences: Directly penalizing per-frame reconstruction discrepancies (\(\|F_t - Q_t\|\)) causes texture blurring; aligning first-order frame differences (\(\Delta F_t \leftrightarrow \Delta Q_t\)) mathematically cancels static quantization offsets and isolates temporal transitions.
- Robust generalizability of bound ensembling: Ensembled bounds derived from only 9 REDS sequences generalize cleanly to cross-domain datasets (SPMCS, UDM10, VideoLQ) and transfer smoothly across diverse architectures from RealViformer transformers to RealBasicVSR CNNs.
Limitations & Future Work¶
- Author-admitted limitations: The temporal refinement objective relies purely on first-order temporal differences between adjacent frames without explicitly computing optical flow or dense motion vectors. In sequences with aggressive global camera rotation or extensive occlusions, large motion shifts may dilute the alignment gradient.
- Potential extensions: The current ensembling strategy assumes uniform sequence weights (\(1/|\mathcal{S}|\)). Future work could explore motion-adaptive or degradation-guided prior weights for calibration aggregation, as well as extending static temporal-aware quantization to emerging video diffusion models.
Related Work & Insights¶
- vs PTQ4SR / 2DQuant: Both represent leading PTQ frameworks tailored for single-image super-resolution. However, concatenating video frames into a global pool ignores temporal drift and clip multimodality. TAQ preserves their lightweight benefits while outperforming them by wide margins in temporal metrics at low bit-widths.
- vs QBasicVSR: QBasicVSR depends on inference-time optical flow estimation for dynamic bit-width selection and activation scaling, which prevents TensorRT from performing static kernel optimizations (\(0.79\times\) on Jetson Orin Nano). In contrast, TAQ operates in a strictly static regime, achieving higher restoration fidelity under a lower average bit budget and delivering a \(3.56\times\) real-world speedup.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Introduces an offline specialize-then-ensemble paradigm that preserves static execution while fully capturing video temporal dynamics]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive multi-bit evaluations across four diverse benchmarks, full ablations, and real hardware profiling on Jetson Orin Nano]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear problem formulation, tightly coupled motivation and methodology, accompanied by elegant mathematical derivations]
- Value: ⭐⭐⭐⭐⭐ [Provides an immediately practical, high-throughput deployment blueprint for edge video super-resolution]