NanoVSR: Towards Real-Time Video Super-Resolution on Edge Devices¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Video Generation
Keywords: Video Super-Resolution, Edge AI, Structural Reparameterization, Bidirectional Recurrent Network, Implicit Alignment
TL;DR¶
NanoVSR eliminates expensive optical flow and spatial-temporal attention in favor of a fully convolutional bidirectional recurrent framework with direct additive propagation and structural reparameterization, collapsing into a single stream of standard convolutions at inference to achieve 27.20 FPS on an NVIDIA Jetson Orin NX (25W) with 28.64 dB PSNR on REDS4.
Background & Motivation¶
Video super-resolution (VSR) aims to restore sharp high-resolution (HR) frames from degraded low-resolution (LR) video sequences by leveraging redundant complementary spatial-temporal cues across consecutive frames. Contemporary state-of-the-art VSR architectures, ranging from optical flow-based recurrent models like BasicVSR++ and deformable convolution networks like EDVR to attention-heavy transformers like RVRT and VRT, have continuously set new benchmarks in restoration fidelity. However, these leading methods impose immense computational overhead and memory bandwidth pressure: they depend on dense optical flow estimation, quadratic-complexity attention mechanisms, or custom CUDA kernels that fail to fuse efficiently on hardware accelerators. Consequently, while they run reasonably on enterprise cloud GPUs, their latency exceeds real-time thresholds on resource-constrained embedded edge devices.
Existing efficiency-oriented explorations face a structural bottleneck. Standard recurrent VSR designs rely on feature concatenation to merge incoming frame features with propagating hidden states. This concatenation doubles the channel dimension for subsequent convolutions, creating significant memory access overhead and inflating runtime. Meanwhile, keeping an explicit motion compensation module (such as SPyNet) incurs massive parameter and latency penalties that overshadow the super-resolution backbone itself. Conversely, naively eliminating motion alignment degrades reconstruction quality and introduces temporal flickering under rapid camera motion or complex object trajectories.
The key entry point of this work stems from the realization that throughput on edge accelerators (such as TensorRT) is primarily bounded by fragmented memory accesses and small kernel launches rather than raw parameter count alone. Core idea: build a purely convolutional bidirectional recurrent architecture using direct element-wise additive temporal propagation and structural reparameterization, which trains as an expressive multi-branch topology but collapses into a streamlined chain of plain convolutions at inference, complemented by a two-stage progressive curriculum to implicitly learn motion compensation without explicit flow.
Method¶
Overall Architecture¶
NanoVSR is designed as a purely convolutional, bidirectional recurrent network constructed entirely from standard tensor operations to ensure seamless ONNX compatibility and maximize TensorRT deep engine fusion. Given a low-resolution sequence \(x \in \mathbb{R}^{T \times 3 \times H \times W}\) across temporal window \(T\), the execution pipeline proceeds through three distinct stages: shallow feature extraction, bidirectional additive propagation, and high-resolution reconstruction.
In the initial stage, the temporal dimension is folded into the batch dimension, allowing all frames to be processed concurrently through a single reparameterizable convolutional block that maps the RGB input to a \(C\)-channel latent representation. After unfolding the temporal dimension, the features are forwarded through bidirectional recurrent networks where hidden states are updated sequentially using direct element-wise addition instead of concatenation. Finally, the forward and backward hidden representations are concatenated channel-wise, compressed back to \(C\) channels via a \(1 \times 1\) fusion convolution, and upscaled to the HR space through a cascaded sub-pixel convolution (PixelShuffle) block before adding a globally upsampled bilinear residual baseline.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Low-Resolution Input Video x"] --> B["Shallow Feature Extraction<br/>RepVGG Block mapping to C channels"]
B --> C["Bidirectional Additive Temporal Propagation<br/>Forward & Backward RepVGG sequences"]
C --> D["Bidirectional Feature Fusion<br/>1x1 Conv bottlenecking to C channels"]
D --> E["Cascaded Sub-Pixel Reconstruction<br/>Two-stage PixelShuffle 4x upsampling"]
A -->|Fast Bilinear Upsampling 4x| F["Global Residual Base"]
E --> G["Element-wise Residual Addition"]
F --> G
G --> H["High-Resolution Video Sequence y"]
Key Designs¶
1. Direct Element-Wise Additive Propagation: Eliminating Channel Concatenation Bottlenecks Standard recurrent VSR models (such as BasicVSR and FRVSR) fuse the prior hidden state with the current frame feature via channel-wise concatenation. This doubles the channel width entering subsequent convolutions, heavily inflating memory traffic and floating-point operations. To overcome this limitation, NanoVSR adopts a direct element-wise additive update formulation for the forward and backward hidden states: $\(h^{\rightarrow}_i = \mathcal{H}_{\text{forward}}(f_i + h^{\rightarrow}_{i-1}), \quad h^{\leftarrow}_i = \mathcal{H}_{\text{backward}}(f_i + h^{\leftarrow}_{i+1})\)$ where the boundary hidden states \(h^{\rightarrow}_0\) and \(h^{\leftarrow}_{T+1}\) are initialized as zero tensors, and \(\mathcal{H}_{\text{forward}}\) and \(\mathcal{H}_{\text{backward}}\) denote sequential structural blocks with LeakyReLU activations. This additive formulation entirely bypasses channel dimension inflation, drastically cuts intermediate memory access overhead, and compels the network to implicitly model residual spatio-temporal displacements across time steps.
2. Structural Reparameterization: Multi-Branch Training Decoupled from Single-Stream Deployment Multi-branch modules offer superior representational capacity and stable gradient dynamics during optimization, but their branched topologies fragment memory bandwidth and trigger frequent CUDA kernel launch penalties during edge deployment. NanoVSR addresses this dilemma via structural reparameterization. During training, every fundamental building unit is instantiated as a multi-branch block comprising parallel \(3 \times 3\) convolution, \(1 \times 1\) convolution, and identity mapping pathways, each coupled with a Batch Normalization (BN) layer. Prior to deployment, the BN affine parameters and running statistics are absorbed into the corresponding convolution kernels and bias vectors. The parallel pathways are then zero-padded and algebraically merged into a single, mathematically equivalent \(3 \times 3\) convolution kernel: $\(W_{\text{fused}} = W'_{3\times3} + \text{Pad}(W'_{1\times1}) + \text{Pad}(W'_{\text{identity}})\)$ Consequently, the multi-branch structure completely collapses into a plain, continuous convolutional stream during edge execution, enabling hardware accelerators like TensorRT to unlock optimal kernel fusion and peak inference throughput.
3. Two-Stage Progressive Curriculum: Guiding Implicit Temporal Alignment Without Flow Training an implicit-alignment recurrent network from scratch over long sequences with substantial motion is prone to numerical instability and gradient explosion. NanoVSR employs a progressive curriculum training strategy: during Phase 1, the model is pre-trained for 50,000 iterations on short 7-frame sequences from Vimeo-90K, focusing on stable spatial feature representation and short-range temporal fusion. In Phase 2, the training corpus transitions to REDS with temporal windows expanded to 30 continuous frames via sliding-window sampling for an additional 100,000 iterations. This long-sequence regime forces the bidirectional recurrent units to learn long-range motion dependencies implicitly, achieving high temporal stability without explicit optical flow guidance.
Loss & Training¶
The network is optimized end-to-end for 150,000 iterations with a global batch size of 12 and low-resolution patch dimensions of \(256 \times 256\). Training utilizes the smooth Charbonnier loss penalty (\(\epsilon = 1 \times 10^{-6}\)): $\(\mathcal{L} = \sqrt{\| \hat{y} - y_{\text{GT}} \|^2 + \epsilon^2}\)$ Optimization is performed via the Adam optimizer (\(\beta_1 = 0.9, \beta_2 = 0.99\)) accompanied by a unified Cosine Annealing learning rate schedule decaying from \(3 \times 10^{-4}\) down to \(1 \times 10^{-7}\). To preserve high training efficiency, BFloat16 Automatic Mixed Precision (AMP) and gradient clipping with a maximum norm of 0.5 are enforced throughout, alongside geometric augmentations and the CutBlur strategy.
Key Experimental Results¶
Main Results¶
NanoVSR is thoroughly benchmarked against both heavy state-of-the-art and compact efficiency-focused baselines across REDS4 (RGB), Vid4 (Y-channel), and Vimeo-90K-T (Y-channel). Execution latency is benchmarked on an NVIDIA H100 GPU in FP32 on \(180 \times 320\) inputs; edge device throughput (FPS) is evaluated on an NVIDIA Jetson Orin NX (16GB, 25W TDP) compiled with TensorRT 10.3 in FP16 precision.
| Model | Params | H100 Latency (ms) | REDS4 (dB / SSIM) | Vid4 (dB / SSIM) | Vimeo-90K-T (dB / SSIM) | Orin NX (25W) FPS |
|---|---|---|---|---|---|---|
| NanoVSR-226k | 226k | 1.910 | 28.23 / 0.8057 | 25.26 / 0.7252 | 34.31 / 0.9130 | 43.86 |
| NanoVSR-644k (Flagship) | 644k | 2.982 | 28.64 / 0.8215 | 26.05 / 0.7761 | 35.00 / 0.9226 | 27.20 |
| NanoVSR-1.7M | 1.7M | 4.268 | 29.15 / 0.8364 | 26.44 / 0.7964 | 35.49 / 0.9294 | 19.58 |
| NanoVSR-5.4M | 5.4M | 8.547 | 29.73 / 0.8526 | 26.76 / 0.8089 | 35.85 / 0.9335 | 8.66 |
| MobileRNN (MAI'22) | 474k | 5.676 | โ / โ | 25.42 / 0.7369 | 34.45 / 0.9156 | โ |
| SPAN (SISR Baseline) | 498k | 5.820 | 28.48 / 0.8109 | 25.43 / 0.7357 | 35.06 / 0.9213 | โ |
| EDVR-M | 3.3M | 27.343 | 30.53 / 0.8699 | 27.10 / 0.8186 | 37.09 / 0.9446 | โ |
| BasicVSR | 6.2M | 15.160 | 31.42 / 0.8909 | 27.24 / 0.8251 | 37.18 / 0.9450 | 8.04 |
| RVRT (SOTA) | 10.8M | 47.867 | 32.75 / 0.9113 | 27.99 / 0.8462 | 38.15 / 0.9527 | Non-real-time |
On perceptual and spatio-temporal metrics, NanoVSR-644k achieves an LPIPS of 0.2546 and an ST-RRED of \(6.64 \times 10^{-5}\) on REDS4, decisively outperforming the frame-by-frame SISR baseline SPAN (LPIPS 0.2757, ST-RRED \(8.91 \times 10^{-5}\)) and validating the superior temporal coherence generated by implicit recurrence.
Ablation Study¶
Comprehensive ablation experiments conducted on the NanoVSR-226k baseline across different structural configurations (measured on H100 GPU):
| Config | Params | Latency (ms) | REDS4 PSNR/SSIM | Vid4 PSNR/SSIM | Note & Key Observation |
|---|---|---|---|---|---|
| NanoVSR-226k (Reference) | 226k | 1.910 | 28.23 / 0.8057 | 25.26 / 0.7252 | Full fused deployment; exact same quality with \(1.82\times\) faster runtime |
| w/o Reparameterization (-NOFUSE) | 245k | 3.471 | 28.23 / 0.8057 | 25.26 / 0.7252 | Retains 3 branches at test time; latency jumps by 81.7% due to fragmented kernels |
| w/o Multi-Branch Training (-SINGLE) | 226k | 1.885 | 28.13 / 0.8027 | 25.25 / 0.7296 | Trained as plain conv from scratch; drops 0.10 dB on REDS4 due to limited capacity |
| w/o Curriculum Pre-training (-NOPRET) | 226k | 1.887 | 28.23 / 0.8054 | 25.17 / 0.7271 | Omits Vimeo-90K Phase 1; zero-shot Vid4 PSNR drops by 0.09 dB |
| w/ Explicit Alignment (+SPYNET) | 1.7M | 3.953 | 29.44 / 0.8460 | 25.95 / 0.7701 | Flow boosts PSNR by 1.21 dB, but parameters jump \(7.5\times\) and latency doubles |
| Unidirectional Inference (-ONEWAY) | 151k | 1.470 | 28.04 / 0.8005 | 25.00 / 0.7107 | Removes backward recurrence for zero-latency streaming; drops 0.26 dB on Vid4 |
Key Findings¶
- Reparameterization delivers asymmetric deployment dividends: Multi-branch topologies enrich gradient flow during training, while algebraic fusion at test time achieves identical restoration fidelity while cutting inference latency from 3.471 ms to 1.910 ms.
- Explicit optical flow is computationally prohibitive on edge devices: While integrating SPyNet raises REDS4 PSNR by 1.21 dB, it introduces ~1.5M parameters exclusively for motion estimation and more than doubles inference latency, immediately rendering real-time streaming infeasible on edge platforms.
- Capacity scaling exhibits diminishing returns beyond 5.4M: Progressively scaling network width and depth lifts PSNR from 27.62 dB (22k) to 28.64 dB (644k) and 29.73 dB (5.4M); however, scaling further to 9.6M yields only an additional 0.17 dB, highlighting 644k as the Pareto-optimal operating point for resource-constrained edge inference.
Highlights & Insights¶
- Hardware-aligned plain convolution philosophy: NanoVSR acknowledges that FLOPs and theoretical parameter counts poorly reflect edge execution speed; standard single-stream convolutions maximize hardware kernel fusion and memory locality in TensorRT engines.
- Direct additive hidden states over channel concatenation: Substituting standard channel concatenation with element-wise addition cuts convolution channel footprints in half, resolving intermediate memory bandwidth bottlenecks in recurrent architectures.
- Curriculum-driven implicit alignment: Progressing from short 7-frame sequences to 30-frame temporal windows allows bidirectional recurrent networks to implicitly capture complex motion trajectories without explicit warping or flow fields.
Limitations & Future Work¶
- High-frequency degradation under extreme non-rigid motion: Lacking explicit geometric alignment, the network exhibits softer edge reconstructions when processing rapid scene pans or turbulent non-rigid movements compared to large transformer models.
- Buffer latency inherent to bidirectional recurrence: The bidirectional temporal window requires buffering 15 look-ahead frames (\(T = 15\)), introducing pipeline delay that must be bypassed via the unidirectional mode (-ONEWAY) for hard real-time interactive constraints.
- Validation on higher input resolutions: Benchmarks currently focus on \(180 \times 320\) and \(270 \times 480\) inputs upscaled \(4\times\); extending evaluations to higher-resolution regimes (\(720\text{p} \to 4\text{K}\)) under edge thermal and memory throttling warrants further investigation.
Related Work & Insights¶
- vs BasicVSR / BasicVSR++: BasicVSR establishes the standard bidirectional recurrent paradigm but requires optical flow and feature concatenation that block real-time edge execution; NanoVSR demonstrates that additive recurrent propagation with curriculum learning achieves strong temporal coherence without explicit flow.
- vs RepNet-VSR / RepVGG: Inherits structural reparameterization concepts but integrates them into bidirectional recurrent temporal aggregation, overcoming memory bottlenecks specific to video super-resolution.
- vs SPAN / MobileRNN: Compared to the lightweight SISR model SPAN, NanoVSR-644k runs at half the latency while dramatically improving spatio-temporal consistency and fidelity; compared to MobileRNN from the Mobile AI 2022 challenge, NanoVSR achieves nearly \(2\times\) higher throughput without custom recurrent operator overheads.
Rating¶
- Novelty: โญโญโญโญโ Elegantly merges structural reparameterization with direct additive recurrent propagation to resolve edge VSR memory bandwidth bottlenecks.
- Experimental Thoroughness: โญโญโญโญโญ Rigorous evaluations on three public benchmarks, complete ablation analyses, and concrete profiling on NVIDIA Jetson Orin NX (8GB/16GB) hardware.
- Writing Quality: โญโญโญโญโญ Clear structural exposition, precise hardware profiling analysis, and well-motivated architectural decisions.
- Value: โญโญโญโญโญ Offers an exceptionally practical and high-throughput blueprint for deploying real-time video super-resolution on edge computing platforms.