Skip to content

Improving Image-to-Image Translation via a Rectified Flow Reformulation

Conference: ECCV 2026
Paper: ECCV Official
Full Text Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5655.txt
Area: Image Generation
Keywords: image-to-image translation, rectified flow, regression reformulation, ODE-based refinement, continuous-time transport

TL;DR

This paper introduces Image-to-Image Rectified Flow Reformulation (I2I-RFR), a practical plug-in framework that converts standard discriminative I2I regression backbones into continuous-time transport refiners by concatenating noise-corrupted targets, optimizing a time-reweighted pixel loss, and analytically inducing a velocity field, achieving high-fidelity 3-step explicit Euler integration without architectural redesign or distillation.

Background & Motivation

Image-to-image (I2I) translation represents a foundational cornerstone in low-level computer vision, encompassing diverse tasks such as single-image super-resolution, deblurring, low-light image enhancement, underwater restoration, and old film video recovery. In both research and deployment scenarios, the dominant and most practical paradigm formulates I2I translation as supervised regression: given an input image \(x\), a deep neural network predicts the clean target image \(y\) and optimizes pixel-wise metrics like \(\ell_1\) or mean squared error. This framework is remarkably stable, easy to optimize, and has nurtured a rich ecosystem of degradation-specific architectures (such as U-Nets, SwinIR, Restormer, and LEDNet), which embody years of inductive bias engineering. However, real-world image degradation is inherently ill-posed: severe blur, heavy noise, and structural information loss frequently render the conditional posterior distribution \(p(y | x)\) multi-modal. When confronted with multi-modal targets, standard pixel-wise regression mathematically collapses toward a point estimate (the conditional mean or median), causing severe over-smoothing, textural detail loss, and unnatural regression artifacts.

Generative objectives offer a principled path to bypass the conditional mean collapse, yet existing generative formulations carry substantial practical friction. Conditional generative adversarial networks (cGANs) rely on a discriminator to enforce realistic high-frequency distributions, but the delicate minimax game between generator and discriminator is notorious for optimization instability, sensitive hyperparameters, and hallucinations. Denoising diffusion probabilistic models (such as SR3 and Palette) excel at modeling complex multi-modal distributions, yet they typically demand diffusion-specific parameterizations, delicate noise schedules, time-step conditioning modules, and iterative sampling trajectories requiring dozens to hundreds of denoising steps. Latent diffusion models and pretrained foundation models further introduce external autoencoders and billion-parameter generative priors, incurring prohibitive inference latencies while dismantling the lightweight inductive biases and native simplicity of task-specialized paired regression networks.

Rectified Flow provides an appealing alternative by learning a continuous-time transport velocity field over linear trajectories between two distributions, allowing deterministic ODE integration with controllable speed-quality trade-offs. Nevertheless, existing rectified flow models have primarily been designed for unconditional or text-conditioned generation, relying on generic time-conditioned Transformer backbones that neglect the strong, pixel-aligned conditioning present in paired I2I restoration. This paper asks: can we harness the progressive refinement benefits of continuous-time transport while preserving the simplicity, stability, and inductive biases of existing paired regression backbones? Core idea: reformulate standard I2I regression networks into continuous-time rectified flow refiners (I2I-RFR) by concatenating noise-corrupted targets, optimizing a time-reweighted pixel loss that analytically induces an RF-consistent velocity field, and achieving high-quality progressive refinement in merely 3 explicit Euler integration steps without distillation.

Method

Overall Architecture

The design philosophy of I2I-RFR is to equip any established paired I2I regression architecture with continuous-time transport capability with minimal structural intrusion. The system consists of two synergistic phases: a time-reweighted clean-target training pipeline and a few-step explicit ODE inference solver. During training, given an input conditioning image \(x\) and ground-truth target \(y\), a linear noisy target state \(y_t\) is constructed and concatenated with \(x\) along the channel dimension. The backbone network \(f_{\theta'}\) receives \([x; y_t]\) and directly predicts the clean target image \(\hat{y}\). During inference, starting from standard Gaussian noise, the model evaluates the induced velocity field from the network prediction and the current state, integrating backward along the time trajectory via 3 explicit Euler discretization steps to progressively refine the output.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Input image x and target y"] --> D1["Noisy Target State Concatenation<br/>Form linear path y_t and channel-concat [x; y_t]"]
    D1 --> D2["Beta-Distributed Time Sampling<br/>Sample t ~ Beta(p+1, 1) to offset 1/t variance"]
    D2 --> D3["Clean Target Prediction & Induced Velocity<br/>Predict clean image and induce analytical RF velocity"]
    D3 --> D4["Few-Step Explicit Euler Integration<br/>Solve ODE backward in 3 steps from Gaussian noise"]
    D4 --> Out["High-fidelity output image y_hat"]

Key Designs

1. Noisy Target State Concatenation: minimal extension of existing regression backbones Prior attempts to incorporate generative dynamics into I2I tasks often forced engineers to modify intermediate feature layers to insert time-embedding modulation or cross-attention blocks. I2I-RFR removes this friction by capitalizing on the strict pixel-aligned spatial correspondence between the degraded input and the target image. Specifically, defining the clean target endpoint as \(x_0 = y\) and the Gaussian noise endpoint as \(x_1 = \varepsilon \sim \mathcal{N}(0, I)\), a linear trajectory yields the intermediate noisy target: $\(y_t = (1 - t)y + t\varepsilon, \quad t \in (0, 1]\)$ The existing network architecture is preserved entirely, needing only an expansion of its input layer channels from \(C_{\text{in}}\) to \(C_{\text{in}} + C_{\text{out}}\) to accommodate the concatenated input \([x; y_t]\). The conditioning image \(x\) provides strong structural spatial anchors, while \(y_t\) directly exposes the corruption level and high-frequency residuals to the network. This eliminates the necessity of dedicated time-embedding MLPs, allowing the model to naturally infer the noise scale while directly reusing specialized restoration backbones.

2. Beta-Distributed Time Sampling: counterbalancing weighting variance and ensuring noise coverage Training continuous-time flow models with uniform time sampling \(t \sim \mathcal{U}(0, 1)\) induces severe optimization instabilities because the analytical loss function incorporates a scale factor of \(t^{-p}\). As \(t \to 0\), this weight explodes, producing excessive gradient variance. To stabilize training, I2I-RFR introduces a non-uniform sampling strategy based on a Beta distribution: $\(t \sim \text{Beta}(p+1, 1), \quad t \leftarrow \max(t, t_{\min})\)$ Under the default \(\ell_1\) loss (\(p=1\)), time steps are drawn from \(\text{Beta}(2, 1)\), whose probability density is proportional to \(t^1\). This density analytically cancels out the \(t^{-1}\) weighting term in the loss expectation, rendering the effective gradient magnitude approximately uniform across time steps. Sampling is implemented efficiently via inverse transform sampling (\(t = u^{1/(p+1)}\) with \(u \sim \mathcal{U}(0, 1)\)), clamped at \(t_{\min} = 10^{-3}\) to guard against boundary singularities. Crucially, this distribution provides dense sampling coverage in the high-noise regime (\(t \approx 1\)), which is vital for robust initialization during backward ODE integration.

3. Clean Target Prediction & Induced Velocity: avoiding parameterization mismatch in restoration networks Standard Rectified Flow and Flow Matching frameworks routinely adopt explicit velocity regression (\(v\)-prediction), training the network to directly predict \(v^* = x_1 - x_0 = \varepsilon - y\). However, modern image restoration backbones are engineered with multi-scale feature hierarchies, residual connections, and frequency priors specifically tailored to remove noise and reconstruct clean photographic content. Forcing these models to output a white-noise-dominated velocity vector severely degrades their feature representations. I2I-RFR preserves the clean target prediction paradigm: the network continues to output the reconstructed clean target \(f_{\theta'}([x; y_t])\), and the corresponding rectified flow velocity field is induced analytically: $\(v_{\theta'}(x, y_t, t) = \frac{y_t - f_{\theta'}([x; y_t])}{t}\)$ Because the linear interpolation identity yields \(y_t - y = t(\varepsilon - y) = t v^*\), minimizing the time-reweighted clean-image regression loss: $\(\mathcal{L}_{\text{RFR}} = \mathbb{E}\left[ \left\| \frac{y - f_{\theta'}([x; y_t])}{t} \right\|_1 \right]\)$ is mathematically identical to minimizing the true velocity regression objective \(\mathbb{E}[\| v^* - v_{\theta'}(x, y_t, t) \|_1]\). This parameterization allows the model to leverage its native inductive biases for degradation suppression, bypassing the catastrophic failures observed with explicit velocity regression.

4. Few-Step Explicit Euler Integration: fast distillation-free inference via spatial conditioning Generic diffusion and flow generative models typically require 25 to 100 sampling steps, and accelerating them to 1โ€“4 steps generally requires multi-stage distillation pipelines that add massive training complexity. In I2I-RFR, because the conditioning image \(x\) provides explicit spatial and low-frequency structure, the ODE state \(y_t\) functions as a localized refinement variable rather than an unconstrained generative latent. Consequently, the trajectory can be integrated accurately with a low-order numerical solver. Starting at \(y_{t=1} \sim \mathcal{N}(0, I)\), the backward-time ODE is solved using \(N=3\) uniform explicit Euler steps with \(\Delta t = 1/N\): $\(y_{t_{n+1}} = y_{t_n} - \Delta t \cdot v_n = y_{t_n} + \frac{1}{N} \frac{f_{\theta'}([x; y_{t_n}]) - y_{t_n}}{t_n}\)$ Each integration step shifts the current noisy state toward the model's clean prediction. In only 3 steps, the solver eliminates over-smoothing artifacts and recovers sharp textures with minimal runtime overhead, achieving high-speed inference without requiring any distillation stage.

Loss & Training

The standard training objective uses the \(t^{-1}\)-reweighted \(\ell_1\) distance. When integrated into domain-specific pipelines utilizing composite objectives (denoted as I2I-RFR-CL), the framework offers seamless backward compatibility: all high-level perceptual or feature-space losses are preserved intact, while only the low-level pixel regression term is replaced by the reweighted noisy-target counterpart. Optimization uses the Adam optimizer with an initial learning rate of \(10^{-4}\) decayed to \(10^{-6}\) via a cosine schedule.

Key Experimental Results

Main Results

To establish the broad applicability of I2I-RFR across distinct vision domains and architectures, the table below consolidates quantitative benchmarks on single-image super-resolution (\(\times 4\) on Urban100), real-world image deblurring (RealBlur-J), and low-light deblurring (LOLBlur). The evaluation spans CNNs, Transformers, classical U-Nets, and the diffusion-oriented Palette backbone retrained with I2I-RFR.

Task / Benchmark Backbone Paradigm / Variant PSNR (dB) โ†‘ SSIM โ†‘ LPIPS โ†“
Super-Resolution (Urban100, ร—4) NLSN Standard \(\ell_1\) Regression 27.22 0.8176 0.1945
NLSN + I2I-RFR 26.83 (-0.39) 0.8079 (-0.0097) 0.1996 (+0.0051)
SwinIR Standard \(\ell_1\) Regression 27.24 0.8199 0.1948
SwinIR + I2I-RFR 27.19 (-0.05) 0.8180 (-0.0019) 0.1897 (-0.0051)
U-Net Standard \(\ell_1\) Regression 25.97 0.7810 0.2414
U-Net + I2I-RFR 25.55 (-0.42) 0.7707 (-0.0103) 0.2218 (-0.0196)
Palette Backbone Original DDPM (Multi-step) 19.58 0.5193 0.3035
Palette Backbone + I2I-RFR (3 steps) 26.40 (+6.82) 0.7973 (+0.2780) 0.1914 (-0.1121)
Image Deblurring (RealBlur-J) NAFNet Standard Regression 28.62 0.8683 0.1557
NAFNet + I2I-RFR 28.74 (+0.12) 0.8737 (+0.0054) 0.1449 (-0.0108)
Restormer Standard Regression 28.73 0.8677 0.1522
Restormer + I2I-RFR 29.04 (+0.31) 0.8812 (+0.0135) 0.1408 (-0.0114)
FFTFormer Standard Regression 28.19 0.8535 0.1691
FFTFormer + I2I-RFR 28.58 (+0.39) 0.8685 (+0.0150) 0.1487 (-0.0204)
Low-Light Deblurring (LOLBlur) LEDNet Baseline Regression 26.04 0.8450 0.1590
LEDNet + Adversarial Training (GAN) 24.55 0.8110 0.1530
LEDNet + I2I-RFR 27.42 (+1.38) 0.8740 (+0.0290) 0.1380 (-0.0210)
DarkIR Baseline Model 27.31 0.8960 0.1110
DarkIR + I2I-RFR-CL (Composite Loss) 27.72 (+0.41) 0.9000 (+0.0040) 0.1030 (-0.0080)

Ablation Study

The ablation experiments isolate the impact of the formulation's parameterization (direct clean target prediction vs. explicit \(v\)-prediction) and the effect of time-embedding modules. The table below presents quantitative comparisons on low-light enhancement (LOLBlur) and old film restoration (DAVIS video test set):

Task / Evaluation Split Configuration / Ablation Variant PSNR (dB) โ†‘ SSIM โ†‘ LPIPS โ†“ Empirical Phenomenon & Mechanism Analysis
Low-Light Deblurring (LOLBlur) LEDNet + I2I-RFR (Default: Target Pred) 27.42 0.8740 0.1380 Converges smoothly, effectively suppressing heavy noise while sharpening edges
LEDNet + I2I-RFR (\(v\)-prediction) 11.57 (-14.47) 0.0600 (-0.7850) 0.8030 (+0.6440) Catastrophic collapse due to severe parameterization mismatch with restoration inductive bias
DarkIR + I2I-RFR (Default) 27.19 0.8970 0.1060 Consistently strong distortion metrics with clear perceptual gain
DarkIR + I2I-RFR (\(v\)-prediction) 16.73 (-10.58) 0.4310 (-0.4650) 0.5120 (+0.4010) Drastic degradation; frequency-filtering modules fail on high-variance noise targets
DarkIR + I2I-RFR + Time Embedding 26.80 (-0.51) 0.8770 (-0.0190) 0.1270 (+0.0160) Explicit time conditioning introduces redundant features and destabilizes optimization
Old Film Video Restoration (DAVIS) RTN + I2I-RFR (Default: Target Pred) 25.05 0.8330 0.2100 Preserves temporal consistency and suppresses severe historical scratches
RTN + I2I-RFR (\(v\)-prediction) 20.76 (-3.78) 0.7127 (-0.1167) 0.3025 (+0.1463) Fails to maintain inter-frame temporal coherence, amplifying visual flicker
RRTN + I2I-RFR (Default) 25.87 0.8659 0.1699 Massive +3.21 dB PSNR improvement over baseline RRTN (22.66 dB)
RRTN + I2I-RFR (\(v\)-prediction) 24.72 (-1.15) 0.8566 (-0.0093) 0.1565 Preserves plausible textures but falls behind direct target prediction on fidelity

Key Findings

  • Clean target prediction is essential for restoration backbones: Across both LEDNet and DarkIR, switching from direct target prediction to explicit \(v\)-prediction triggered catastrophic performance degradation (PSNR drops of 14.47 dB and 10.58 dB). Task-tailored restoration backbones are engineered to filter out degradations and recover sharp signals; forcing them to fit stochastic velocity targets containing raw Gaussian noise contradicts their architectural inductive bias.
  • Explicit time embeddings are redundant in paired settings: Adding sinusoidal time embeddings into DarkIR degraded PSNR by 0.51 dB and increased LPIPS by 0.016. Because the input condition \(x\) and the noisy state \(y_t\) are concatenated directly at the input, the network naturally perceives the noise scale, rendering dedicated time embeddings unnecessary and potentially harmful to optimization stability.
  • Perception-distortion trade-offs in super-resolution: In pixel-aligned single-image super-resolution, where minor hallucinations can reduce pixel-level matching despite improving visual realism, I2I-RFR yields substantial perceptual gains (lower LPIPS) while maintaining comparable or slightly lower PSNR/SSIM. In contrast, on severely ill-posed restoration tasks (deblurring and low-light enhancement), I2I-RFR delivers simultaneous leaps in both perceptual fidelity and distortion metrics.

Highlights & Insights

  • Plug-and-play mathematical equivalence: The core mathematical elegance lies in demonstrating that time-reweighted pixel regression on noise-augmented inputs analytically induces a Rectified Flow velocity field, establishing an unforced bridge between discriminative regression and continuous-time transport.
  • Distillation-free 3-step inference: By harnessing the strong spatial guidance of the conditioning image, the framework compresses the sampling trajectory from the typical 25โ€“100 steps of standard flow/diffusion models down to just 3 explicit Euler steps without needing any auxiliary distillation stages.
  • Universal architectural compatibility: With a trivial expansion of the first-layer channel dimension, I2I-RFR seamlessly enhances CNNs, Transformers, standard U-Nets, and recurrent video architectures, unlocking generative refinement across legacy restoration codebases.

Limitations & Future Work

  • Unsuitability for extreme unaligned translation: The framework heavily presumes pixel-level or structural spatial alignment between the conditioning input \(x\) and the target \(y\). In tasks involving radical geometric shifts (e.g., cross-view synthesis or unpaired style transfer), channel concatenation and 3-step Euler integration are insufficient to guide trajectory transport.
  • Slight distortion drop in mild super-resolution: In mild degradation settings where the mapping is nearly deterministic and one-to-one, continuous-time stochastic sampling can introduce subtle high-frequency variance that slightly penalizes pixel-level PSNR.
  • Future directions: Adapting the reformulation to cross-attention conditioning for loosely paired multi-modal tasks, and investigating its potential in 3D radiance field restoration and non-photographic rendering domains.
  • vs. DDPM / Palette (Conditional Diffusion): Palette requires tuning complex variance noise schedules and executing dozens to hundreds of iterative denoising steps. Retraining the Palette backbone under I2I-RFR with 3-step Euler sampling delivers dramatic quality boosts over native DDPM training (e.g., LOLBlur PSNR surges from 8.03 dB to 26.88 dB while slashing sampling latency).
  • vs. cGAN / SRGAN (Adversarial Approaches): Adversarial learning requires delicate balancing of generator-discriminator dynamics and is prone to collapse or artificial artifacts. I2I-RFR dispenses with the discriminator entirely, achieving superior realism and detail preservation solely through stable supervised regression (outperforming LEDNet+Adv. by 2.87 dB PSNR and 0.015 LPIPS).
  • vs. Large Foundation Generative Restorers (DiffBIR / SeeSR / FlowIE): Billion-parameter models leveraging Stable Diffusion 2.1 priors produce strong perceptual hallucination but require 1600Mโ€“2500M parameters and 6โ€“8 seconds per image. SwinIR + I2I-RFR operates at 12M parameters and 0.14 seconds, providing an accessible, faithful paired restoration solution for resource-constrained deployments.

Rating

  • Novelty: โญโญโญโญยฝ [Elegant analytical induction connecting paired regression to continuous-time rectified flow without structural overhaul]
  • Experimental Thoroughness: โญโญโญโญโญ [Evaluated across 5 diverse vision tasks, 8+ backbones, rigorous parameterization ablations, and reference comparisons]
  • Writing Quality: โญโญโญโญโญ [Extremely lucid motivation, mathematically grounded derivations, and clean experimental exposition]
  • Value: โญโญโญโญโญ [Provides an immediate, distillation-free generative upgrade path for the vast ecosystem of existing image restoration backbones]