Skip to content

Controllable Generative Reference for Stereo Image Compression via Reliability-Aware Gating

Conference: ECCV 2026
Paper: ECCV 2026
Area: 3D Vision
Keywords: Stereo Image Compression, Diffusion Models, Generative Prior, Reliability-Aware Gating, Entropy Coding

TL;DR

Addressing the vulnerabilities of traditional stereo image compression in out-of-view (OOV) regions and the propagation of heavy quantization artifacts, this work leverages a pre-trained latent diffusion model steered by joint semantic-spatial control to synthesize a high-quality target-view reference, adaptively regulated via reliability-aware gating within the entropy model to achieve state-of-the-art rate-distortion performance.

Background & Motivation

Stereo image compression is foundational for binocular vision systems in safety-critical autonomous driving, robotic navigation, and immersive 3D telepresence. The core principle lies in eliminating inter-view redundancy across binocular pairs to maximize rate-distortion efficiency. Over the past decade, learned methodologies have progressively evolved from explicit geometric alignment via disparity-based warping or optical flow to implicit cross-view interactions using bidirectional attention and masked image modeling. Across all these frameworks, cross-view context prediction fundamentally depends on establishing explicit correspondences from a compressed reference view.

However, under stringent bandwidth constraints and low-bitrate regimes, this correspondence-driven paradigm encounters two insurmountable bottlenecks. First, the compressed reference view itself suffers from aggressive quantization, resulting in heavily degraded latent representations; forcing feature alignment onto these degraded representations inevitably propagates quantization noise and blurring artifacts to the target view. Second, due to the physical stereo baseline, certain portions of the target view are completely occluded or absent in the reference view—a phenomenon known as out-of-view (OOV) regions. Because feature warping and cross-attention rely strictly on finding co-visible correspondences, these non-overlapping areas suffer from total contextual starvation, causing severe prediction collapse. While unconstrained generative models such as GANs or diffusion models can hallucinate realistic details, their stochastic nature introduces structural inconsistency that severely penalizes objective rate-distortion metrics like PSNR and MS-SSIM.

This paper tackles this fundamental conflict by re-conceptualizing inter-view prediction: rather than struggling to extract explicit correspondences from degraded references, it reframes target-view prediction as a controllable generative completion task while deploying an adaptive feature gate to isolate structural inaccuracies. Core idea: leverage a pre-trained diffusion prior steered by a geometric anchor from the compressed left view and an ultra-compact semantic condition from the target view to synthesize a high-quality reference, then dynamically filter structural inconsistencies using a reliability-aware gating module to strictly anchor rate-distortion optimization.

Method

Overall Architecture

The framework operates as an asymmetric three-stage compression pipeline. In Stage I (Anchor View Compression), the left view \(x_L\) is compressed using a standard learned image compression (LIC) network to reconstruct \(\hat{x}_L\), serving as both the final left-view output and a robust geometric anchor. In Stage II (Dual-Control Conditional Generation), a frozen Latent Diffusion Model (LDM) executes single-step latent inference steered by Joint Semantic-Spatial Control, synthesizing a high-quality target-view reference \(x_{gen}\) equipped with complete OOV structures and restored high-frequency textures. In Stage III (Target View Compression with Gated Prior), features from the geometric anchor \(\hat{x}_L\) and the generative reference \(x_{gen}\) are extracted into a hybrid context; a Reliability-Aware Gating module dynamically assesses the feature-space trustworthiness of \(x_{gen}\) to refine the Gaussian entropy model parameters, enabling rate-distortion optimal coding and reconstruction of the target view \(\hat{x}_R\).

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    subgraph S1["Stage I: Anchor View Compression"]
        direction TB
        XL["Left View Input xL"] --> L_Enc["Standard LIC Codec"]
        L_Enc --> XL_Hat["Reconstructed Anchor x̂L"]
    end

    subgraph S2["Stage II: Dual-Control Conditional Generation"]
        direction TB
        XR["Right View Input xR"] --> Sem_Enc["Lightweight Semantic Codec<br/>(~0.00065 bpp)"]
        Sem_Enc --> F_Sem["Semantic Condition f̂sem"]
        XL_Hat --> CN_Spa["Spatial Control Network CNimg"]
        F_Sem --> CN_Lat["Latent Control Network CNlat"]
        CN_Spa & CN_Lat --> Fuse["Element-wise Feature Fusion"]
        Fuse --> LDM["Frozen LDM U-Net<br/>(Single-step Latent Inference)"]
        LDM --> VAE_Dec["VAE Decoder"]
        VAE_Dec --> X_Gen["Synthesized Reference xgen"]
    end

    subgraph S3["Stage III: Target View Compression with Gated Prior"]
        direction TB
        X_Gen & XL_Hat --> Feat_Ext["Shared Feature Extractor grefa"]
        Feat_Ext --> Gating["Reliability-Aware Gating Module<br/>(Spatial Confidence Map G)"]
        Gating --> Entropy["Conditional Gaussian Entropy Model<br/>(Refined μFinal and σFinal)"]
        XR --> R_Enc["LIC Encoder & Quantization ŷR"]
        R_Enc & Entropy --> R_Dec["Entropy Decoder & LIC Decoder"]
        R_Dec --> XR_Hat["Reconstructed Target x̂R"]
    end

    XL_Hat -.-> CN_Spa
    XL_Hat -.-> Feat_Ext
    X_Gen -.-> Feat_Ext

Key Designs

1. Spatial Control via Geometric Anchor: Constraining Co-visible Baseline Geometry Rather than directly transferring heavily quantized reference features that induce severe blurring artifacts, the framework uses the reconstructed left view \(\hat{x}_L\) strictly as a spatial conditioning field for diffusion. Processed through an image-domain control network \(CN_{img}\) alongside diffusion timestep \(t\) and noisy latent \(z_t\), it produces multi-scale spatial guidance features: $\(F_{ctl}^{spa} = \text{CN}_{img}(\hat{x}_L, z_t, t)\)$ This feature map supplies deterministic spatial boundaries, edges, and co-visible geometric constraints, ensuring that the single-step generative reference generation maintains rigid stereoscopic alignment and prevents macro-structural hallucinations.

2. Semantic Control via Ultra-Compact Condition: Guiding OOV Inpainting at Negligible Overhead Because out-of-view (OOV) regions and heavily occluded zones in the target view have zero physical correspondence in the left anchor, purely spatial guidance cannot resolve cross-view ambiguities. The framework extracts an ultra-compact semantic condition \(\hat{f}_{sem}\) from the target view \(x_R\) using a lightweight semantic codec operating on the VAE latent space, requiring an average bitrate of merely \(\sim 0.00065\text{ bpp}\). A lightweight Latent Adapter aligns feature dimensions and restores spatial context for a modified latent control network \(CN_{lat}\): $\(F_{ctl}^{sem} = \text{CN}_{lat}(\text{Adapter}(\hat{f}_{sem}), z_t, t)\)$ The spatial and semantic control features are summed element-wise and injected into the frozen LDM decoder blocks via zero-convolutions. This minimal semantic seed acts as an anchor for target disparity and macro-layout, prompting the diffusion prior to accurately inpaint missing OOV regions with high perceptual fidelity to produce \(x_{gen} = \text{VAE}_{dec}(z_{gen})\).

3. Reliability-Aware Gating: Mitigating Structural Inconsistency to Preserve Objective Fidelity While diffusion priors generate highly realistic textures, their stochastic nature can introduce local pixel-level discrepancies relative to the ground truth, causing severe penalties in PSNR and MS-SSIM if naively incorporated into the entropy model. To resolve this tension, a Reliability-Aware Gating module is built into the Gaussian entropy model. A shared extractor computes geometric features \(F_{geo} = g_a^{ref}(\hat{x}_L)\) and generative features \(F_{gen} = g_a^{ref}(x_{gen})\). A residual geometric branch first produces a baseline mean offset \(\Delta\mu_{geo}\), while a generative branch computes a richer but riskier candidate offset \(\Delta\mu_{gen}\). A lightweight gating subnetwork dynamically estimates a spatial confidence map \(G \in [0, 1]\): $\(G = \sigma(\mathcal{F}_{Gate}([\mu_{Base}, F_{gen}]))\)$ The final refined mean parameter is computed via adaptive gating: $\(\mu_{Final} = \mu_{Base} + \Delta\mu_{geo} + (G \odot \Delta\mu_{gen})\)$ When generative features exhibit uncertainty or structural misalignment, \(G\) approaches zero, seamlessly suppressing the generative path and falling back to the reliable geometric context. The scale parameter \(\sigma_{Final}\) is refined jointly across all contextual branches, anchoring strict objective rate-distortion behavior.

Loss & Training

The framework adopts a decoupled three-stage training strategy to guarantee convergence and prevent gradient interference: - Stage I (Anchor View Optimization): The left-view LIC codec is trained using standard rate-distortion loss \(\mathcal{L}_{\text{Stage I}} = \lambda D(x_L, \hat{x}_L) + R(\hat{y}_L) + R(\hat{z}_L)\), where \(D(\cdot)\) denotes MSE distortion. - Stage II (Joint Generative & Semantic Optimization): With the pre-trained LDM and VAE strictly frozen, the dual control networks and semantic codec are optimized jointly: $\(\mathcal{L}_{\text{Stage II}} = \mathcal{L}_{Diff} + \lambda_{codec} \mathcal{L}_{Codec}\)$ $\(\mathcal{L}_{Codec} = R(\hat{f}_{sem}) + \beta \|f_{sem} - \hat{f}_{sem}\|_2^2\)$ Gradients from the diffusion loss \(\mathcal{L}_{Diff}\) backpropagate through \(CN_{lat}\) directly into the semantic codec, compelling the ultra-compact representation \(\hat{f}_{sem}\) to prioritize crucial macro-layout cues. - Stage III (Target View Compression with Gated Prior): Preceding stages are frozen while the right-view LIC codec and gating module are trained on the target rate-distortion objective. To ensure the gating module robustly detects generative failure, a targeted corruption scheme is applied during training: \(F_{gen}\) is randomly dropped with probability \(p_{Drop}\) or corrupted with Gaussian noise with probability \(p_{Noise}\), forcing the gate to identify uninformative features and shut down unreliable information flow.

Key Experimental Results

Main Results

The proposed method was comprehensively evaluated against standard traditional codecs (BPG, HEVC, VVC, MV-HEVC) and recent state-of-the-art learned stereo compression models (LDMIC, ECSIC, BiSIC, CAMSIC) on InStereo2K and Cityscapes. Table 1 reports Bjøntegaard Delta Bitrate (BDBR) savings relative to the BPG anchor (lower / more negative indicates superior performance):

Dataset Metric Ours Prev. SOTA Gain
InStereo2K PSNR BDBR (↓) -67.64% -48.08% (BiSIC) -19.56% bitrate savings
InStereo2K MS-SSIM BDBR (↓) -65.22% -61.13% (BiSIC) -4.09% bitrate savings
Cityscapes PSNR BDBR (↓) -65.90% -59.74% (CAMSIC) -6.16% bitrate savings
Cityscapes MS-SSIM BDBR (↓) -74.53% -69.67% (CAMSIC) -4.86% bitrate savings

In terms of inference runtime measured on an NVIDIA A100 GPU for learned models and an AMD EPYC 7742 CPU for traditional codecs, the total per-pair runtime is 1.60 seconds (0.78s for the core compression pipeline and 0.82s for single-step reference synthesis). This remains on par with state-of-the-art learned codecs (ECSIC: 1.47s, BiSIC: 1.29s, CAMSIC: 1.34s) while being over two orders of magnitude faster than VVC (208.42s).

Ablation Study

Ablation experiments conducted on InStereo2K and Cityscapes quantify the contribution of each core component in terms of Bjøntegaard Delta PSNR (BD-PSNR) relative to the full model (negative values indicate performance degradation):

Config InStereo2K BD-PSNR Cityscapes BD-PSNR Note
Full Model (Ours) 0.00 dB 0.00 dB full model
w/o Generative Reference -1.31 dB -1.23 dB Removing generative prior causes massive performance collapse
w/o Reliability-Aware Gate -1.07 dB -0.91 dB Ungated integration of generative features incurs severe objective penalty
w/o Semantic Control -0.76 dB -0.83 dB Removing semantic condition leads to spatial ambiguity and misalignment

Key Findings

  • Generative prior is the primary performance driver: Removing the synthesized reference degrades BD-PSNR by 1.31 dB on InStereo2K and 1.23 dB on Cityscapes, validating that replacing degraded correspondence matching with generative synthesis fundamentally elevates coding efficiency.
  • Reliability-aware gating acts as an essential objective stabilizer: Incorporating generative features without the dynamic gate leads to a 1.07 dB drop on InStereo2K, demonstrating that unconstrained generative priors introduce structural inconsistencies that damage objective rate-distortion metrics unless strictly regulated.
  • Pronounced low-bitrate dominance: On InStereo2K at approximately 0.07 bpp, the model attains 35.75 dB PSNR, outperforming BiSIC and CAMSIC by over 1.0 dB. On Cityscapes at roughly 0.03 bpp, it surpasses ECSIC by 1.5 dB in PSNR.

Highlights & Insights

  • From Correspondence Search to Controllable Generation: Moves beyond a decade-long reliance on explicit optical flow, disparity warping, and cross-attention matching in stereo compression, reframing target-view prediction as a conditional generative completion task that inherently overcomes out-of-view regions.
  • Micro-Bitrate Semantic Seeding: Demonstrates that an ultra-compact semantic representation consuming merely \(\sim 0.00065\text{ bpp}\), optimized via end-to-end diffusion loss gradients, provides sufficient macro-layout and disparity guidance to deterministically steer a foundation diffusion model.
  • Harmonizing Perception and Distortion via Gating: Elegant resolution of the perception-distortion tradeoff in multi-view compression: stochastic generative synthesis provides rich perceptual detail, while the feature-level reliability gate dynamically suppresses hallucinations to safeguard objective rate-distortion performance.

Limitations & Future Work

  • Computational Footprint and Memory Requirements: Although reference synthesis is accelerated via single-step latent inference (0.82s), hosting a pre-trained Stable Diffusion U-Net requires notable GPU VRAM and compute, presenting deployment challenges for low-power embedded stereo systems.
  • Diminishing Returns at High Bitrates: Under abundant bandwidth where fine textures can be faithfully reconstructed directly via quantized transform coefficients, the marginal benefit of hallucinated generative priors naturally contracts.
  • Future Directions: Exploring distilled lightweight diffusion backbones (e.g., consistency models or compact student networks) to reduce parameter and compute overhead, and extending the dual-control gated paradigm to multi-view video compression.
  • vs BiSIC [Liu et al., ECCV 2024] / CAMSIC [Zhang et al., AAAI 2025]: Both BiSIC (cross-dimensional 3D convolution entropy model) and CAMSIC (masked image modeling transformer) rely on extracting valid context from heavily quantized reference features, degrading severely when the left view is aggressively compressed or in OOV regions. The proposed method bypasses degraded correspondence search by synthesizing clean target-view references via diffusion priors.
  • vs Generative Codecs [Mentzer et al., NeurIPS 2020 / Agustsson et al., CVPR 2023]: Unconstrained generative image compression models optimize primarily for perceptual metrics and suffer significant PSNR drops. In contrast, this paper integrates stereo geometric anchors with reliability-aware gating, successfully harnessing generative richness while adhering strictly to objective rate-distortion optimization.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering paradigm shifting stereo reference prediction to controlled diffusion generation with dual semantic-spatial conditioning and reliability-aware gating]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Thorough evaluation across two standard benchmarks against 9 traditional and deep learned baselines, comprehensive BDBR, ablations, and runtime benchmarks]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Exemplary technical clarity, well-structured multi-stage formulation, and precise architectural schematics]
  • Value: ⭐⭐⭐⭐⭐ [Sets a new benchmark for learned stereo compression and offers a compelling template for integrating generative priors into bandwidth-constrained multi-view vision tasks]