StereoEdit: A Diffusion-Based Framework for Stereo-Consistent Image Editing¶
Conference: ECCV 2026
Paper: CVF Open Access
Area: 3D Vision
Keywords: Stereo Image Editing, Stereo Consistency, Diffusion Models, Epipolar Attention, Latent Alignment
TL;DR¶
Addressing cross-view texture drift and geometric distortions in independent 2D diffusion edits, StereoEdit introduces a training-free framework that incorporates Geometrically-Gated Attention Reference (GGAR) at the feature level and Geometry-Aware Latent Alignment (GALA) at the latent level to enforce rigid stereo consistency without retraining.
Background & Motivation¶
The rapid proliferation of head-mounted displays (such as AR/VR headsets) and dual-lens mobile devices has renewed interest in stereoscopic 3D content consumption. However, generating and manipulating high-quality stereoscopic imagery remains a formidable challenge. Traditional stereoscopic editing techniques predominantly rely on handcrafted geometric warping, salient low-distortion retargeting, or constrained neural style transfer. These pipelines lack semantic flexibility and cannot accommodate open-vocabulary, natural-language-guided content creation. Meanwhile, modern text-to-image diffusion models such as Stable Diffusion demonstrate exceptional 2D editing capabilities, but their generative priors operate strictly within an image-centric latent space absent explicit multi-view geometric grounding. Naively applying monocular diffusion models independently to the left and right views causes independent noise evolutions to diverge, precipitating severe binocular texture mismatches, epipolar violations, and visual fatigue during 3D viewing.
To enforce cross-view coherence, a plausible workaround is adapting text-guided video diffusion editing frameworks (e.g., TokenFlow, VidToMe). Yet, a fundamental physical gap separates temporal video sequences from stereoscopic image pairs. Video editing frameworks model smooth temporal dynamics captured by a moving viewpoint over time, optimizing for progressive inter-frame continuity. In sharp contrast, stereoscopic pairs represent simultaneous planar projections of the exact same 3D scene from dual viewpoints along a static horizontal baseline, governed strictly by rigid epipolar geometry and exact pixel-level parallax correspondences. Enforcing temporal attention priors across stereo views overlooks horizontal epipolar lines, introducing out-of-plane texture drift and geometric artifacts. Furthermore, explicit 3D editing methods built on NeRF or 3D Gaussian Splatting (3DGS) rely heavily on multi-view camera pose initialization (e.g., via COLMAP); under the extreme sparsity of small-baseline two-view stereo inputs, such pipelines frequently suffer from pose failure, depth collapse, and geometry degradation.
Consequently, the central dilemma of stereoscopic image editing lies in harnessing the semantic manipulation power of pretrained 2D diffusion models while strictly enforcing rigid epipolar constraints and disparity reliability across denoising steps, without incurring costly retraining or fragile 3D reconstructions. Core idea: incorporate pre-estimated bidirectional disparity priors and geometric reliability confidence directly into the diffusion inference pipeline, gating out occluded and non-epipolar features in U-Net self-attention while adaptively correcting drifting latents via asymmetric bidirectional warping errors.
Method¶
Overall Architecture¶
StereoEdit operates on a paired stereoscopic input \((I_L, I_R)\) and a target text editing prompt within a unified, training-free, geometry-aware diffusion framework. As illustrated in the pipeline, the system executes across three coordinated stages: geometric prior estimation, shared DDIM inversion, and joint geometry-constrained denoising. First, a pre-trained stereo disparity estimator (such as MASt3R) extracts bidirectional disparity maps \((D_{R \to L}, D_{L \to R})\) along with cyclic geometric reliability maps \((R_{R \to L}, R_{L \to R})\). Second, both views are mapped into the shared latent space at timestep \(T\) (\(z_T\)) via DDIM inversion. Throughout reverse diffusion from \(t = T\) down to \(0\), the Geometrically-Gated Attention Reference (GGAR) module enforces epipolar constraints and suppresses unreliable occlusions within the U-Net self-attention layers. Concurrently, after each denoising step, the Geometry-Aware Latent Alignment (GALA) module measures cross-view warping discrepancies and adaptively corrects latent drift, collectively securing local texture fidelity and global binocular consistency.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Stereo Input Pair<br/>(I_L, I_R) & Text Prompt"] --> B["Geometric Prior Estimation<br/>Bidirectional Disparities & Reliability Maps"]
B --> C["DDIM Latent Inversion<br/>Initial Latent States z_T"]
C --> D["Geometrically-Gated Attention Reference (GGAR)<br/>Disparity Warping + Epipolar Mask + Log Gating"]
D --> E["U-Net Step Denoising<br/>Coarse Latents (z_L,t-1, z_R,t-1)"]
E --> F["Geometry-Aware Latent Alignment (GALA)<br/>Asymmetric Error Estimation + Confidence-Weighted Refinement"]
F -->|Iterative Denoising t=T..0| D
F --> G["Stereo-Consistent Output Pair<br/>(I'_L, I'_R)"]
Key Designs¶
1. Geometrically-Gated Attention Reference (GGAR): Epipolar Soft Banding and Logarithmic Reliability Gating Standard self-attention mechanisms and unconstrained cross-view attention permit query tokens to attend globally across all key/value locations of the opposing view, which easily contaminates non-epipolar regions and allows occluded features to bleed across views. GGAR explicitly binds feature interactions to valid epipolar geometries. Formally, for the left-view branch (with an identical symmetric operation applied from right to left), the key and value projections \((K_R, V_R)\) are warped into the left coordinate frame via the estimated disparity map \(D_{R \to L}\), yielding \((K_{R \to L}, V_{R \to L})\). To account for slight disparity estimation errors while suppressing off-epipolar noise, a 2D epipolar Gaussian mask matrix \(M \in \mathbb{R}^{N \times N}\) is formulated based on the vertical row distance between token \(i\) and warped token \(j\): \(M_{i,j} = \exp\left(-\frac{(\text{row}_i - \text{row}_j)^2}{2\sigma^2}\right)\). Furthermore, utilizing the cyclic forward-backward disparity consistency map \(R_{R \to L}\), GGAR modulates the cross-attention affinity matrix:
Through this logarithmic scaling, highly reliable stereo co-visible regions (\(R \approx 1\)) preserve full attention transmission without penalization. Conversely, occluded areas, boundary discontinuities, or estimation discrepancies (\(R \to 0\)) produce an unbounded negative term, driving cross-view attention weights to zero after Softmax normalization. This ensures that cross-view attention strictly references reliable, epipolarly compliant features.
2. Geometry-Aware Latent Alignment (GALA): Asymmetric Error Estimation and Dynamic Latent Refinement While GGAR maintains localized feature correspondence inside U-Net attention layers, global structural disparities can still accumulate across multi-step denoising trajectories. GALA acts as an external post-step geometric regularizer on the predicted latents \((z_{L,t-1}, z_{R,t-1})\). At each timestep, bidirectional reconstruction absolute errors are computed via warping: \(E_{L \to R} = | \text{Warp}(z_{L,t-1}, D_{L \to R}) - z_{R,t-1} |\) and \(E_{R \to L} = | \text{Warp}(z_{R,t-1}, D_{R \to L}) - z_{L,t-1} |\). An observed discrepancy \(E_{L \to R} > E_{R \to L}\) reveals that projecting left into right incurs higher inconsistency, signifying that the right-view latent has drifted further from geometric truth. To address this imbalance without over-smoothing, GALA assigns an asymmetric dynamic strength weight:
where \(\sigma\) denotes the Sigmoid function, and \(k\) and \(\beta_{\text{base}}\) control sensitivity and baseline correction scale. The latent update is then weighted by the reliability map:
An identical symmetric correction is applied to \(z'_{L,t-1}\). This asymmetric pulling forces the more deviant view toward the reliable view's projection while leaving the more consistent view essentially intact, steering the dual latents toward rigid binocular alignment over 50 denoising iterations.
3. Mask-Guided Content Editing: Spatial Disentanglement for Local Edits When users desire localized semantic modifications on specific foreground objects while holding the surrounding scene invariant, StereoEdit integrates a mask-conditioned inpainting pipeline. Spatial binary masks \((M_L, M_R)\) are generated across dual views using Segment Anything Model (SAM) or user bounding boxes. During reverse diffusion, unmasked background regions are precisely replaced by the inverted source latents, while GGAR and GALA concentrate geometric filtering and alignment exclusively within the edited regions and their transition boundaries. This spatial disentanglement prevents unwanted background alterations and suppresses spurious artifacts in complex regions.
Key Experimental Results¶
Main Results¶
The framework was evaluated on 120 stereo image pairs compiled from NBU-SIRQA, Flickr1024, Middlebury 2014, and synthetic benchmarks, covering portraits, wildlife, and natural landscapes. Quantitative evaluation combines three automatic metrics: CLIP-T (text-image semantic fidelity), CLIP-F (cross-view cosine similarity of CLIP embeddings), and Warp-Err (pixel-level reprojection error between the warped view and target view using ground-truth disparities). Given the absence of prior open-source stereoscopic diffusion editing models, five state-of-the-art text-guided video diffusion editing methods (TokenFlow, VidToMe, VideoGrain, VideoDirector, ReAtCo) served as baselines. In addition, a user study with 25 expert participants assessed 50 test pairs on Edit Accuracy (Edit-Acc), Stereo Consistency (Stereo-Con), and Overall Quality (Overall) on a 1-10 scale.
| Method | CLIP-T โ | CLIP-F โ | Warp-Err โ | Edit-Acc โ (Human) | Stereo-Con โ (Human) | Overall โ (Human) |
|---|---|---|---|---|---|---|
| TokenFlow | 31.87 | 96.14 | 5.53 | 6.52 | 7.31 | 6.19 |
| VidToMe | 32.14 | 96.38 | 6.88 | 6.30 | 7.64 | 6.64 |
| VideoGrain | 33.77 | 94.62 | 6.51 | 7.62 | 7.06 | 7.18 |
| VideoDirector | 33.08 | 96.09 | 8.07 | 7.08 | 6.96 | 6.89 |
| ReAtCo | 33.72 | 95.04 | 7.20 | 6.99 | 6.76 | 6.39 |
| StereoEdit (Ours) | 33.82 | 97.36 | 3.92 | 7.65 | 8.01 | 8.09 |
Ablation Study¶
The ablation experiments, built upon Stable Diffusion v1.5 (Base), isolate the performance gains from GGAR, the logarithmic reliability factor \(R\), GALA, and the video editing cross-attention baseline (WAttn):
| Config | Base | GGAR w/o R | GGAR | GALA | WAttn | CLIP-F โ | CLIP-T โ | Warp-Err โ |
|---|---|---|---|---|---|---|---|---|
| Vanilla SD without stereo adaptation | โ | 94.97 | 32.29 | 8.78 | ||||
| GGAR without reliability gating (\(R\)) | โ | โ | 97.19 | 33.19 | 4.55 | |||
| Full GGAR only | โ | โ | 97.23 | 33.19 | 4.41 | |||
| GALA only | โ | โ | 95.97 | 32.84 | 5.15 | |||
| GALA + Video Cross-Attention (WAttn) | โ | โ | โ | 96.31 | 33.25 | 5.04 | ||
| Full Model | โ | โ | โ | 97.36 | 33.82 | 3.92 |
Key Findings¶
- 29.11% reduction in Warp-Err versus the best baseline: StereoEdit cuts reprojection error from 5.53 (TokenFlow) down to 3.92, verifying that explicit epipolar gating and disparity-weighted latent updates eliminate binocular disparity mismatch.
- Complementarity between GGAR and GALA: Operating GGAR alone achieves a 4.41 Warp-Err but lacks global trajectory realignment; deploying GALA alone achieves a 5.15 Warp-Err with a lower CLIP-F (95.97). Their integration produces strong synergy across both local textures and global stereo structures.
- Video attention fails on stereo geometry: Replacing GGAR with temporal cross-attention (WAttn) degrades Warp-Err to 5.04 and lowers cross-view coherence, corroborating that temporal smoothing cannot substitute for strict horizontal epipolar geometry.
Highlights & Insights¶
- Logarithmic geometric reliability gating: Fusing disparity warping with epipolar Gaussian masks confines cross-view attention within horizontal epipolar bands, while the logarithmic confidence term penalizes occluded boundaries smoothly, avoiding boundary tearing.
- Asymmetric error-driven latent correction: Rather than imposing symmetrical feature blending, GALA diagnoses the more distorted view via bidirectional projection residuals and directs stronger correction toward the drifting latent, preserving generated content stability.
- Direct extensibility to stereoscopic video editing: Embedding GGAR and GALA into video diffusion architectures with concatenated stereo channels preserves inter-frame temporal continuity while maintaining binocular stereo correspondence, providing a foundation for future immersive 3D video manipulation.
Limitations & Future Work¶
- Dependency on front-end disparity estimation quality: StereoEdit relies on accurate disparity and occlusion maps generated by upstream estimators like MASt3R. Extreme specular highlights, transparent surfaces, or severe depth discontinuities can induce disparity noise that translates into subtle geometric artifacts.
- Limitation to rectified horizontal baselines: The current formulation assumes rectified stereo pairs with strictly horizontal epipolar lines. Handling non-rectified binocular setups, significant vertical disparity, or camera rotation would require generalizing the banded Gaussian mask to arbitrary fundamental matrix formulations.
Related Work & Insights¶
- vs Video Editing Methods (TokenFlow, VidToMe): Temporal video editors prioritize inter-frame continuity along an unconstrained motion trajectory, lacking horizontal baseline constraints; StereoEdit enforces epipolar line masking and disparity alignment, preventing artifacts like mismatched cup handles or broken petal contours.
- vs Explicit 3D Editing (3DGS, NeRF): Explicit 3D scene editing methods (e.g., GaussianEditor, Instruct-NeRF2NeRF) require dense multi-view initialization and accurate camera poses; StereoEdit operates directly on small-baseline dual-view pairs in a training-free 2D diffusion framework, showing superior robustness when multi-view reconstruction fails.
Rating¶
- Novelty: โญโญโญโญ [Pioneering training-free diffusion framework combining epipolar-gated attention with asymmetric latent alignment for stereo editing]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive quantitative benchmarks, professional human evaluations, thorough module ablations, and extension to stereoscopic video]
- Writing Quality: โญโญโญโญโญ [Clear problem formulation, elegant mathematical formulations tightly coupled with diagrams, and cohesive narrative flow]
- Value: โญโญโญโญโญ [Provides a practical zero-shot solution for stereoscopic content manipulation across emerging AR/VR headsets and 3D displays]