Skip to content

Unpaired Geometry-Guided Sim2Real Translation for Autonomous Driving

Conference: ECCV 2026
Paper: ECCV 2026
Area: Autonomous Driving
Keywords: Sim2Real / Diffusion Models / Autonomous Driving / Geometry Guidance / Unpaired Translation

TL;DR

To tackle the lack of paired geometric supervision and the prohibitive parameter overhead of multi-condition control in autonomous driving Sim2Real transfer, this paper proposes Cross-Domain Control Transfer (CDCT), which injects domain-agnostic geometry guidance derived from synthetic score differences into a real-domain appearance prior, and designs the Lightweight Multi-Condition Adapter (LMCA) to achieve high-fidelity unpaired translation and substantially improve downstream perception.

Background & Motivation

Simulation platforms are indispensable for autonomous driving perception models because they offer scalable data generation with perfect 3D ground truth and facilitate the rigorous evaluation of rare, safety-critical corner cases. However, synthetic imagery exhibits a pronounced visual Sim2Real gap compared to physical-world sensor data. Prior studies disentangle this discrepancy into a content gapโ€”manifesting in scene layouts and object category distributionsโ€”and an appearance gap, which stems from low-level image formation factors such as sensor noise, photometric statistics, illumination distributions, and material reflectance properties. Modern deep neural networks are acutely vulnerable to these low-level discrepancies, suffering catastrophic performance drops when models trained purely on synthetic data are deployed in the real world.

While modern diffusion models offer remarkable generative capacity to bridge the appearance gap, adapting them to autonomous driving Sim2Real translation exposes two critical dilemmas. First is the dilemma of unpaired conditional guidance: obtaining pixel-aligned, real-world geometric supervision (e.g., G-buffers) in dynamic driving environments is practically impossible, and direct application of a control module trained on synthetic geometry to the real domain fails due to severe cross-domain statistical shifts. Second is the dilemma of computational redundancy in multi-condition control: autonomous driving demands dense structural constraints spanning depth, surface normals, semantics, and materials, yet conventional multi-branch paradigms like ControlNet incur prohibitive memory and computational overhead when scaling to multiple conditions.

Prior methods relying on probability flow ODE inversion or cyclic consistency constraints frequently distort fine-grained geometric layouts, whereas forcing control networks to generalize across domains degrades real-world appearance fidelity. The key insight of this paper is that underlying 3D scene geometry is inherently domain-agnostic, meaning the synthetic domain alone can supply clean, differential geometric guidance without requiring paired real-world geometric annotations. Core Idea: decouple Sim2Real geometry control into a real-domain appearance prior and synthetic-domain geometric score differences via cross-domain score composition, and introduce a single-branch Lightweight Multi-Condition Adapter (LMCA) to achieve parameter-efficient, geometry-consistent Sim2Real synthesis without paired supervision.

Method

Overall Architecture

CDCT operates via a principled three-stage pipeline built upon a Multi-Modal Diffusion Transformer (MM-DiT, using Stable Diffusion 3 as the base backbone) optimized with the Rectified Flow velocity matching objective. Stage 1 fine-tunes the base model on unpaired synthetic and real driving images using a discrete domain indicator to establish distinct appearance priors. Stage 2 freezes the backbone and trains a lightweight control module exclusively on paired synthetic G-buffers. Stage 3 conducts inference sampling by transferring geometric control into the real appearance prior via cross-domain score composition.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    InSynth["Input: Synthetic Images & Multimodal G-buffers<br/>(BaseColor, Depth, Normal, Seg, Material)"] --> LMCA["Lightweight Multi-Condition Adapter LMCA<br/>LoRA KV Projections + Zero Linear Gating"]
    LMCA --> Stage2["Synthetic Conditional Branch<br/>Velocity Prediction v_theta(x_t, t, d_synth, g)"]
    InPrior["Input: Real and Synthetic Domain Indicators<br/>(d_real / d_synth)"] --> UnifiedBB["Unified MM-DiT Backbone<br/>Domain Embedding Projection for Appearance Priors"]
    UnifiedBB --> Stage1Real["Real Unconditional Branch<br/>Velocity Prediction v_theta(x_t, t, d_real)"]
    UnifiedBB --> Stage1Synth["Synthetic Unconditional Branch<br/>Velocity Prediction v_theta(x_t, t, d_synth)"]
    Stage2 --> ScoreComp["Cross-Domain Score Composition<br/>Inject Synthetic Geometric Guidance into Real Prior"]
    Stage1Real --> ScoreComp
    Stage1Synth --> ScoreComp
    ScoreComp --> OutReal["Output: High-Fidelity Real-Domain Images<br/>Strict Geometric Alignment & Superior Downstream Utility"]

Key Designs

1. Unified Domain-Specific Prior Learning: Disentangling Geometry-Marginalized Distributions To enable a single generative backbone to master both real-world photometric statistics and synthetic visual representations, Stage 1 establishes a shared conditional space guided by discrete domain tokens. Domain indicators \(d \in \{d_{\text{real}}, d_{\text{synth}}\}\) are mapped through a lightweight linear projection into a continuous embedding vector and broadcast across all MM-DiT attention blocks. Under the Rectified Flow framework, the model jointly optimizes the velocity prediction objective \(v_\theta(z_t, t, d)\) over unpaired synthetic and real datasets, fitting the marginal data distributions \(p(x \mid d_{\text{real}})\) and \(p(x \mid d_{\text{synth}})\). This unified parameterization avoids the computational redundancy of maintaining two independent generative models while ensuring that switching to \(d_{\text{real}}\) reliably activates the photorealistic street-view prior.

2. Cross-Domain Score Composition: Unpaired Geometry Guidance via Synthetic Gradient Differences Without real-world paired G-buffers, feeding synthetic geometric conditions directly into a real-domain conditional network triggers severe out-of-distribution failure. To overcome this limitation, the authors introduce a conditional independence assumption: given the image \(x\), its domain indicator \(d\) and its underlying geometry \(g\) are conditionally independent, i.e., \(p(d, g \mid x) = p(d \mid x) p(g \mid x)\). By Bayes' rule, the conditional score in the real domain expands as: $\(\nabla_x \log p(x \mid d_{\text{real}}, g) = \nabla_x \log p(x \mid d_{\text{real}}) + \nabla_x \log p(g \mid x)\)$ Because the pure geometry score \(\nabla_x \log p(g \mid x)\) is domain-agnostic, it can be estimated in the synthetic domain using the difference between conditional and unconditional scores: \(\nabla_x \log p(g \mid x) \approx \nabla_x \log p(x \mid d_{\text{synth}}, g) - \nabla_x \log p(x \mid d_{\text{synth}})\). During sampling, this translates directly into a Rectified Flow velocity composition rule: $\(\tilde{v}(x_t, t) = v_\theta(x_t, t, d_{\text{real}}) + \lambda \left[ v_\theta(x_t, t, d_{\text{synth}}, g) - v_\theta(x_t, t, d_{\text{synth}}) \right]\)$ where \(\lambda\) controls the guidance strength. This formulation anchors photorealistic textures using the real branch while isolating pure structural corrections from the synthetic domain, eliminating the need for paired real-world geometric supervision.

3. Lightweight Multi-Condition Adapter (LMCA): Linear-Scaling Attention Condition Injection Autonomous driving scenes require comprehensive geometric constraints comprising surface normals, metric depth, semantic segmentation, base color, and material properties. Replicating the full transformer encoder for each condition as in ControlNet is computationally infeasible, while simply concatenating condition tokens induces quadratic \(O(N^2)\) attention complexity. LMCA resolves this by attaching condition-specific Low-Rank Adaptation (LoRA) branches directly to the frozen attention projection layers. For the \(i\)-th geometry condition \(g^i\), the auxiliary key-value pairs are computed as: $\(K_g^i = (W_K + \Delta W_K^i) g^i, \quad V_g^i = (W_V + \Delta W_V^i) g^i\)$ To maintain training stability and dynamically balance multi-modal importance, a zero-initialized linear layer \(\alpha^i = \text{ZeroLinear}(g^i)\) scales only the Value features via \(V_g^i \leftarrow \alpha^i \cdot V_g^i\), leaving Key features \(K_g^i\) unscaled to preserve spatial layout representations. The query \(Q_{\text{x\&t}}\) attends across the aggregated key tokens \([K_{\text{x\&t}}, K_g^1, \dots, K_g^N]\), scaling linearly as \(O(N)\) and adding only 2.8% parameters per condition.

Loss & Training

The framework follows a progressive two-stage optimization strategy based on the Rectified Flow velocity matching loss: - Stage 1 updates all MM-DiT backbone weights with a learning rate of \(1 \times 10^{-5}\) on balanced batches of unpaired synthetic and real data; - Stage 2 freezes the backbone and trains only the condition-specific LoRA branches (rank \(r=4\)) and Zero Linear layers with a learning rate of \(1 \times 10^{-4}\), setting the domain indicator to \(d_{\text{synth}}\); - Inference sampling utilizes an Euler solver with 25 discretization steps and a guidance scale \(\lambda = 6\).

Key Experimental Results

Main Results

On the CARLA-to-Cityscapes benchmark, CDCT outperforms prior unpaired translation methods across controllability, generative quality, and cross-domain similarity metrics.

Method Controllability SSIM โ†‘ Controllability LPIPS โ†“ Generative Quality NIQE โ†“ Generative Quality MANIQA โ†‘ Similarity FID โ†“ Similarity KID (\(\times 10^3\)) โ†“
EPE (PAMI 2022) 0.82 0.26 4.19 0.32 34.78 25.23
VSAIT (ECCV 2022) 0.89 0.18 3.26 0.41 29.45 27.67
HIE (arXiv 2023) 0.84 0.26 4.87 0.33 32.76 42.05
UNSB (ICLR 2024) 0.73 0.34 5.44 0.37 41.62 39.46
CACTIF (CVIU 2025) 0.69 0.39 6.63 0.31 50.91 36.06
CDCT (Ours) 0.85 0.17 3.21 0.44 26.45 22.17

In zero-shot downstream perception benchmarks, perception networks (Mask2Former for semantic segmentation, YOLOv8 for object detection, NeWCRFs for monocular depth estimation) trained solely on translated data and evaluated on the Cityscapes validation set yield substantial performance improvements:

Training Set Semantic Seg. mIoU โ†‘ Object Det. mAP50 โ†‘ Depth AbsRel โ†“ Depth SILog โ†“ Depth \(\delta < 1.25\) โ†‘
Raw CARLA 14.6 20.2 0.56 44.18 0.29
EPE 36.9 33.9 0.20 25.04 0.73
VSAIT 33.8 35.7 0.19 21.15 0.78
HIE 30.5 31.8 0.21 28.93 0.68
UNSB 26.4 28.6 0.25 39.11 0.51
CACTIF 32.1 31.7 0.23 18.06 0.60
CDCT (Ours) 48.3 40.5 0.16 19.24 0.81
Cityscapes (Real Bound) 65.2 48.8 0.11 14.92 0.87

Furthermore, joint data augmentation experiments confirm that adding CDCT translated data to Cityscapes raises YOLOv8 detection from 48.8 to 52.5 mAP50 (+3.7) and Mask2Former segmentation from 65.2 to 67.0 mIoU (+1.8), consistently outperforming the real-data-only baseline.

Ablation Study

The ablation experiments isolate the impact of cross-domain score composition and evaluate the architectural efficiency of LMCA compared to standard ControlNet paradigms.

Ablation Config / Architecture Multi-Cond. Support Additional Params FID โ†“ KID (\(\times 10^3\)) โ†“ Empirical Mechanism
w/o Score Comp. โœ” 57M \(\times N\) 63.77 57.11 Direct real-domain conditional inference overfits G-buffers and collapses real appearance priors
Full Model (w/ Score Comp.) โœ” 57M \(\times N\) 26.45 22.17 Real appearance prior and synthetic geometric guidance are orthogonally integrated
Standard ControlNet (SD3) โœ˜ 1.1B (+55%) - - Duplicating transformer encoder branches causes catastrophic GPU memory exhaustion
LMCA (Ours) โœ” 57M (+2.8%) \(\times N\) - - Cross-condition LoRA KV projections scale attention memory strictly linearly

Key Findings

  • Cross-domain score composition is the indispensable mechanism for unpaired control: removing score composition deteriorates the FID score from 26.45 to 63.77, as the network prioritizes raw G-buffer features and loses real-world distribution alignment.
  • Geometric guidance substantially mitigates cross-domain semantic degradation: CDCT boosts Mask2Former zero-shot transfer from 14.6 mIoU (raw CARLA) to 48.3 mIoU, achieving huge gains on safety-critical classes such as road (83.9 vs 22.1), car (79.5 vs 36.1), and pedestrian (38.4 vs 1.6).
  • Synthetic data effectively expands real-world training diversity: joint training with CDCT data surpasses models trained solely on real Cityscapes data, validating that translated synthetic samples provide valuable complementary geometric layouts and visual diversity.

Highlights & Insights

  • Principled Score-Space Transfer: avoids brittle adversarial objectives or cyclical loss constraints by proving that conditional independence enables clean extraction of geometric guidance from synthetic score differences.
  • Scalable Multi-Condition Conditioning: LMCA integrates multiple G-buffer modalities directly into multi-modal attention key-value projections, modulating only the Value path with zero-initialized gating and scaling linearly at only 2.8% extra parameters per condition.
  • Thorough Perception Verification: validates the framework across image quality metrics (FID, KID, NIQE, MANIQA), structural metrics (SSIM, LPIPS), and three downstream autonomous driving perception tasks (segmentation, detection, depth estimation).

Limitations & Future Work

  • Inherent Content Gap: CDCT bridges the appearance gap by generating photorealistic textures, but discrepancies in geographic traffic layouts, regional road signs, and object frequencies between CARLA and Cityscapes persist.
  • Video Temporal Consistency: the framework currently operates on static frames; extending it to multi-view temporal sequences requires introducing cross-frame attention or optical flow priors to suppress temporal flicker.
  • Future directions include scaling CDCT to surround-view multi-camera consistency and incorporating 3D bounding box layout controls for full-stack autonomous driving simulation pipelines.
  • vs EPE (PAMI 2022): EPE relies on convolutional refinement with intermediate rendering buffers, which often produces high-frequency artifacts; CDCT leverages an MM-DiT diffusion backbone to synthesize photorealistic textures with superior FID and KID.
  • vs VSAIT (ECCV 2022): VSAIT enforces vector symbolic structural binding but fails to capture specular highlights and fine material properties; CDCT explicitly injects normal and material G-buffers, yielding higher visual realism and downstream accuracy.
  • vs CACTIF (CVIU 2025): CACTIF explores diffusion style transfer without dense geometric guidance, leading to structural distortions in complex traffic scenes; CDCT maintains strict spatial layouts, outperforming CACTIF by a large margin on segmentation (48.3 vs 32.1 mIoU) and detection (40.5 vs 31.7 mAP50).

Rating

  • Novelty: โญโญโญโญโญ An elegant cross-domain score composition formulation paired with the highly efficient LMCA module, systematically resolving unpaired multi-condition diffusion Sim2Real transfer.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive validation covering image generation metrics, zero-shot perception benchmarks across three modalities, and joint data augmentation experiments.
  • Writing Quality: โญโญโญโญโญ Clear mathematical derivations, cohesive architectural diagrams, and rigorous ablation analyses.
  • Value: โญโญโญโญโญ Delivers an efficient, high-performance data generation paradigm for closing the Sim2Real loop in autonomous driving.