Co-evolving Representations in Joint Image-Feature Diffusion¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/zelaki/CoReDi
Area: Image Generation
Keywords: Joint Image-Feature Diffusion, Co-evolving Representations, Flow Matching, Feature Collapse Regularization, Pixel-Space Diffusion
TL;DR¶
Addressing the bottleneck of static, frozen semantic representation spaces in joint image-feature diffusion, this paper introduces Co-evolving Representation Diffusion (CoReDi), which jointly optimizes a lightweight learnable linear projection alongside the diffusion model; equipped with stop-gradient targets, batch normalization, and feature variance regularization to prevent collapse, CoReDi achieves up to a 13x convergence acceleration and superior sample quality in both latent and pixel spaces.
Background & Motivation¶
Diffusion models and flow matching techniques have become the dominant paradigm for high-fidelity visual synthesis, predominantly operating either within compressed VAE latent spaces or directly in pixel space to model low-level image distributions. Nevertheless, standard diffusion backbones lack explicit high-level semantic priors, leaving them vulnerable to spatial incoherence and structural distortion. The recently emerging joint image-feature diffusion paradigm (such as ReDi and REG) addresses this by concurrently modeling low-level image latents alongside high-level semantic features extracted from pretrained visual encoders (e.g., DINOv2), forcing the network to balance fine-grained texture synthesis with macro-level semantic layout.
However, existing joint diffusion formulations suffer from a fundamental disconnection: the semantic representation space that guides the generative process is constructed entirely independently of the generative objective and kept strictly fixed throughout training. In practice, high-dimensional visual representations are compressed offline via static PCA or shallow autoencoders into a low-dimensional target space, which the diffusion model is forced to passively fit. Neither the encoder projection nor the feature target receives any task-specific feedback or adaptive updates from the image synthesis loss. This asymmetric formulation—using an inflexible representation space to steer dynamic generation—severely curtails the complementary potential between high-level semantics and low-level image latents.
This paper approaches the problem with a compelling question: should the semantic space guiding diffusion remain static, or should it actively adapt to the needs of the generative model? The core idea is to introduce Co-evolving Representation Diffusion (CoReDi), replacing fixed dimensional reduction with a lightweight learnable linear projection that co-evolves with the diffusion backbone under joint flow matching, stabilized against degenerate collapse via stop-gradient targets, batch normalization, and explicit feature variance regularization.
Method¶
Overall Architecture¶
The foundational insight of CoReDi is to integrate representation projection into the generative optimization loop, forming a dual-modality joint flow matching architecture. Given an input image, a frozen pretrained visual encoder (e.g., DINOv2) extracts dense high-dimensional semantic tokens, which are immediately mapped to a lower-dimensional manifold via a learnable linear projection matrix and stabilized by batch normalization without affine parameters. The perturbed image latents (or downsampled pixel tokens) and the perturbed co-evolving representation tokens are merged into unified sequence embeddings and processed by a Diffusion Transformer (DiT) backbone to predict velocity vector fields for both modalities simultaneously. To eliminate degenerate trivial solutions and channel collapse during end-to-end training, stop-gradient is enforced on the clean representation target in the velocity loss, accompanied by an explicit feature diversity regularization term.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image x0 & Pretrained Feature z0"] --> B["Adaptive Learnable Linear Projection<br/>g_phi(z0) = z0 W_phi replacing static PCA"]
B --> C["Affine-Free Batch Normalization<br/>Stabilizing feature scale against noise schedule drift"]
C --> D["Stop-Gradient Flow Matching Velocity Target<br/>Cutting target backprop to prevent trivial zero solutions"]
D --> E["Feature Variance & Decorrelation Regularization<br/>Enforcing channel diversity to eliminate feature collapse"]
E --> F["Dual-Modality Velocity Prediction & Sampling Output"]
Key Designs¶
1. Adaptive Learnable Linear Projection: Specializing the Semantic Space for Synthesis
Conventional joint diffusion approaches (e.g., ReDi) rely on an offline, precomputed PCA matrix to project high-dimensional representations (such as DINOv2 1024-d features) down to 8 or 16 channels. However, PCA purely preserves global maximum variance directions, which often retain irrelevant background clutter or uninformative dominant patterns while dropping subtle geometric boundaries critical for generative guidance. CoReDi completely eliminates static projections by parameterizing the mapping as a learnable linear transformation \(W_\phi \in \mathbb{R}^{D \times d}\): $\(\tilde{\mathbf{z}}_0 = g_\phi(\mathbf{z}_0) = \mathbf{z}_0 W_\phi\)$ where \(D\) denotes the feature dimension of the visual encoder and \(d\) is the compact target dimensionality (\(d=8\) for latent diffusion, \(d=16\) for pixel diffusion). As gradient backpropagation updates \(W_\phi\), the projection subspace dynamically reorients itself to isolate spatial cues and structural layouts that are most complementary to image synthesis.
2. Stop-Gradient Target & Affine-Free Batch Normalization: Dual Anchors Against Numerical Collapse
Directly minimizing the joint flow matching objective end-to-end with respect to \(\phi\) induces a severe numerical pitfall: because the clean representation target \(\tilde{\mathbf{z}}_0\) depends on \(\phi\), the network can trivially minimize the velocity prediction loss by collapsing \(W_\phi\) to an all-zero matrix or a degenerate constant vector. CoReDi circumvents this trivial minimum by applying a stop-gradient operator to the clean target in the representation loss: $\(\mathcal{L}_{\text{rep}} = \left\| \mathbf{v}_\theta^z(\mathbf{x}_t, \tilde{\mathbf{z}}_t, t) - (\boldsymbol{\epsilon}_z - \text{sg}(\tilde{\mathbf{z}}_0)) \right\|^2\)$ Furthermore, continuous-time diffusion processes are exceptionally sensitive to feature scaling; unconstrained optimization of \(\phi\) causes feature variance to drift or explode, distorting the predefined noise schedule and destabilizing training. CoReDi incorporates Batch Normalization immediately following the linear projection, stripped of trainable affine parameters (\(\gamma\) and \(\beta\)). Using solely running exponential moving averages for mean and variance, it strictly centers and normalizes activations to zero mean and unit variance, perfectly harmonizing representation scales with standard Gaussian noise schedules.
3. Feature Variance & Decorrelation Regularization: Structurally Suppressing Channel Redundancy
While stop-gradient prevents scalar collapse to zero and batch normalization avoids multi-sample point collapse, the projected representation can still undergo "feature collapse", wherein all \(d\) channels mirror each other, becoming redundant or carrying zero informative variance across spatial locations. To compel each channel to encode distinct, informative generative cues, CoReDi introduces Feature Variance Regularization, penalizing spatial token positions \(i\) whose standard deviation across the channel dimension falls below a threshold \(\gamma\) via a hinge loss: $\(\mathcal{L}_{\text{var}}(\tilde{\mathbf{z}}_0) = \frac{1}{L} \sum_{i=1}^L \max\left(0, \gamma - \sqrt{\text{Var}(\tilde{\mathbf{z}}_0^i) + \epsilon}\right)\)$ where \(L\) is the number of spatial tokens and \(\gamma=1\). This ensures that individual channels remain active and non-redundant. The authors also investigate projection weight orthogonality \(\mathcal{L}_{\text{orth}} = \|W_\phi^\top W_\phi - \mathbf{I}\|_F^2\) and off-diagonal covariance regularization \(\mathcal{L}_{\text{cov}}\), proving that explicit regularization is indispensable and that feature variance regularization yields the best generative fidelity.
4. Decoupled Extension to Pixel-Space Diffusion: Overcoming the VAE Bottleneck
Beyond latent diffusion, CoReDi seamlessly extends to direct pixel-space synthesis, eliminating the inherent reconstruction bottleneck of VAE encoders. Building upon the frequency-decoupled DeCo architecture, CoReDi feeds the downsampled noisy image \(\hat{\mathbf{x}}_t\) together with the noisy co-evolving semantic tokens \(\tilde{\mathbf{z}}_t\) into a DiT encoder to produce joint conditioning features \(\mathbf{c}_{\text{joint}}\): $\(\mathbf{c}_{\text{joint}} = \text{Enc}_\theta(\hat{\mathbf{x}}_t, \tilde{\mathbf{z}}_t, t)\)$ Full-resolution pixel velocities are reconstructed by a lightweight pixel decoder \(\mathbf{v}_{\text{pred}}^x = \text{Dec}_\theta(\mathbf{x}_t, \mathbf{c}_{\text{joint}}, t)\), while semantic velocities are predicted by a linear projection head \(\mathbf{v}_{\text{pred}}^z = \mathbf{W}_{\text{dec}} \mathbf{c}_{\text{joint}}\). Because raw pixels and high-level representations reside in vastly distinct dimensional regimes, tuning the representation loss weighting to \(\lambda_z=0.1\) and allocating 16 channels delivers optimal pixel diffusion performance.
Loss & Training¶
CoReDi trains the diffusion parameters \(\theta\) and projection parameters \(\phi\) jointly via an end-to-end multi-task objective: $\(\mathcal{L}(\theta, \phi) = \mathcal{L}_{\text{image}}(\theta, \phi) + \lambda_z \mathcal{L}_{\text{rep}}(\theta, \phi) + \lambda_{\text{reg}} \mathcal{L}_{\text{reg}}(\phi)\)$ where \(\lambda_z = 1.0\) for latent diffusion, \(\lambda_z = 0.1\) for pixel diffusion, and \(\lambda_{\text{reg}} = 1.0\) by default. The projection matrix \(W_\phi\) is initialized with random orthogonal weights. For large-scale XL models trained over millions of iterations, a dedicated cosine decay learning rate scheduler optimizes the projection layer to guarantee long-term stability and convergence.
Key Experimental Results¶
Main Results¶
Evaluated on the ImageNet \(256 \times 256\) class-conditional benchmark, CoReDi demonstrates commanding advantages over prior latent and pixel diffusion models. In the unguided (w/o CFG) regime, CoReDi achieves equal or superior fidelity while dramatically reducing computational overhead.
| Model | #Params | Training Iterations / Epochs | FID ↓ | sFID ↓ | IS ↑ | Precision ↑ | Recall ↑ |
|---|---|---|---|---|---|---|---|
| SiT-B/2 | 130M | 400K iter | 33.0 | - | - | - | - |
| ReDi-B/2 | 130M | 400K iter | 21.4 | - | - | - | - |
| CoReDi-B/2 | 130M | 200K iter | 24.7 | 6.7 | 57.4 | 0.60 | 0.61 |
| CoReDi-B/2 | 130M | 400K iter | 16.4 | - | - | - | - |
| SiT-XL/2 | 675M | 7M iter | 8.3 | - | - | - | - |
| REPA-XL/2 | 675M | 4M iter | 5.9 | - | - | - | - |
| ReDi-XL/2 | 675M | 4M iter | 3.3 | - | - | - | - |
| CoReDi-XL/2 | 675M | 2M iter | 3.4 | - | - | - | - |
| DiT-XL/2 (CFG) | 675M | 1400 epochs | 2.27 | 4.60 | 278.2 | 0.83 | 0.57 |
| SiT-XL/2 (CFG) | 675M | 1400 epochs | 2.06 | 4.50 | 270.3 | 0.82 | 0.59 |
| REPA-XL/2 (CFG) | 675M | 800 epochs | 1.80 | 4.50 | 284.0 | 0.81 | 0.61 |
| ReDi-XL/2 (CFG) | 675M | 800 epochs | 1.72 | 4.68 | 278.7 | 0.77 | 0.63 |
| CoReDi-XL/2 (CFG) | 675M | 400 epochs | 1.58 | 4.33 | 297.2 | 0.63 | 0.78 |
In pixel-space diffusion experiments built upon the DeCo framework, CoReDi exhibits identical convergence acceleration and fidelity gains:
| Pixel Diffusion Model | Iterations | #Params | FID ↓ | Relative Convergence Speedup |
|---|---|---|---|---|
| DeCo-L/16 (reproduced) | 100K iter | 426M | 46.0 | 1.0× |
| DeCo-L/16 | 200K iter | 426M | 31.3 | 1.0× |
| CoReDi-L/16 | 100K iter | 426M | 31.5 | 2.0× (100K matches 200K) |
| CoReDi-L/16 | 200K iter | 426M | 21.5 | +9.8 FID improvement |
Ablation Study¶
Systematic ablations conducted on CoReDi-B/2 at 200K steps isolate the contributions of stabilization components (Stop-Gradient, Batch Normalization), regularization variants, and penalty weights.
| Configuration / Ablation | FID ↓ | sFID ↓ | IS ↑ | Precision ↑ | Recall ↑ | Note |
|---|---|---|---|---|---|---|
| CoReDi (Full model, VF) | 24.7 | 6.7 | 57.4 | 0.60 | 0.61 | Full pipeline: SG + BN + Feature Variance (\(\lambda_{\text{reg}}=1.0\)) |
| w/o Stop-Gradient (SG) | 50.8 | 8.1 | 29.5 | 0.49 | 0.53 | Target collapses; projection trivially minimizes loss (+26.1 FID) |
| w/o Batch Normalization (BN) | 223.9 | 150.6 | 3.6 | 0.06 | 0.01 | Scale explosion completely shatters noise schedule |
| w/o Regularization (-) | 37.2 | 7.2 | 39.8 | - | - | Severe feature collapse; underperforms static PCA baseline (30.9) |
| Orthogonality Reg. (Ortho) | 25.6 | 6.8 | 57.0 | - | - | Enforces weight orthonormality; successfully prevents collapse |
| Covariance Reg. (Cov) | 25.9 | 7.0 | 54.8 | - | - | Penalizes off-diagonal cross-channel covariance |
| Feature Variance (\(\lambda_{\text{reg}}=0.5\)) | 25.9 | 6.6 | 55.2 | 0.59 | 0.60 | Robust under moderate regularization weighting |
| Feature Variance (\(\lambda_{\text{reg}}=1.5\)) | 24.3 | 6.6 | 58.9 | 0.61 | 0.61 | Stronger channel diversity yields optimal generative score |
Furthermore, testing representation encoder variation across different foundation vision encoders at 200K steps confirms that CoReDi consistently outperforms ReDi's static PCA across all backbones: DINOv2 (30.9 → 24.7 FID), MOCOv3 (38.2 → 33.5 FID), SigLIPv2 (36.2 → 29.1 FID), and MAE (40.3 → 37.1 FID).
Key Findings¶
- Dual stabilization is strictly non-negotiable: Omitting the stop-gradient leads to severe target degradation (FID drops to 50.8), while eliminating batch normalization causes catastrophic failure (FID 223.9, near-zero recall), demonstrating that strict scale bounds and target isolation are vital for co-evolving diffusion systems.
- Regularization is mandatory to unlock performance: Without regularization, joint projection optimization underperforms fixed PCA (37.2 vs 30.9 FID) due to channel redundancy. Feature variance regularization achieves superior results by directly enforcing per-token spatial diversity across channels.
- Emergence of self-organized spatial structure: Tracking local distance similarity (LDS), correlogram decay slope (CDS), and root mean square contrast (RMSC) reveals that co-evolving representations autonomously develop substantially richer spatial contours and boundary definitions compared to static PCA, explaining the underlying mechanism for accelerated diffusion synthesis.
Highlights & Insights¶
- From Static Fitting to Dynamic Specialization: Overturns the prevalent dogma that auxiliary semantic representations must remain frozen read-only targets, proving that a single learnable linear layer adapting under generative gradients can unlock richer generative priors.
- Minimalist Design with Extreme Speedup: By adding only a few thousand projection parameters alongside stop-gradient, BN, and hinge variance losses, CoReDi achieves a 2x to 13x convergence speedup over SOTA baselines (converging ~13x faster than REPA and halving ReDi's required epochs).
- Universal Latent and Pixel Synergy: Demonstrates identical efficacy across both latent-based diffusion transformers and direct pixel-space generators (DeCo), confirming the fundamental applicability of co-evolving semantic guidance.
Limitations & Future Work¶
- Linear Mapping Expressivity: To preserve training stability and mathematical tractability, the mapping \(W_\phi\) is constrained to a single linear layer; non-linear adapters or cross-attention projections remain an unexplored frontier due to heightened risks of high-dimensional collapse.
- Frozen Encoder Backbone: The underlying large vision encoder (e.g., ViT-L/14) remains frozen throughout training; exploring parameter-efficient fine-tuning (e.g., LoRA) for end-to-end representation adaptation warrants investigation.
- Evaluation Scope: Benchmarks focus primarily on ImageNet \(256 \times 256\) class-conditional generation; scaling co-evolving representations to high-resolution (e.g., \(1024 \times 1024\)) and open-domain text-to-image synthesis remains future work.
Related Work & Insights¶
- vs ReDi (Kouzelis et al., 2025): ReDi introduced joint image-feature flow matching but relied on precomputed static PCA projections; CoReDi establishes adaptive co-evolution of the representation space, yielding superior spatial structure and faster convergence.
- vs REPA / iREPA (Yu et al., 2025; Singh et al., 2025): REPA aligns intermediate diffusion features with pretrained visual representations, and iREPA highlights the critical role of spatial structure; CoReDi unifies image and representation generation in an end-to-end flow matching system where spatial structure spontaneously improves during training.
- vs DeCo (Ma et al., 2025): DeCo decouples frequency components to enable efficient pixel-space diffusion without a VAE; CoReDi incorporates co-evolving semantic features into DeCo's encoder-decoder pipeline, accelerating pixel-space training by 2x.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering framework that enables semantic representation spaces to co-evolve with generative models, resolving long-standing end-to-end collapse dilemmas.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous validation across latent and pixel diffusion, multiple model scales, 4 foundation visual encoders, comprehensive ablations, and spatial self-similarity analysis.
- Writing Quality: ⭐⭐⭐⭐⭐ Exceptionally articulate and structured; the three indispensable stabilization pillars are systematically motivated and validated.
- Value: ⭐⭐⭐⭐⭐ Highly impactful paradigm shift for representation-guided generative modeling, with open-source code and immediate broad applicability.