PixelDiT2: Representation-Grounded Pixel Diffusion Transformers¶
Conference: NeurIPS2026
arXiv: 2609.24919
Area: Image Generation
Keywords: pixel-space diffusion, representation grounding, frozen vision encoder, spatial adaptive layer normalization, classifier-free guidance
TL;DR¶
PixelDiT2 re-encodes the current noisy image at each pixel-denoising evaluation, using frozen DINO patch representations and a timestep adapter as spatial conditioning without an autoencoder, achieving FID 1.46 on ImageNet-256 after 600 epochs and 1.48 on ImageNet-512 after 680 epochs.
Background & Motivation¶
Pixel-space generation avoids an autoencoder's reconstruction bottleneck but also loses the visual structure organized beforehand in a latent space. Before latent diffusion learns denoising, its encoder has already arranged semantics and some low-level details in a more compact coordinate system. A pixel model instead starts from RGB and must learn both how to represent an image internally and how to turn noise into accurate pixels. PixelDiT allocates different pathways to semantics and detail, while JiT simplifies modeling through clean-image prediction, but representation learning and generation still largely occur together inside the denoiser.
Pretrained vision models such as DINO offer an obvious source of knowledge, although using pretrained representations can mean quite different things. REPA supervises intermediate denoiser features with clean-image features and removes the teacher during sampling. RAE generates in a pretrained representation space and decodes back to images. Latent Forcing jointly generates pixel and representation streams. This paper takes another route: it neither changes the pixel-space generation variables nor generates an additional representation variable, but extracts conditioning from the current noisy image at every evaluation.
The difficulty is that DINO was trained on clean images, so its features drift substantially on highly noisy inputs. Moving all adaptation into DINO could damage its pretrained prior, while providing only a global semantic vector would discard patchwise correspondence. Core idea: preserve the frozen visual prior, assign noise-level adaptation to a timestep-conditioned projector, and keep patch representations active through spatial AdaLN throughout pixel denoising rather than using them only as training targets.
Method¶
Overall Architecture¶
The inputs are the current RGB state, diffusion timestep, and class label; the output is a clean-image estimate. The backbone is a single-path, patch-level pixel diffusion transformer, using 16ร16 patches in the main configurations. It does not treat DINO features as latent variables to generate, but uses them to modulate its own patchwise hidden states.
Every conditional denoising evaluation passes the current noisy image through frozen DINO and a timestep-conditioned projector to produce representation-grounding tokens. Spatial AdaLN injects these tokens into denoiser blocks, preserving correspondence between a representation's location and the generation location it influences. After the ODE solver updates the pixel state, the next evaluation extracts features again; encoding only the initial noise is not sufficient.
Training also retains REPA: a separate frozen teacher encodes clean images to supervise an intermediate backbone representation. Its input, role, and lifecycle differ from the grounding pathway. REPA supplies clean-image training supervision, whereas the grounding encoder actually runs in conditional sampling. The classifier-free guidance (CFG) unconditional branch uses the null-grounding state seen during training and bypasses the grounding encoder and projector.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
X["Current noisy image<br/>timestep and class"] --> G["Noise-aware representation grounding<br/>frozen DINO โ timestep projector"]
X --> S["Spatial AdaLN injection<br/>pixel denoising backbone"]
G --> S
T["Clean image โ REPA teacher<br/>training only"] -.->|intermediate feature supervision| S
S -->|conditional prediction| N["Matched null-grounding branch<br/>unconditional prediction and CFG"]
U["Timestep + null class<br/>bypass DINO and projector"] --> N
N --> O["ODE update of RGB<br/>re-encode at next evaluation"]
Key Designs¶
1. Noise-aware representation grounding: preserve the visual prior and adapt outside the encoder
Frozen DINO receives the current noisy image, not the target clean image. Its patch grid aligns with the backbone: at 256 resolution, there are 16ร16, or 256, positions; at 512 resolution, there are 32ร32, or 1024, positions. Different timesteps depart from DINO's pretraining distribution to different degrees, so a linear projection that ignores time cannot interpret features according to their noise level.
The projector first lifts DINO features to the backbone width, then applies self-attention and an MLP in one DiT-style transformer block. The timestep embedding modulates normalization through AdaLN, and a final linear layer emits grounding tokens. This allows the model to learn how to use the encoder output at a particular noise level without making the frozen encoder relearn clean-image representation recovery. The default grounding encoder is DINOv3-S/16 at 256px and DINOv3-B/16 at 512px; the projector remains trainable.
The adapter should not be interpreted as accurately recovering a target object from pure noise. The internal analysis measures mean patchwise cosine similarity for the same image across noise levels. Raw encoder features drift sharply, while projected outputs maintain similarity above 0.77 to their own clean-input outputs. This supports conditioning stability, not equivalence to true clean DINO features, and not the claim that pure noise already contains the target image's semantics.
2. Spatial AdaLN injection: modulate generation by location instead of collapsing the prior into a global label
Grounding tokens and backbone tokens share patch counts and positions, allowing each location's condition to modulate the corresponding transformer state element-wise. Unlike adding features only at the input, spatial AdaLN supplies conditioning throughout backbone blocks. Unlike token concatenation, it does not ask the denoiser to rediscover spatial correspondence from another token sequence. The controlled experiments favor this injection mechanism, but do not separately establish whether normalization, repeated injection, or spatial structure causes its advantage.
The backbone still predicts RGB directly, with neither a DINO-feature decoder nor a second jointly denoised representation trajectory. REPA additionally supervises intermediate representations using clean-image features. In the standard recipe, the REPA teacher is DINOv2-B/14; the H model aligns DiT block 8 with weight 0.5. The dedicated interaction experiment instead uses DINOv3-S/16 for both representation pathways, so its numbers must not be merged into the default model's training trajectory.
This distinction explains why retaining REPA does not make representation grounding a renaming of REPA. One changes the training signal; the other changes the computation graph of every conditional inference evaluation. The controlled experiment shows that benefits are not always additive early in training, but become complementary after longer training.
3. Matched null-grounding branch: prevent CFG's unconditional branch from retaining class information
Dropping only the class label leaves DINO representations that may still carry class-discriminative information, so the nominal unconditional branch is not a genuinely null-conditioned state. Grounding dropout therefore skips the frozen encoder and projector and broadcasts timestep and null-class embeddings into null-grounding tokens. CFG's unconditional branch uses the same construction during sampling; it is not a second forward pass that removes the label while retaining DINO conditioning.
The default recipe independently samples class dropout and grounding dropout from initialization, both with probability 0.1. Independence does not mean the two conditions are always dropped together. Different combinations expose the model to labeled/unlabeled inputs and real/null grounding. During sampling, the conditional branch retains both the class and grounding from the current image, while the unconditional branch uses the null class and null grounding; CFG then combines their velocity predictions.
The paper also tests delayed grounding dropout: train first with grounding retained and class dropout active, then enable independent dual dropout at epoch 160. This curriculum first learns to use the visual prior and subsequently learns the matched null-conditioned branch. It improves intermediate convergence at 256px, but is not the default recipe behind the 600-epoch headline result, and a single delayed trajectory cannot establish a universally optimal switching time.
A Worked Example¶
Consider one class-conditional ImageNet-256 sampling trajectory. Its initial RGB state is noise, and the main patch grid contains 256 positions. The first conditional evaluation extracts DINOv3-S/16 features at those positions, and the timestep projector turns them into conditioning appropriate to the current noise level. The backbone simultaneously reads noisy RGB patches and receives spatial AdaLN modulation to predict the clean image.
The same state also passes through the null-class/null-grounding branch, which does not run DINO or the projector. Within the guidance interval, CFG combines the two predictions, and Heun updates the RGB state. Predictor and corrector operations require additional network evaluations, whose grounding features follow their evaluation inputs. At lower noise levels, DINO inputs approach familiar natural images, but no autoencoder decoder is invoked anywhere in the trajectory.
Training instead constructs intermediate states from real images and noise, while the REPA teacher additionally reads real clean images for supervision. That clean image exists only in the training example; it must not be drawn as an external reference supplied to the grounding pathway during sampling.
Loss & Training¶
The model uses pixel-space rectified flow, with time running from noise to image rather than from image to noise:
Thus \(t=0\) is noise and \(t=1\) is the clean image. The network predicts the clean image, but its prediction error is measured in velocity space:
The denominator floor prevents numerical divergence near the clean endpoint. Outside the clipped region, this corresponds to time-dependent reweighting of clean-image prediction error. Training additionally uses REPA alignment with weight 0.5; the cache does not expand its full mathematical definition, so no exact REPA formula is reconstructed here.
Timesteps follow a logit-normal distribution with \(\mu=-0.8,\sigma=0.8\). Training uses AdamW, learning rate \(2\times10^{-4}\) held constant after a 5-epoch linear warmup, batch size 1024, weight decay 0, EMA 0.9999, gradient clipping 1.0, and bf16. The appendix specifies noise scale 2.0 at 512px versus 1.0 at 256px, a distinction that should be preserved in reproduction.
The backbone, projector, and REPA head are trained from scratch under one optimizer, while teacher encoders stay frozen. Sampling uses Heun-50. The FID-optimal default 256px configuration at epoch 600 uses CFG scale 2.4 on \([0.125,0.9]\); guidance settings are swept separately for each model and checkpoint.
Key Experimental Results¶
Main Results¶
The following representative results come from Tables 2, 3, and 14. Lower FID and higher IS are better. Architectures, pretrained assets, and epoch budgets differ, so these are not iso-FLOP or wall-clock comparisons.
| Resolution | Method | Epochs | Parameters M | FID | IS |
|---|---|---|---|---|---|
| 256 | PixelDiT-XL | 320 | 797 | 1.61 | 292.7 |
| 256 | PixelDiT-XL | 800 | 797 | 1.54 | 297.0 |
| 256 | JiT-G/16 | 600 | 2000 | 1.82 | 292.6 |
| 256 | SiD2 | 1280 | Not reported | 1.38 | Not reported |
| 256 | RAE-XL (latent space) | 800 | 839 | 1.13 | 262.6 |
| 256 | PixelDiT2-H/16 | 600 | 1008 | 1.46 | 301.6 |
| 256 | PixelDiT2-H/16 (delayed dropout) | 480 | 1008 | 1.48 | 299.1 |
| 512 | PixelDiT-XL | 850 | 797 | 1.81 | 278.6 |
| 512 | JiT-G/16 | 600 | 2000 | 1.78 | 306.8 |
| 512 | RAE-XL (latent space) | 400 | 839 | 1.13 | 259.6 |
| 512 | PixelDiT2-H/16 | 680 | 1074 | 1.48 | 295.7 |
PixelDiT2 has the lowest FID among the listed pixel-space methods at 512px. At 256px, however, SiD2's 1.38 remains better than its 1.46, so the results cannot be summarized as a universal best across all pixel diffusion methods. At 512px, epoch 200 reaches 1.78, already below PixelDiT's 850-epoch result of 1.81. The 4.25ร figure is an epoch-budget ratio, not a measured training speedup.
The reported parameter counts include the frozen grounding encoder used during sampling and exclude the training-only REPA teacher. Appendix Table 10 lists 568/2438 GFLOPs for the H model at 256px/512px; these forward-computation figures are not the complete cost of Heun-50 sampling with CFG.
Ablation Study¶
Table 4 controls the two representation pathways using the same backbone and DINOv3-S/16 for both active encoders. It tests whether training supervision and inference conditioning are complementary, rather than repeating the default DINOv2 REPA recipe's results.
| Config | FID@200 | FID@320 | FID@400 | FID@480 |
|---|---|---|---|---|
| Bare backbone | 2.53 | 2.30 | 2.13 | 2.19 |
| REPA only | 1.89 | 1.75 | 1.71 | 1.73 |
| Grounding only | 2.28 | 2.08 | 1.95 | 1.83 |
| REPA + grounding | 1.90 | 1.70 | 1.63 | 1.59 |
At epoch 200, the combined setting's 1.90 is slightly worse than REPA alone at 1.89; consistent gains begin at epoch 320. At epoch 480, the combination lowers FID by 0.14 relative to REPA alone and by 0.24 relative to grounding alone. This supports complementarity, not replacement of alignment supervision.
The following diagnostics consolidate Tables 5, 7, and 8. The first two groups use H/16 at 256px for 600 epochs; the projector group uses B/16 for 200 epochs with 143M trainable parameters. FID must not be compared across these groups as if the settings were identical.
| Diagnostic | Config | FID | Experimental setting |
|---|---|---|---|
| Injection | Spatial AdaLN | 1.462 | H/16, 600 epochs |
| Injection | Input addition | 1.521 | Same setting |
| Injection | Layer 0 token concatenation | 1.621 | Same setting |
| Injection | Layer 8 token concatenation | 1.603 | Same setting |
| Encoder adaptation | Frozen encoder | 1.462 | H/16, 600 epochs |
| Encoder adaptation | Joint timestep adapter | 1.623 | Same setting |
| Encoder adaptation | Two-stage timestep adapter | 1.581 | Same setting |
| Encoder adaptation | Joint first DINO block | 1.620 | Same setting |
| Encoder adaptation | Two-stage first DINO block | 1.625 | Same setting |
| Projector | Linear projection, 13 backbone layers | 6.79 | B/16, 200 epochs, 143M |
| Projector | Timestep DiT block, 12 backbone layers | 5.29 | Same setting |
Key Findings¶
- Adaptation placement matters beyond adding parameters: the parameter-matched projector study retains a 1.50-FID gap. Allocating the same capacity to the backbone does not replace explicit processing of noisy features.
- Larger encoders are not always better. At 256px and epoch 600, DINOv3-S/B/L obtain FID 1.462/1.547/1.552. At 512px, L beats B at epoch 200 (1.71 vs 1.78), but trails at epoch 600 (1.54 vs 1.52), ruling out a resolution-independent rule that smaller encoders always win.
- Enabling delayed dropout at epoch 160 yields 1.66/1.50/1.48 at epochs 200/320/480 at 256px, versus 1.75/1.59/1.54 when grounding is retained. This supports the curriculum and null-condition matching together, rather than isolating a single regularization mechanism.
- At NFE=100, excluding the common CFG factor, Heun-50 and Euler-100 obtain FID 1.50/1.54. At NFE=50, Heun-25/Euler-50 obtain 1.76/1.72. A higher-order solver is not better at every budget.
- A small source inconsistency remains: for delayed dropout at epoch 320, Table 14 reports IS 292.4, while Appendix B.4's Heun-50 description reports 292.2. These are not silently reconciled. Appendix C's encoder-scale experiment is actually Table 15, although its prose mistakenly cites Table 16; the table content identifies the experiment.
Highlights & Insights¶
- A visual prior need not become the generation space. A frozen encoder can observe the current pixel state and supply structure throughout inference while all generation variables remain RGB.
- Noise adaptation can be moved outside the foundation model. An explicit timestep tells a lightweight projector how unreliable current features may be, preserving the prior without writing the high-noise distribution into DINO parameters.
- CFG's unconditional branch requires checking every side channel. Removing labels does not remove all information; spatial representation conditioning also needs a matched null state, a lesson transferable to other diffusion architectures with auxiliary visual conditions.
Limitations & Future Work¶
- The authors acknowledge that the strongest latent baselines remain ahead and that sampling incurs frozen-encoder overhead. Fewer epochs do not establish lower end-to-end training cost or inference latency.
- Evaluation is limited to class-conditional ImageNet generation. Text-to-image is future work, with no demonstrated complex language conditioning, compositional concepts, or open-domain generation results.
- Stable projector outputs do not directly establish semantic correctness, and clean-input linear probes are not a complete explanation of generation. More targeted spatial interventions could test conditioning behavior at the pure-noise stage.
- Delayed dropout is evaluated under limited curricula, and larger-encoder benefits depend on resolution and training stage. Future studies could vary adaptive switching, projector capacity, and encoder evaluation frequency without assuming a universally optimal recipe.
Related Work & Insights¶
- vs REPA: REPA uses clean features as training targets and does not run its teacher during sampling. PixelDiT2 produces conditioning from the current noisy input and still retains REPA. The two pathways affect training supervision and denoising computation differently, with controlled evidence for later-stage complementarity.
- vs RAE: RAE diffuses pretrained feature representations and decodes them into images. PixelDiT2 neither generates feature latents nor uses an image decoder, but repeatedly pays encoder cost during sampling.
- vs Latent Forcing: Latent Forcing jointly generates pixels and spatial representations, whereas PixelDiT2 advances only a pixel trajectory. At the same nominal 200-epoch H/16 budget, the appendix reports FID 1.822 vs 2.287, but Latent Forcing has higher IS and each method receives its own guidance sweep.
- vs PixelDiT / JiT: PixelDiT2 inherits large-patch modeling and clean-image prediction, then supplies a frozen external prior for representations otherwise learned internally. The contribution changes how representations are used, rather than proving that pixel space inherently outperforms latent space.
Rating¶
- Novelty: 4/5 โ Frozen representations become stepwise, patchwise inference conditions, distinct from training-only alignment and joint representation generation.
- Experimental Thoroughness: 4/5 โ Two resolutions, pathway controls, parameter matching, and encoder adaptation studies are substantial, but wall-clock costs and open-domain validation are missing.
- Writing Quality: 4/5 โ Method boundaries and causal claims are reasonably cautious, although appendix references and a few IS values remain inconsistent.
- Value: 4/5 โ A reusable design for pixel generation without autoencoders, with the cost of the active visual prior still a practical constraint.