Skip to content

Revisiting Diffusion Fine-Tuning for Unsupervised Domain Adaptation

Conference: NeurIPS2026
arXiv: 2609.33716
Paper: Project page
Area: Unsupervised Domain Adaptation / Diffusion Models
Keywords: multi-target domain adaptation, shared-private adapters, gradient decoupling, bridge data, parameter-efficient fine-tuning

TL;DR

MUSE separates diffusion-generator adaptation into a shared semantic branch trained only with source-label supervision and a target-conditioned style branch, allowing one source-involved fine-tuning process to generate bridge data for multiple targets; it improves average downstream UDA accuracy on three closed-set classification benchmarks while reducing generator fine-tuning time to roughly one-half to two-fifths of Terra's cost.

Background & Motivation

Unsupervised domain adaptation (UDA) provides labeled source images and unlabeled target images, but no target training labels. Conventional methods help classifiers transfer through feature alignment, adversarial training, or pseudo-labeling; diffusion-based methods instead make training data more target-like before supplying labeled synthetic images to downstream adaptation algorithms. The challenge is not merely to generate convincing target-style images: generated samples must also preserve their classes, or their assigned training labels become noise.

Methods such as Terra and DCDM primarily construct a generation process for one sourceโ€“target pair. When one labeled source must serve several targets, pairwise adaptation repeatedly learns the same source-class knowledge and maintains separate generator-adaptation parameters for each target. Simply mixing all targets is also problematic: appearance updates can interfere across targets, while target reconstruction without class-specific prompts can overwrite shared class information. The paper therefore studies multi-target data generation, rather than directly proposing one unified classifier for all targets.

Source supervision is reusable, but target appearance still requires specialization. The paper expresses this division through parameter sharing and gradient routing, without assuming that semantics and style are strictly independent in real images. Core idea: encode source-class knowledge in a shared semantic branch that accepts only source gradients, let targets share style projections while maintaining small private cores, and generate labeled bridge data through the corresponding target branch to reuse source learning without erasing target differences.

Method

Overall Architecture

The inputs are one labeled source and multiple unlabeled targets with a shared class set. This is a closed-set assumption, not target-data-free domain generalization. MUSE freezes SDXL's original U-Net, VAE, and both text encoders, inserting shared-private adapters only into the query, key, value, and output attention projections of the U-Net. Each source undergoes one joint multi-target fine-tuning process; afterward, selecting a target index produces that target's bridge data.

Four connected designs define the pipeline: shared-private adapters determine parameter reuse; two-pass gradient decoupling protects source-class information; adaptive target sampling allocates target-update budgets; and dual-route bridge generation converts the adapted generator into downstream classification supervision. Each target still uses its own unlabeled images and bridge set to train MCC, ELS, or SSRT, rather than obtaining class predictions directly from the diffusion denoiser.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Labeled source<br/>Multiple unlabeled targets"] --> B["Shared-private adapters"]
    B --> C["Two-pass gradient decoupling"]
    C -->|Training: target-loss EMA| D["Adaptive target sampling"]
    D -.->|Training: select next target| C
    C -->|Fine-tuning complete; select target core for inference| E["Dual-route bridge generation"]
    E -->|Source inversion or noise generation; inherit labels| F["Target-specific bridge set"]
    A -.->|Downstream training: unlabeled target images| G["MCC / ELS / SSRT"]
    F -->|Downstream training: synthetic-label supervision| G

Key Designs

1. Shared-private adapters: share large projections and distinguish targets through small cores

Each attention projection adds two branches to its frozen weights. The semantic branch is a conventional low-rank update with one parameter set reused by every target. The style branch does not assign a complete LoRA to every target: all targets share its input and output projections, with a target-private small square matrix between them. The adapted weight for target \(m\) is:

\[ \widetilde W_{\ell}^{(m)}=W_{0,\ell}+U_{\ell,\mathrm{sem}}V_{\ell,\mathrm{sem}}+U_{\ell,\mathrm{sty}}\Omega_{\ell}^{(m)}V_{\ell,\mathrm{sty}}. \]

The semantic rank is 32 and the style rank is 64, so each target adds a \(64\times64\) private core per layer rather than another complete set of low-rank projections. Source conditioning disables the entire style term; target conditioning activates only the corresponding core, with other target cores excluded from that sample's denoising forward pass. This avoids fully mixing targets while allowing them to reuse parameters in a common style subspace.

Appendix A uses a specific parameter-count reference: independent per-target adaptation with the same adapter form. With \(L\) adapted layers and \(S=\sum_\ell(d_\ell^{\mathrm{in}}+d_\ell^{\mathrm{out}})\), that reference requires \(M[(r_{\mathrm{sem}}+r_{\mathrm{sty}})S+Lr_{\mathrm{sty}}^2]\) trainable parameters, whereas MUSE requires \((r_{\mathrm{sem}}+r_{\mathrm{sty}})S+MLr_{\mathrm{sty}}^2\). Adding one target adds only \(Lr_{\mathrm{sty}}^2\) core parameters. This counts adapters, not the entire SDXL model, optimizer states, activation memory, or full-pipeline storage, and is not a universal superiority proof over all PEFT methods.

2. Two-pass gradient decoupling: target forward passes use semantics without updating them

Each optimization iteration contains one source denoising pass and one target denoising pass. Source images use the class prompt a y with target-private cores disabled, and their denoising loss updates only the shared semantic factors. Target images have no labels and use the generic prompt an image with the selected target core enabled. The semantic branch still contributes to the forward prediction but is stop-gradient, preventing target reconstruction from directly overwriting source-trained class parameters.

Target back-propagation updates shared style projections and private cores. Gradients accumulate on disjoint parameter subsets before a single optimizer update. This does not alternate source and target updates over the same parameters; nor is the semantic branch frozen throughout training, since it continuously receives source supervision. Shared style projections still receive updates from different targets, so routing removes direct sourceโ€“target gradient interference on semantic parameters, not all cross-target coupling.

The target objective also regularizes every private core for diversity and magnitude. A core can therefore receive regularization gradients even when its target is not sampled:

\[ \mathcal L_{\mathrm{orth}}=\sum_{\ell}\sum_{i<j}\left\|(\Omega_{\ell}^{(i)})^\top\Omega_{\ell}^{(j)}\right\|_F^2,\qquad \mathcal L_{\mathrm{core}}=\sum_{\ell}\sum_m\|\Omega_{\ell}^{(m)}\|_F^2. \]

The first term penalizes correlation between target-core coefficients; the second limits core magnitude. Shared style projections are not constrained to be orthonormal, so orthogonality regularization in core-coefficient space does not enforce exact orthogonality of the induced full weight updates. Likewise, a small core norm is not a hard constraint on the norm of the complete style update.

3. Adaptive target sampling: increase effective loss weights for targets that are harder to fit

A target pass reads a minibatch from one target rather than evaluating all target branches simultaneously. Sampling begins uniformly. Each iteration updates only the sampled target's exponential moving average of denoising loss, leaving other targets' statistics unchanged. A larger recent loss indicates that the current generator finds that domain harder to reconstruct, so it receives more subsequent target updates. Sampling probabilities are:

\[ p_m=\frac{(\bar\ell_m+\tau)^\alpha}{\sum_{k=1}^M(\bar\ell_k+\tau)^\alpha}. \]

The implementation uses an EMA coefficient of 0.9, \(\tau=10^{-6}\), and \(\alpha=1\), recomputing probabilities every 20 optimization steps. Difficulty here means denoising reconstruction difficulty, not target classification error: target labels are unavailable for directly measuring classification difficulty.

Crucially, there is no \(1/p_m\) importance correction. Conditional on the current sampling distribution, the expected target denoising term is \(\sum_m p_m\lambda_m\mathcal L_{\mathrm{sty}}^{(m)}\), not the fixed \(\sum_m\lambda_m\mathcal L_{\mathrm{sty}}^{(m)}\) written in main-text Eq. (6); the implementation sets every \(\lambda_m=1\). Adaptive sampling therefore changes both visitation frequency and effective domain weights. Appendix C explicitly explains this distinction: it should not be described as an unbiased acceleration of a fixed uniform objective. Core regularization still covers all targets at every iteration.

4. Dual-route bridge generation: instance structure and class diversity provide different supervision

After fine-tuning, semantic and style projections remain shared and generation switches only the target core. The first route performs DDIM inversion on a labeled source image and denoises it with the target branch, producing a target-style image that inherits the source label. Concrete instance structure helps limit class drift. The second route starts from Gaussian noise and uses the class prompt a y with the same target branch; the prompt supplies the label, and 50 samples per class per target add within-class variation beyond the source instances.

The two sets are concatenated into a bridge set. The paper's union symbol means dataset concatenation, not image blending, latent addition, or deduplication. The downstream UDA learner receives labeled synthetic bridge data together with real unlabeled target data. Synthetic labels are inherited from source labels or prompts, rather than indicating human verification of every generated image. Source inversion and noise generation are complementary, but label preservation still depends on the diffusion prior and domain shift.

A Worked Example

Consider miniDomainNet with Real as the source and Clipart, Painting, and Sketch as three targets. A source photograph labeled lion uses a lion in the source pass, updating only the semantic branch. If Sketch is sampled, a batch of unlabeled Sketch images uses an image in the target pass. Semantic parameters contribute to the forward pass without receiving gradients; the Sketch core and shared style projections receive denoising gradients, and all cores receive regularization gradients.

After training, DDIM inversion of the lion source image followed by denoising with the Sketch core produces a bridge image inheriting the lion label. Noise-based generation then produces 50 Sketch-style samples per class. Both sets are concatenated and used with real unlabeled Sketch images to train the downstream classifier. Switching to Clipart or Painting reuses the generator's shared components and changes the corresponding core, but still requires target-specific data generation and downstream training. The saving is repeated source-involved generator fine-tuning, not these later stages.

Loss & Training

Both denoising passes operate in the frozen VAE's latent space, using squared prediction error with timestep-dependent weighting. The pretrained scheduler determines the denoising target; in standard noise prediction it is the injected noise. The source objective contains only semantic denoising loss. The target objective combines the sampled target's denoising loss with core regularizers weighted by \(\lambda_{\mathrm{orth}}=10^{-4}\) and \(\lambda_{\mathrm{core}}=10^{-5}\).

Images have resolution \(1024\times1024\). Training runs for 10,000 steps with per-device batch size 8 and gradient accumulation 1; every step contains one source batch and one target batch. Optimization uses 8-bit AdamW and a constant learning rate without warmup. Semantic factors and shared style projections use \(5\times10^{-5}\), while private cores use \(5\times10^{-3}\), a 100-fold difference. Weight decay is \(10^{-4}\) and the maximum gradient-clipping norm is 1.0.

The VAE runs in fp32, while the U-Net uses fp16 mixed precision and gradient checkpointing. Appendix D specifies one A100 80GB and 16 CPU cores. Validation and final sampling use a DPMSolver multistep scheduler, with 25 denoising steps by default for qualitative samples. This qualitative setting should not automatically be treated as a complete step specification for every inversion operation.

Key Experimental Results

Main Results

Office-31 has 3 domains and 6 ordered transfers; Office-Home and miniDomainNet each use 4 domains and 12 transfers. The metric is top-1 accuracy (%) on real target images, and the table preserves the source tables' reported task averages. MCC and ELS are CNN-based, whereas SSRT is transformer-based. Comparing the same downstream learner better isolates generation strategy; differences between backbones cannot be attributed to MUSE.

Dataset Downstream learner No diffusion Terra pairwise generation DCDM pairwise generation MUSE multi-target generation
Office-31 MCC 89.61 91.34 91.01 91.77
Office-31 ELS 90.21 91.28 91.80 92.36
Office-31 SSRT 93.50 93.73 95.29 95.53
Office-Home MCC 72.24 75.45 73.71 76.53
Office-Home ELS 71.84 76.26 73.64 77.57
Office-Home SSRT 85.43 85.58 86.54 88.15
miniDomainNet MCC 61.35 61.57 62.98 64.02
miniDomainNet ELS 60.09 61.16 63.60 65.04
miniDomainNet SSRT 80.78 81.31 81.26 84.37

Source: main-text Tables 1โ€“3. Diffusion methods use a common synthetic-supervision protocol concatenating 50 class-conditioned samples per class and source-to-target inversion images. MUSE has the highest average in all nine datasetโ€“learner combinations, but does not win every transfer. For example, Office-Home Arโ†’Cl with SSRT is 79.09 for MUSE and 79.34 for DCDM.

Fine-tuning efficiency compares only Terra and MUSE, both of which adapt pretrained diffusion foundation models. DCDM does not fine-tune a pretrained foundation model such as SDXL and is excluded from this time and adapter-parameter comparison.

Dataset Targets per source Terra fine-tuning time (h) MUSE fine-tuning time (h) Speedup Terra / MUSE per-process peak VRAM (GiB)
Office-31 2 96.78 51.74 1.87ร— 32.433 / 27.849
Office-Home 3 196.36 80.79 2.43ร— 32.433 / 27.835
miniDomainNet 3 197.68 81.04 2.44ร— 32.433 / 27.835

Source: main-text Table 4 and Appendix A Table 5. Times are cumulative generator fine-tuning costs for the complete benchmark protocol, not one source run, and exclude bridge generation and downstream training. VRAM is the peak memory of PyTorch-allocated tensors per training process. Summed memory for concurrent independent jobs is not Terra's per-process memory; sequential Terra still peaks at 32.433 GiB.

Ablation Study

All results below are average accuracy (%) over the 12 Office-Home tasks; no ablation results from other datasets are mixed in.

Config MCC ELS SSRT Note
Full MUSE 76.53 77.57 88.15 Both bridge routes, gradient decoupling, adaptive sampling, and core regularization
No core orthogonality regularizer 75.22 75.58 87.04 Core magnitude regularization retained
Target pass also updates semantic branch 74.56 74.98 87.12 Target-side stop-gradient removed
Uniform target sampling 75.74 76.78 86.94 Other settings unchanged
Inversion bridge images only 73.20 74.14 85.73 No noise-based class-conditioned images
Class-conditioned images only 73.16 73.97 85.16 No inversion bridge images

Source: Appendix F Tables 9โ€“11 and 13. Removing one bridge-data route causes the largest drops: ELS loses 3.43 and 3.60 percentage points with inversion-only and class-conditioned-only data, respectively. Updating semantic parameters on target batches costs 2.59 points, removing orthogonality regularization costs 1.99 points, and uniform sampling costs 0.79 points. Concatenating both routes also increases sample count, so this ablation alone cannot establish that all gains arise from complementary sample types rather than data volume.

Key Findings

  • SSRT+MUSE improves over SSRT+Terra by 1.80, 2.57, and 3.06 percentage points across the three benchmarks. This comparison is stronger on miniDomainNet, but does not establish that every learner and reference has its largest gain there.
  • In Appendix F's prompt-only analysis, ELS+SDXL with class prompts averages 69.92%, detailed target-style prompts reach 73.23%, and confidence filtering reaches 73.99%, still below ELS+MUSE at 77.57%. These controls do not fine-tune the generator with sourceโ€“target data; this does not imply that their downstream learner avoids target data.
  • Inversion generation takes 4.23 seconds per image and noise-based class-conditioned generation takes 2.08 seconds (Table 12). Adapter and fine-tuning savings do not remove generation latency or establish a 2.44ร— end-to-end training speedup.
  • Appendix E reports per-task standard deviations across three independent runs, not confidence intervals for average accuracy. A cross-table consistency issue remains unresolved: Office-31 SSRT+MUSE reports 100.00% for both Dโ†’W and Wโ†’D, but corresponding standard deviations are 0.29 and 0.62; some SSRT+DCDM entries at 100.00% also have nonzero deviations. The original values are retained rather than corrected. An exact mean of 100% cannot coexist with these deviations, and whether rounding or reporting is responsible requires author clarification.

Highlights & Insights

  • The sharing boundary is defined by supervision provenance, not just matrix structure. Source inputs carry class prompts whereas target inputs use a generic prompt, and only the former can modify semantic parameters, giving reusable class knowledge a concrete training interpretation.
  • Sharing large style projections with small private cores lowers the incremental parameter cost per target. A transferable idea is to identify a common cross-domain variation subspace and modulate it with low-dimensional domain coefficients, while validating whether that shared subspace is sufficient for the task.
  • Adaptive sampling is dynamic reweighting, not a neutral acceleration trick. Reusing it should involve reporting target probabilities, target quality, and the relative regularization scale, rather than equating denoising difficulty with classification difficulty.

Limitations & Future Work

  • Evaluation covers only closed-set image classification. Unknown classes, missing classes, and shifted class priors can bias supervision generated from source labels; open-set or partial-set extensions need transferable-class identification rather than unchanged use of all class prompts.
  • Operational separation of source-supervised semantics and target appearance is not a proof of strict semanticโ€“style disentanglement. Viewpoint shifts, occlusion, strong abstraction, or class-dependent styles can make the style branch alter discriminative structure, and artifacts or semantic drift remain possible.
  • Joint multi-target generation is evaluated with only 2 or 3 targets per source. The incremental parameter formula does not establish that new targets require no joint retraining, or that quality and time remain stable with many targets or sequential target arrivals.
  • High-resolution SDXL adaptation, inversion, and per-target downstream training remain expensive. Future audits should separate fine-tuning, generation, storage, and classifier-training costs instead of measuring deployment cost solely through adapter size.
  • Qualitative images and t-SNE provide supporting interpretations, but this task's text cache contains only captions and descriptions, so actual visual quality cannot be verified here. Equal-sample-count bridge controls, generated-label error rates, and further statistical testing would be useful.
  • vs Terra: Terra uses time-varying low-rank adaptation for sourceโ€“target pairs; MUSE shares source-supervised semantics and style projections across targets. Its advantage concerns multi-target reuse, not a claim that every single-target task requires MUSE.
  • vs DCDM: Both improve downstream UDA through target-relevant synthetic supervision, but their generator-adaptation forms differ. Accuracy can be compared under matching downstream protocols; SDXL PEFT fine-tuning times cannot be mechanically applied to DCDM.
  • vs MCC / ELS / SSRT: These methods perform discriminative downstream adaptation. MUSE changes their synthetic supervision data, acting as a composable data-side complement rather than replacing the classifier.
  • Research direction: Shared style ranks could adapt to structural differences among targets, and sampling budgets could reflect label-preservation quality rather than denoising loss alone. This is a direction proposed in this note, not a result validated by the paper.

Rating

  • Novelty: 4/5 โ€” Implements multi-target reuse through shared projections, private cores, and supervision-specific gradient routing.
  • Experimental Thoroughness: 4/5 โ€” Three benchmarks, three learners, and major-component ablations are substantial, but target counts and cost/data-volume controls remain limited.
  • Writing Quality: 4/5 โ€” Appendices clarify effective sampling weights and orthogonality boundaries; perfect accuracies and their cross-table standard deviations need clarification.
  • Value: 4/5 โ€” Provides a reusable generator for multi-target diffusion UDA, but is neither low-cost adaptation nor adaptation without target data.