Skip to content

D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Amplitude and Pixel Spaces

Conference: NeurIPS2026
arXiv: 2511.11286
Code: https://github.com/RapidsAtHKUST/D-GAP
Area: Others (Out-of-Domain Robustness / Domain Adaptation)
Keywords: out-of-domain robustness, gradient guidance, amplitude mixing, dual-space augmentation, unsupervised domain adaptation

Version note: This note follows arXiv v4 dated 2026-09-30. The earlier list used โ€œFrequency and Pixel Spaces,โ€ whereas v4 uses โ€œAmplitude and Pixel Spacesโ€; this is a title update under the same arXiv ID, and the existing file path is retained.

TL;DR

D-GAP uses task-loss gradients with respect to Fourier amplitudes to determine frequency-wise cross-domain mixing strengths, then fuses the result with pixel mixing, improving over the respective best generic methods by an average of 5.3 percentage points on four real-world datasets; however, not using target labels does not mean not accessing target data during training.

Background & Motivation

Camera locations, histopathology staining, recording devices, and telescopes can change input appearance statistics without changing the categories a model must recognize. When background or acquisition features correlate with classes in the training domain, models can exploit these readily learned cues and fail after a domain shift. Generic augmentations such as RandAugment and MixUp do not explicitly address these shifts; Copy-Paste and Stain Color Jitter are more targeted but require foreground segmentation, staining knowledge, or prior analysis. The problem is not simply a shortage of training examples, but whether augmentation perturbs domain-related cues or destroys information needed for recognition.

Fourier augmentation offers an alternative interface: change the amplitude spectrum while preserving the source phase that carries spatial structure, modifying appearance while attempting to retain content. FACT, FDA, and related methods already use this idea, but fixed frequency bands or random mixing ratios do not identify the frequencies on which the current model actually relies. The same frequency can play different roles across tasks, datasets, and training stages, so uniform perturbation may not effectively disrupt the current model's spectral dependence. Furthermore, amplitude-only mixing can blur reconstructions or introduce artifacts, removing local recognition details.

D-GAP therefore uses the model's own task gradients as augmentation feedback rather than manually deciding which frequency bands represent background or style. Here, dataset-agnostic means that a handcrafted augmentation rule is not required for each dataset; it does not eliminate the need for cross-domain examples, hyperparameters, or validation data. Its primary setting is domain adaptation, and access to target images must be stated separately for each experiment. Core idea: use amplitude-gradient sensitivity to regulate cross-domain amplitude interpolation, then supplement it with a pixel branch, making augmentation responsive to the current model rather than permanently tied to fixed spectral rules.

Method

Overall Architecture

The training inputs are a labeled source image and another cross-domain image; the latter supplies input information without requiring its class label as a mixing target. D-GAP first computes amplitude sensitivity through the supervised task on the source example and converts it into a bounded mixing map. The amplitude branch interpolates the two amplitude spectra according to this map and reconstructs an image with the source phase; the pixel branch directly blends the original images. The two augmented results are then fused again to train the classifier.

The โ€œcross-domain imageโ€ need not come from the eventual deployment domain: iWildCam, Camelyon17, and BirdCalls sample from other training domains, whereas Galaxy10 and the three common benchmarks allow unlabeled target-domain images. Thus, the cross-domain input in the diagram is a protocol-dependent data source, not a supervision branch containing test labels. Sensitivity computation and augmentation take place during training; inference directly uses the trained classifier without executing this dual-branch augmentation pipeline.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    S["Source image + source label"] --> G["Amplitude Sensitivity Guidance"]
    G --> A["Phase-Preserving<br/>Amplitude Mixing"]
    S --> P["Pixel Compensation<br/>and Dual-Space Fusion"]
    T["Cross-domain image<br/>its label is not used for mixing"] --> A
    T --> P
    A --> P
    P --> L["Task training with source labels"]
    L --> M["Trained classifier<br/>direct prediction at inference"]

Key Designs

1. Amplitude Sensitivity Guidance: let the current task response determine perturbation strength

The method applies the Fourier transform to each RGB channel, separating amplitude and phase. Amplitude describes the strength of frequency components, while phase preserves source spatial structure; this is an empirical division of roles for augmentation, not a strict decomposition in which amplitude contains only style and phase contains only class information. Within the square amplitude-mixing window \(\Omega_r\), it computes the absolute gradient of the source task loss with respect to each amplitude location. A large gradient means that the current model's loss is locally sensitive to perturbing that location; it does not directly prove that the frequency represents a spurious correlation.

\[ G(u,v)=\left|\frac{\partial\mathcal{L}_{\mathrm{task}}(f(x_1),y)}{\partial A(x_1)(u,v)}\right|,\qquad (u,v)\in\Omega_r. \]

The sensitivity map is standardized using its mean and standard deviation, passed through a sigmoid, and clipped to a predefined interval. This produces a continuous frequency-wise mixing map rather than a single random scalar for the entire window. Clipping here constrains mixing ratios to an interval; it is not multi-step adversarial projected-gradient optimization of the input. The inspected text does not specify numerical values for \(d_{\min}\) and \(d_{\max}\), so this note does not invent defaults.

\[ \widetilde G=\frac{G-\mu(G)}{\sigma(G)+\varepsilon},\qquad D=\operatorname{clip}\bigl(\operatorname{Sigmoid}(\widetilde G),d_{\min},d_{\max}\bigr). \]

This mapping injects more cross-domain amplitude at relatively high-sensitivity locations and preserves more source amplitude at low-sensitivity locations. The motivation is to perturb spectral components on which the model currently relies strongly, requiring it to accommodate cross-domain changes during training. However, sensitive components can also contain task-relevant information; this is an experimentally supported sensitivity heuristic, not a causal criterion that automatically identifies domain bias.

2. Phase-Preserving Amplitude Mixing: vary appearance statistics while attempting to retain source structure

Given the mixing map, each frequency inside the window interpolates between source and cross-domain amplitudes. Unlike replacing an entire low-frequency region, replacement strength is continuous, location-dependent, and responsive to the model's task gradients. The window limits the intervention region; \(r\) denotes its side length relative to image dimensions, not its area ratio.

\[ A_{\mathrm{mix}}(u,v)=\bigl(1-D(u,v)\bigr)A(x_1)(u,v)+D(u,v)A(x_2)(u,v),\qquad (u,v)\in\Omega_r. \]

The mixed amplitude is combined with the source phase and inverse-transformed into the frequency-augmented image \(\hat x_f\). The source phase is not mixed with target phase; this is the branch's main constraint for retaining source-label semantics. The cross-domain image need not belong to the same class: it supplies appearance statistics, and the method does not use its label for same-class pairing. This enables the use of unlabeled target images, while retaining the risk that mixing different-class content changes semantics.

Preserving phase can improve the likelihood of structural fidelity but cannot guarantee label consistency for every task. The paper shows examples of blurring and artifacts from frequency mixing and consequently does not treat this branch alone as the complete solution. This explains why pixel information is still needed, rather than merely adding a regularizer unrelated to the spectrum.

3. Pixel Compensation and Dual-Space Fusion: avoid making the spectral reconstruction the only training view

The pixel branch blends the two original images using one weight, then fuses its result with the frequency-augmented image using a second weight. The first weight controls the cross-domain image's influence within the pixel branch; the second controls the final image's reliance on that branch. They serve different roles and should not be collapsed into a single โ€œtarget-domain injection ratio.โ€

\[ \hat x_p=(1-\lambda_1)x_1+\lambda_1x_2,\qquad \hat x=(1-\lambda_2)\hat x_f+\lambda_2\hat x_p. \]

The pixel branch provides complementary information from the original pixels, so the final view does not depend entirely on Fourier reconstruction. It is not segmentation-based foreground pasting and does not explicitly detect or restore blurred edges; โ€œdetail compensationโ€ is the design motivation and empirical interpretation. Pixel blending can also introduce content from the other image, so its effectiveness depends on mixing strength rather than inherently guaranteeing unchanged semantics. In the ablation, pixel-only mixing is substantially worse than no augmentation on Camelyon17, showing that dual-space fusion is not simply the addition of arbitrary augmentations.

A Worked Example

For Galaxy10, the source input is a labeled DECaLS image, and the cross-domain input is randomly sampled from the unlabeled SDSS test set. Amplitude gradients are first computed from the DECaLS image and its source label; the SDSS class label does not participate in sensitivity computation. The default \(r=0.5\) defines the mixing window, and the sensitivity map assigns different interpolation strengths to different frequencies within it.

The DECaLS phase is retained to reconstruct the frequency-augmented image, while a pixel-blended DECaLSโ€“SDSS image is generated in parallel. Finally, \(\lambda_2\) combines them into a view used for source-label training. This example explains the paper's sampling protocol, not a numerical trajectory reported for a particular image. It illustrates target-input access in unsupervised domain adaptation, not pure domain generalization with no exposure to SDSS data.

Loss & Training

The method uses the supervised task loss to generate sensitivity feedback; augmentation does not depend on target labels and does not require an additional loss explicitly identifying spurious features. Real-world datasets use linear probing followed by fine-tuning (LP-FT): the pretrained encoder is initially frozen while training the classifier, after which encoder and classifier are fine-tuned jointly. This stabilizes the classifier before adapting representations to augmented views, avoiding strong changes to pretrained features at the outset.

The appendix specifies 10 epochs of linear probing for all four real-world datasets; fine-tuning lasts 20 epochs for iWildCam and Camelyon17, 30 for BirdCalls, and 15 for Galaxy10. Galaxy10 fine-tuning uses Adam, learning rate \(10^{-4}\), batch size 8, and augmentation probability 0.5. PACS, Office-Home, and Digits-DG omit LP-FT and train directly under the comparison methods' settings. Appendix B recommends \(r=0.5\), \(\lambda_1\in[0.2,0.6]\), and \(\lambda_2\in[0.2,0.6]\); these are recommended ranges, not disclosure of one identical exact weight pair used in every experiment.

Key Experimental Results

Main Results

Real-world results below come from the Section 4.2 textual discussion of Figure 5; benchmark averages come from Tables 2 and 3. Values are percentages and gains are percentage points; F1 and accuracy across different datasets are not one unified metric. The comparison method is the best generic method identified for that dataset or the listed benchmark method, not a globally optimal method across the field.

Dataset Metric Comparison Method Comparison Value D-GAP Gain (percentage points)
iWildCam OOD Macro F1 FACT 34.7 36.8 +2.1
Camelyon17 OOD accuracy SAM 92.2 96.4 +4.2
BirdCalls OOD Macro F1 FACT 35.1 40.7 +5.6
Galaxy10 OOD accuracy SAM 74.1 83.4 +9.3
PACS Average accuracy FACT 87.88 89.03 +1.15
Office-Home Average accuracy Su et al. 67.71 70.22 +2.51
Digits-DG Average accuracy Su et al. 82.6 84.5 +1.9

The mean gain is 5.3 percentage points across the four real-world datasets and approximately 1.9 percentage points across the three benchmarks. PACS and Office-Home use ResNet18, and Digits-DG uses ConvNet; the real-world datasets respectively use ResNet50, DenseNet121, EfficientNet-B0, and ResNet18.

Protocol boundaries must accompany these numbers: Appendix E.1 states that mixing targets for iWildCam, Camelyon17, and BirdCalls come from other training domains. Galaxy10 uses mixing targets from the unlabeled test set; training on the three common benchmarks permits unlabeled images from the held-out target domain. Thus, even though the tables are titled leave-one-domain-out, those three benchmark results cannot be directly interpreted as pure DG with no target-input access during training.

The main text and appendix state that some FACT, SAM, and other baseline values are taken from source papers; Appendix A provides no standard deviations for DeepAll, FACT, or SAM, while other relevant results report five random seeds. This is not a significance test in which all methods are rerun under fully unified conditions and identical target-access permissions.

Ablation Study

The following table selects the OOD columns from Table 4; all values are percentages. Frequency-only uses only the frequency branch, while Mask-low, Mask-high, and Mask-ring replace gradient guidance with fixed masks.

Config iWildCam F1 Camelyon17 Accuracy BirdCalls F1 Galaxy10 Accuracy
No augmentation 31.2 91.4 30.1 54.3
Pixel-only 29.1 69.8 32.4 68.2
Frequency-only 35.6 95.7 37.6 77.9
Mask-low 36.0 96.1 38.4 80.5
Mask-high 35.4 95.4 36.4 79.5
Mask-ring 34.3 94.0 36.5 77.1
Full D-GAP 36.8 96.4 40.7 83.4

The full method improves over Frequency-only by 1.2, 0.7, 3.1, and 5.5 percentage points, respectively. Its gains over Mask-low are 0.8, 0.3, 2.3, and 2.9 percentage points, supporting adaptive mixing over this fixed low-frequency variant. However, these are differences between complete training pipelines; they do not prove that the pixel branch restores a specific edge or that gradients accurately identify spurious features.

Table 6 also compares the Jaccard similarity of high-gradient location sets, defined as intersection size divided by union size. SC-CD denotes same-class cross-domain pairs, and DC-SD denotes different-class same-domain pairs; the following table retains all three thresholds.

High-Gradient Location Fraction Galaxy10 SC-CD Galaxy10 DC-SD Office-Home SC-CD Office-Home DC-SD
Top 10% 0.5449 0.5551 0.2774 0.3035
Top 20% 0.5485 0.5734 0.3332 0.3554
Top 30% 0.5497 0.5847 0.4372 0.4558

Key Findings

  • Different-class same-domain overlap is consistently higher in Table 6, suggesting domain-related information in high-gradient locations; the gaps are small and lack corresponding significance tests, so they do not establish identification of causal domain features.
  • Connectivity analysis trains binary classifiers to distinguish two classโ€“domain pairs and uses test error to estimate connectivity. \(\alpha\) corresponds to same-class cross-domain pairs, \(\beta\) to different-class same-domain pairs, and \(\gamma\) to different-class cross-domain pairs; \(\alpha/\gamma\) changes from 0.33 to 4.16 on iWildCam and from 7.50 to 40.0 on Camelyon17.
  • These ratios capture empirical cross-domain connectedness, not direct measurement of spurious-feature removal; Camelyon17's \(\beta/\gamma\) also declines from 94.5 to 24.3, suggesting possible perturbation of class-relevant information.
  • ConvNeXt and ViT experiments in Table 5 support applicability beyond the original CNN backbones, but not effectiveness for arbitrary architectures or distribution shifts.
  • Table 8 reports D-GAP iteration time \(90.46\pm2.38\) ms, throughput \(353.9\pm9.2\) img/s, and peak memory 0.86 GB; ERM reports \(24.46\pm0.17\) ms, \(1308.3\pm9.1\) img/s, and 0.86 GB. Equal memory does not imply equal computational cost.

Highlights & Insights

  • Turning model response into augmentation parameters. Gradients generate a frequency-location-dependent interpolation map rather than a worst-case pixel attack, connecting current task dependence to conventional Fourier augmentation.
  • Two mixing weights serve different roles. \(\lambda_1\) controls cross-domain mixing within the pixel view, while \(\lambda_2\) controls the combination of frequency and pixel views, separately regulating domain-information injection and reliance on reconstruction.
  • Phase preservation is a soft semantic constraint. It avoids modifying amplitude and phase simultaneously, but is not a mathematical guarantee of label invariance; this interpretation better explains the pixel branch and ablation findings.

Limitations & Future Work

  • Additional gradient computation is a real cost. The authors suggest periodic sensitivity-map updates, low-resolution spectral proxies, or lightweight sensitivity predictors; these are not validated optimizations in this paper.
  • Target-input access limits extrapolation. Galaxy10 and the common benchmarks use unlabeled target data, so absence of target labels alone does not establish suitability for deployment with a completely unknown target domain.
  • The validation description is ambiguous. The Galaxy10 appendix first states that DECaLS supplies training and validation data, then describes selecting the linear-probe checkpoint by the best OOD validation performance. It does not clearly identify a separate OOD validation set; reproduction requires clarification, and neither use nor non-use of SDSS labels for tuning should be inferred.
  • Sensitivity still lacks a spurious-correlation identification guarantee. Controlled background or acquisition interventions could distinguish useful-but-sensitive frequencies from domain-specific sensitive frequencies and test failure conditions of the gradient heuristic.
  • Detail compensation remains empirical. Augmentation visualizations, accuracy, and connectivity alone cannot establish causal disentanglement; detail-retention metrics, label-preservation evaluation, and controls with strictly matched target-access permissions are needed.
  • vs FACT / FDA: These methods also modify appearance through amplitude while retaining source phase; D-GAP instead derives frequency-wise mixing strengths from task gradients and adds pixel compensation rather than using only fixed low-frequency replacement or predefined interpolation.
  • vs SAM: Here SAM means Semantic-aware Mixup, not Sharpness-Aware Minimization. The paper treats it as a semantic-aware Fourier mixing baseline, whereas D-GAP emphasizes sensitivity maps driven by the current model's response.
  • vs Connect Later: D-GAP retains the pretraining and LP-FT strategy but replaces handcrafted Copy-Paste or staining perturbations with gradient-based augmentation across datasets; suitable sampling protocols and validation selection remain necessary.
  • Research lead: Comparing source-only mixing, mixing with an independent unlabeled target set, and mixing with an unlabeled test set under the same budget could separate augmentation-design gains from target-distribution access gains; this is an experimental suggestion from this note, not a conclusion of the paper.

Rating

  • Novelty: 4/5 โ€” Task-gradient-driven amplitude interpolation and dual-space fusion form a clear combined contribution, building on existing Fourier augmentation.
  • Experimental Thoroughness: 4/5 โ€” Seven datasets, branch and mask ablations, and multiple backbones; target-access protocols and baseline provenance limit strict cross-method conclusions.
  • Writing Quality: 3/5 โ€” v4 adds sensitivity evidence and cost analysis, but Galaxy10 validation and DG terminology require clarification.
  • Value: 4/5 โ€” A practical interface for reducing task-specific augmentation design, with utility depending on target-data availability and training budget.