Skip to content

SFKD: Spatial–Frequency Joint-Aware Heterogeneous Knowledge Distillation via Multi-Level Wavelet Spectral Interaction

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/cpcpWang/SFKD
Area: Model Compression
Keywords: Heterogeneous Knowledge Distillation / Discrete Wavelet Transform / Fast Fourier Transform / Spatial–Frequency Joint Awareness / Model Compression

TL;DR

To resolve representational spatial distribution discrepancies caused by divergent inductive biases across heterogeneous architectures (CNN, Transformer, MLP-Mixer), SFKD couples multi-level discrete wavelet spatial decoupling, dual-stream dual-stage spectral refinement, and Gaussian-filtered Fourier frequency contrastive loss to preserve and transfer global structures and local details without discarding spatial information.

Background & Motivation

Knowledge distillation serves as a cornerstone model compression technique, predominantly investigated in homogeneous settings where teachers and students share identical architectural paradigms (e.g., CNN-to-CNN or Transformer-to-Transformer). However, in resource-constrained edge deployment scenarios, high-performing homogeneous teachers are often scarce. More importantly, strictly confining distillation within homogeneous models fails to capitalize on complementary inductive biases across distinct architectural families—for instance, CNNs inherently provide strong spatial locality and translation invariance, whereas Vision Transformers excel at long-range global contextual reasoning. Transferring knowledge from a heterogeneous teacher into a compact student model often unlocks superior compression potential beyond what homogeneous counterparts can offer.

The primary hurdle in heterogeneous distillation lies in the substantial discrepancy between intermediate feature distributions. Because the underlying operations differ fundamentally (local convolution vs. patch tokenization and global self-attention), representations from teacher and student exhibit severe spatial misalignment. Prior heterogeneous distillation methods (such as OFA and FBT) mitigate this structural collision by intentionally weakening, compressing, or outright discarding spatial features—for example, projecting representations into low-dimensional global embeddings, mapping intermediate features to logit space, or utilizing structure-agnostic objectives. Nevertheless, intermediate feature maps encode essential global semantic topology as well as fine-grained architectural cues. Excessively suppressing spatial information eliminates the student's opportunity to acquire rich cross-architecture geometric structures.

To fully exploit spatial representations in heterogeneous distillation without suffering from architectural spatial bias mismatches, this work introduces a spatial–frequency joint-aware distillation framework. Core idea: exploit the spatial locality of discrete wavelet transforms to explicitly decouple heterogeneous representations into low-frequency and high-frequency sub-bands, adaptively refine the student's sub-band stability via a dual-stream dual-stage module, and align dominant spectral energy and global patterns via a Gaussian-filtered Fourier contrastive loss.

Method

Overall Architecture

The SFKD framework orchestrates three sequential components: the Multi-level Discrete Wavelet Transform (MDWT) module for spatial decoupling, the Dual-Stream Dual-Stage Spectral Refinement (DS2SR) module for sub-band denoising, and the Gaussian-Filtered Frequency Loss (GFFL) for spectral alignment. For teacher representations \(F_t\) and projector-aligned student representations \(F_s\), a \(J\)-level Haar-based discrete wavelet transform first decouples the features into a low-frequency sub-band capturing global semantic contours and multi-level high-frequency sub-band sets encoding directional edge textures. To address structural instability and spectral leakage in student representations caused by limited network capacity and finite-length wavelet filters, the DS2SR module employs complementary Local Convolutional Refinement (LCR) and Global Transformer Refinement (GTR) streams with unidirectional cross-frequency interaction. Finally, 2D Fast Fourier Transform (FFT) projects the refined sub-bands into the frequency domain, where Gaussian low-pass and high-pass filtering masks guide InfoNCE contrastive alignment on Fourier amplitude spectra alongside inverse wavelet reconstruction and task-oriented logits distillation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Features<br/>Teacher Ft and projected student Fs"] --> B["Multi-Level DWT (MDWT)<br/>Spatial decoupling into low-frequency LJ and high-frequency HJ sub-bands"]
    B --> C["Dual-Stream Dual-Stage Spectral Refinement (DS2SR)<br/>LCR local convolution and GTR masked global attention"]
    C --> D["Adaptive Weighted Fusion (AWF)<br/>Integration of local and global enhanced sub-bands"]
    D --> E["Gaussian-Filtered Frequency Loss (GFFL)<br/>FFT amplitude alignment + IMDWT inverse feature reconstruction"]
    E --> F["Total Distillation Objective<br/>Spectral sub-band loss + reconstructed feature loss + Logits KL"]

Key Designs

1. Multi-Level Discrete Wavelet Transform (MDWT): Explicit Spatial Decoupling of Scale and Geometry Directly aligning heterogeneous features in raw pixel coordinate space induces severe conflicts between convolutional local priors and Transformer patch-level tokenization. To isolate transferable structural semantics from architecture-specific local fluctuations, SFKD employs the orthogonal and invertible Haar wavelet transform. Applying \(J\)-level 2D discrete wavelet decomposition recursively on input feature maps \(F \in \mathbb{R}^{B \times C \times H \times W}\) explicitly decomposes the representation into one low-frequency sub-band \(L^J \in \mathbb{R}^{B \times C \times \frac{H}{2^J} \times \frac{W}{2^J}}\) and a high-frequency sub-band set across levels \(\mathcal{H}^j = \{H^{j, LH}, H^{j, HL}, H^{j, HH}\}\) (\(j \in \{1, \dots, J\}\)). The low-frequency sub-band aggregates invariant global geometric outlines, while the high-frequency sub-bands isolate localized horizontal, vertical, and diagonal edge textures. This decoupling allows distinct spectral bands to be supervised independently while preserving full lossless reconstructibility via Inverse Multi-level Discrete Wavelet Transform (IMDWT).

2. Dual-Stream Dual-Stage Spectral Refinement (DS2SR): Complementary Inductive Biases and Asymmetric Cross-Frequency Guidance Because practical wavelet filters operate with finite support kernels, ideal band-limited separation is unattainable, leading to spectral leakage and residual frequency overlap across sub-bands. Furthermore, the limited capacity of compact student networks yields pronounced structural instability and representation distortion in its wavelet sub-bands. DS2SR resolves this via two complementary streams: - Local Convolutional Refinement (LCR): Employs dedicated residual convolutional blocks (ResConvBlock) with multiple strides and weighted skip connections to enhance local geometric consistency and prevent over-smoothing; - Global Transformer Refinement (GTR): First tokenizes sub-band features via convolutional projection and injects a specialized 4D Positional Encoding \([j, d, Pos_h, Pos_w]\), where \(j\) indexes the decomposition level, \(d \in \{0, 1, 2, 3\}\) indicates the sub-band orientation (low-frequency or one of three high-frequency orientations), and \(Pos_h, Pos_w\) denote 2D spatial coordinates, resolving spatial and hierarchical ambiguities; - Dual-Stage Masked Cross-Frequency Interaction: In GTR, stage one processes low-frequency tokens through unmasked self-attention; stage two concatenates low-frequency and high-frequency tokens under an asymmetric attention mask that allows low-frequency tokens to attend to high-frequency tokens (utilizing edge cues to suppress leaked high-frequency noise) while strictly blinding high-frequency tokens to low-frequency tokens to prevent reverse interference; - Adaptive Weighted Fusion (AWF): Low-frequency and high-frequency sub-bands from both streams are fused through separate adaptive gating networks to produce refined student sub-bands \(L_{\mathrm{ref}}^{J, s}\) and \(\mathcal{H}_{\mathrm{ref}}^{s}\).

3. Gaussian-Filtered Frequency Loss (GFFL): Amplitude Spectral Energy Alignment with Smooth Frequency Selection While wavelets provide spatial localization, point-to-point spatial matching remains vulnerable to slight non-rigid spatial distortions and phase shifts. In contrast, Fourier amplitude spectra characterize global energy distributions and topological structures with translation invariance. SFKD computes the 2D FFT of sub-bands and extracts amplitude spectra \(|\mathcal{F}(X)|(u,v) = \sqrt{\mathrm{Re}(X)^2 + \mathrm{Im}(X)^2}\). To circumvent Gibbs ringing artifacts and boundary distortion caused by hard truncation filters, smooth Gaussian low-pass filters \(G_{\mathrm{low}}(u,v) = \exp(-\frac{u^2+v^2}{2\sigma^2})\) and high-pass filters \(G_{\mathrm{high}}(u,v) = 1 - G_{\mathrm{low}}(u,v)\) are constructed. Low-pass masks are applied to central energy-concentrated low-frequency sub-bands, whereas high-pass masks isolate peripheral high-frequency textures, supervised by an InfoNCE contrastive objective:

\[\mathcal{L}_{\mathrm{GFFL}}(X^t, X^s; G) = \mathcal{L}_{\mathrm{InfoNCE}}(G \odot |\mathcal{F}(X^t)|, \, G \odot |\mathcal{F}(X^s)|)\]

Loss & Training

To enforce multi-tier alignment spanning isolated sub-bands, unified whole representations, and downstream task decisions, the overall training objective combines four synergistic terms: 1. Low-Frequency Sub-Band Loss: \(\mathcal{L}_{\mathrm{LF}} = \mathcal{L}_{\mathrm{GFFL}}(L^{J, t}, L_{\mathrm{ref}}^{J, s}; G_{\mathrm{low}})\); 2. High-Frequency Sub-Band Loss: \(\mathcal{L}_{\mathrm{HF}} = \sum_{j=1}^J \mathcal{L}_{\mathrm{GFFL}}(\mathcal{H}^{j, t}, \mathcal{H}_{\mathrm{ref}}^{j, s}; G_{\mathrm{high}})\); 3. Reconstructed Feature Loss: Refined student sub-bands are inverted via IMDWT to produce reconstructed representation \(F_s^{\mathrm{rec}}\), aligned with teacher representation \(F_t\) using unfiltered spectral loss to maintain overall anisotropic consistency: \(\mathcal{L}_{\mathrm{Rec}} = \mathcal{L}_{\mathrm{GFFL}}(F_t, F_s^{\mathrm{rec}}; \mathbf{1})\); 4. Logits KL Divergence Loss: \(F_s^{\mathrm{rec}}\) is passed through a newly initialized classifier head producing logits \(Z_s^{\mathrm{rec}}\), supervised by teacher logits \(Z_t\) via KL divergence: \(\mathcal{L}_{\mathrm{KL}} = \mathcal{L}_{\mathrm{KL}}(Z_s^{\mathrm{rec}}, Z_t)\).

The total training objective is balanced by loss coefficients and task cross-entropy:

\[\mathcal{L}_{\mathrm{total}} = \lambda_1 \mathcal{L}_{\mathrm{LF}} + \lambda_2 \mathcal{L}_{\mathrm{HF}} + \lambda_3 \mathcal{L}_{\mathrm{Rec}} + \lambda_4 \mathcal{L}_{\mathrm{KL}} + \mathcal{L}_{\mathrm{cls}}\]

Key Experimental Results

Main Results

SFKD is extensively benchmarked across 12 heterogeneous teacher-student pairs on CIFAR-100 and 14 heterogeneous pairs plus 2 homogeneous pairs on ImageNet-1K, covering CNNs, Vision Transformers, and MLP-based architectures.

Dataset Teacher \(\to\) Student Teacher Acc Student Scratch Prev. SOTA (FBT/OFA) SFKD (Ours) Gain
CIFAR-100 Swin-T \(\to\) ResNet18 89.26% 74.01% 81.61% (FBT) 83.26% +1.65%
CIFAR-100 ViT-S \(\to\) ResNet18 92.04% 74.01% 81.93% (FBT) 82.10% +0.17%
CIFAR-100 Mixer-B/16 \(\to\) ResNet18 87.29% 74.01% 81.90% (FBT) 81.98% +0.08%
CIFAR-100 Swin-T \(\to\) MobileNetV2 89.26% 73.68% 81.28% (FBT) 82.03% +0.75%
CIFAR-100 ConvNeXt-T \(\to\) DeiT-T 88.41% 68.00% 79.57% (FBT) 83.91% +4.34%
CIFAR-100 Mixer-B/16 \(\to\) DeiT-T 87.29% 68.00% 74.40% (FBT) 81.14% +6.74%
CIFAR-100 ConvNeXt-T \(\to\) Swin-P 88.41% 72.63% 80.73% (FBT) 82.16% +1.43%
CIFAR-100 ConvNeXt-T \(\to\) ResMLP-S12 88.41% 66.56% 78.03% (FBT) 85.08% +7.05%
CIFAR-100 Swin-T \(\to\) ResMLP-S12 87.29% 66.56% 77.20% (FBT) 84.01% +6.81%
ImageNet-1K Swin-T \(\to\) ResNet18 81.38% 69.75% 72.21% (FBT) 72.54% +0.33%
ImageNet-1K Swin-T \(\to\) MobileNetV2 81.38% 68.87% 72.54% (FBT) 72.67% +0.13%
ImageNet-1K ResNet50 \(\to\) DeiT-T 80.38% 72.17% 75.84% (FitNet) 75.89% +0.05%
ImageNet-1K ResNet50 \(\to\) Swin-N 80.38% 75.53% 77.79% (FBT) 78.28% +0.49%
ImageNet-1K ConvNeXt-T \(\to\) ResMLP-S12 82.05% 76.65% 77.33% (FBT) 78.35% +1.02%
ImageNet-1K (Homo) ResNet34 \(\to\) ResNet18 73.31% 69.75% 72.29% (FBT) 72.34% +0.05%
ImageNet-1K (Homo) ResNet50 \(\to\) MobileNet 76.61% 68.58% 73.45% (FBT) 73.81% +0.36%

Ablation Study

1. Contribution of Loss Components (ImageNet-1K: Swin-T \(\to\) ResNet18)

Configuration \(\mathcal{L}_{\mathrm{LF}}\) \(\mathcal{L}_{\mathrm{HF}}\) \(\mathcal{L}_{\mathrm{Rec}}\) \(\mathcal{L}_{\mathrm{KL}}\) Top-1 Accuracy (%) Note
w/o Low-Freq Loss \(\times\) \(\checkmark\) \(\checkmark\) \(\checkmark\) 72.19% Largest drop (-0.35%), proving global structure dominates transfer
w/o High-Freq Loss \(\checkmark\) \(\times\) \(\checkmark\) \(\checkmark\) 72.32% Drops 0.22%, indicating high-frequency boundary cues are vital
w/o Reconstruction Loss \(\checkmark\) \(\checkmark\) \(\times\) \(\checkmark\) 72.31% Drops 0.23%, showing sub-band alignment requires whole-feature constraint
w/o KL Divergence Loss \(\checkmark\) \(\checkmark\) \(\checkmark\) \(\times\) 72.46% Drops 0.08%, task-oriented supervision adds necessary regularization
Full Model (SFKD) \(\checkmark\) \(\checkmark\) \(\checkmark\) \(\checkmark\) 72.54% Joint objective achieves the best performance

2. Effect of DS2SR Refinement Streams (ImageNet-1K: Top-1 Accuracy %)

LCR Stream GTR Stream Swin-T \(\to\) ResNet18 (CNN Student) ResNet50 \(\to\) Swin-N (Transformer Student) Analysis
\(\times\) \(\times\) 71.24% 76.83% Baseline without refinement suffers from spectral leakage and noise
\(\checkmark\) \(\times\) 72.19% 78.24% CNN student +0.95%, Transformer student +1.41%
\(\times\) \(\checkmark\) 72.46% 77.83% CNN student prefers GTR (+1.22%) to acquire global context
\(\checkmark\) \(\checkmark\) 72.54% 78.28% Dual-stream complementary fusion yields optimal accuracy

Key Findings

  • Low-frequency sub-bands form the core transfer foundation: Removing \(\mathcal{L}_{\mathrm{LF}}\) causes the sharpest accuracy decline (72.54% down to 72.19%), confirming that macro-level topological structure and global semantic contours govern cross-architecture knowledge transfer.
  • Cross-architectural inductive bias preference: CNN students (ResNet18) benefit substantially more from the global self-attention GTR stream (+1.22%) than from the local convolutional LCR stream (+0.95%). Conversely, Transformer students (Swin-N), which inherently lack translation invariance and local inductive bias, gain significantly more from the LCR stream (+1.41%) than from GTR (+1.00%). This verifies that heterogeneous distillation thrives when actively compensating for missing architectural biases.
  • Spectral alignment mitigates representational collapse: Standard feature distillation methods utilizing pixel-wise MSE (e.g., FitNet) suffer catastrophic degradation when distilling CNNs into Transformers (e.g., scoring only 24.06% on CIFAR-100 ConvNeXt-T \(\to\) Swin-P), whereas SFKD achieves 82.16% by aligning translation-invariant Fourier amplitude spectra.

Highlights & Insights

  • Unified spatial-frequency decoupling: Rather than collapsing feature maps into 1D embeddings to avoid spatial misalignment, SFKD explicitly decomposes features via wavelets into localized sub-bands and aligns their global spectral distribution in Fourier space, harmonizing spatial locality with spectral robustness.
  • 4D positional encoding for multi-level wavelet tokens: Treating the decomposition level index \(j\) and sub-band orientation type \(d\) as formal structural dimensions alongside spatial coordinates \([j, d, Pos_h, Pos_w]\) provides a clean, generalizable mechanism for standard Transformers to process complex hierarchical multi-band feature grids.
  • Directional cross-frequency interaction mask: Allowing low-frequency tokens to attend to high-frequency tokens to remove leaked noise while blocking the reverse flow prevents destructive interference and establishes an effective asymmetric information-routing pattern.

Limitations & Future Work

  • Computational overhead during training: Evaluating multi-level DWT, dual-stream refinement networks, and 2D FFTs increases per-step GPU memory consumption and training latency compared to simple logit distillation.
  • Filter bandwidth hyperparameter tuning: The Gaussian filter bandwidth \(\sigma\) and wavelet level \(J\) require empirical adjustments across image resolutions (e.g., 32x32 vs. 224x224); adaptive or learnable filtering bandwidths warrant further exploration.
  • Expansion to dense prediction tasks: Current empirical validation centers on image classification; evaluating whether wavelet-frequency decoupling preserves spatial fidelity for pixel-sensitive downstream tasks such as object detection and semantic segmentation remains a valuable future direction.
  • vs OFA (One-for-All): OFA sidesteps spatial discrepancy by projecting intermediate features into non-spatial logit space and necessitates 4-stage feature alignment; SFKD relies solely on the final feature layer and directly unlocks rich spatial structures via wavelet-frequency transforms, achieving higher accuracy with a lighter distillation interface.
  • vs FBT (Fuse Before Transfer): FBT relies on training a heavy hybrid bridging model to reconcile heterogeneous representations; SFKD operates directly on orthogonal mathematical bases (Haar wavelets and Fourier spectra) without auxiliary bridging architectures, outperforming FBT by an average of 2.54% across CIFAR-100 pairs.
  • vs FitNet / FCFD: Direct spatial MSE alignment collapses under cross-architecture inductive bias conflicts; SFKD provides an invariant spectral perspective that stabilizes feature distillation across heterogeneous backbones.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Elegant integration of wavelet spatial decoupling with Fourier amplitude alignment, featuring 4D positional encoding and unidirectional cross-frequency masking.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Extensive evaluation across 26 distinct teacher-student combinations spanning CIFAR-100 and ImageNet-1K with comprehensive ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulation and coherent structural motivation throughout.
  • Value: ⭐⭐⭐⭐⭐ Overturns the prevailing dogma that heterogeneous distillation must discard spatial information, establishing a strong benchmark for cross-architecture compression.