Skip to content

Fourier Self-Supervision for Fine-Grained Generalized Category Discovery

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/SarahRastegar/FourEx
Area: Self-Supervised Learning
Keywords: generalized category discovery, Fourier transform, self-supervised contrastive learning, fine-grained classification, dual-frequency filtering

TL;DR

To overcome the failure of standard contrastive learning in fine-grained Generalized Category Discovery caused by subtle inter-class differences, this paper introduces Fourier Self-Supervision that decouples the latent space into low- and high-frequency sub-dimensions based on adaptive SNR thresholds, harnessing low frequencies for abstract hierarchical discovery and high frequencies for discriminative fine details.

Background & Motivation

Generalized Category Discovery (GCD) seeks to categorize unlabeled data comprising both known classes and undiscovered novel categories, guided by a limited set of labeled known samples. While established GCD methods built upon spatial contrastive learning and prototype-based clustering perform impressively on coarse-grained benchmarks like CIFAR or ImageNet subsets, they suffer severe degradation when confronted with fine-grained visual categories such as bird species, aircraft models, or car trims. In fine-grained domains, inter-class appearance discrepancies diminish dramatically, and conventional aggressive spatial augmentations (e.g., severe color jittering, random cropping) inadvertently destroy fragile discriminative signaturesโ€”such as beak shapes or feather texturesโ€”forcing models to rely on spurious background correlations or coarse silhouettes.

The core tension in fine-grained GCD stems from the contradictory representational demands of category discovery versus category discrimination. Discovering novel classes requires broad, abstract semantics to infer implicit hierarchical structures that constrain the novel category search space and prevent parent-sibling confusion. Conversely, distinguishing subtle, visually adjacent fine-grained categories demands localized, high-frequency boundary and texture cues. Conventional spatial-domain contrastive approaches entangle these disparate spectral properties within an undifferentiated latent space, leading either to overfitting on localized artifacts or to feature collapse where subtle differences are washed out.

Inspired by Fourier optics and the F-principle of neural networksโ€”which dictates that deep networks naturally prioritize lower spectral components before fitting higher frequenciesโ€”this paper recognizes that category semantics are distributed predictably across the Fourier spectrum. Low-frequency reconstructions provide coarse, contextual representations that outline taxonomic relationships, while high-frequency reconstructions isolate sharp structural details. Core idea: decompose images into orthogonal low-pass and high-pass Fourier views, performing dynamic subspace contrastive learning across decoupled latent dimensions to steer novel category emergence via low-frequency abstraction while sharpening fine-grained discrimination via high-frequency details.

Method

Overall Architecture

The proposed Fourier Self-Supervision pipeline operates through input preprocessing and frequency transformation, dual-frequency orthogonal contrastive representation learning, and multi-view classification consistency. Given an input image, the model first applies a lightweight Gaussian blur to suppress Gibbs ringing artifacts, followed by a Fast Fourier Transform (FFT) to convert spatial patterns into frequency magnitudes and phases. Based on a standardized Signal-to-Noise Ratio (SNR) criterion, dataset-dependent cutoff frequencies are computed to construct an ideal low-pass filter (LPF) and a high-pass filter (HPF). Reconstructed low-frequency and high-frequency spatial views are obtained via inverse FFT (iFFT). A Vision Transformer maps original and frequency-filtered views into a shared latent space partitioned into a left half (low frequencies) and a right half (high frequencies), allocating dynamic slice dimensions according to sampling cutoff ratios, and optimizing jointly with labeled classification cross-entropy.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image x_i<br/>Gaussian Smoothing"] --> B["Adaptive SNR Frequency Filtering<br/>FFT Transform + SNR Cutoff Thresholds"]
    B --> C["Frequency Reconstructions<br/>LPF Low-Pass / HPF High-Pass + iFFT"]
    C --> D["Low-Frequency Contrastive Learning<br/>Left Latent Subspace Alignment"]
    C --> E["High-Frequency Contrastive Learning<br/>Right Latent Subspace Texture Separation"]
    D --> F["Frequency Augmentation Classification Consistency<br/>Original/Low/High Joint BCE Objective"]
    E --> F
    F --> G["Fine-Grained Category Discovery & Clustering"]

Key Designs

1. Adaptive SNR Frequency Filtering: establishing objective, dataset-dependent cutoffs

Arbitrary selection of cutoff frequencies introduces severe representational risks: setting the low-pass cutoff too high preserves excessive fine details, rendering the view redundant with the original image, while setting the high-pass cutoff too low permits overwhelming high-frequency sensor noise that masks informative textures. To establish an objective criterion, the framework adopts the 15 dB standard (equivalent to a linear signal ratio of 32) commonly utilized in digital video and subjective image quality assessments. For an image Fourier transform \(X\), the SNR at a cutoff frequency \(f_c\) is defined as the ratio of lower-frequency energy to higher-frequency residual energy:

\[SNR(f_c) = 10 \log_{10} \left( \frac{\sum_{f \le f_c} |X(f)|^2}{\sum_{f > f_c} |X(f)|^2} \right)\]

By averaging the frequency thresholds that satisfy the 15 dB criterion across all categories in the dataset, the framework derives robust low-pass and high-pass cutoff thresholds \(T_l\) and \(T_h\). Additionally, to suppress Gibbs ringing artifacts that naturally emerge around sharp image borders during truncated inverse Fourier transforms, a Gaussian blur kernel of size 5 is applied prior to frequency filtering, guaranteeing clean, artifact-free reconstructed views.

2. Low-Frequency Contrastive Learning: constraining implicit hierarchies for novel category generalization

Human perception reliably identifies broad category memberships (e.g., distinguishing a marine bird from a woodland bird) even when all high-frequency textures and edge details are stripped away. Guided by the F-principle, the model samples a random cutoff frequency \(f_l \in [1, T_l]\) and generates a low-frequency reconstruction \(l_i = \mathcal{F}^{-1}(H_{LP} \odot \mathcal{F}(x_i))\). Rather than projecting onto the full latent dimension \(D\), the left half of the latent space (\(D/2\)) is exclusively assigned to low-frequency contrast, dynamically allocating the first \(d = \frac{f_l}{T_l} \cdot \frac{D}{2}\) leftmost dimensions:

\[\mathcal{L}_{\text{low}} = (1-\lambda) \sum_{i \in \mathcal{B}} \mathcal{L}_{i|d}^u + \lambda \sum_{i \in \mathcal{B}_{\mathcal{L}}} \mathcal{L}_{i|d}^s\]

Here, \(\mathcal{L}_{i|d}^u\) and \(\mathcal{L}_{i|d}^s\) represent unsupervised and supervised contrastive losses evaluated over the designated sub-vector slice. As the sampling frequency progressively moves from coarse to fine, the latent subspace enforces an implicit taxonomic hierarchy, enabling unlabeled novel samples to leverage shared parent-level features from known categories and substantially suppressing category confusion.

3. High-Frequency Contrastive Learning: highlighting nuanced textures for fine-grained discrimination

The definitive features that separate visually similar fine-grained species reside predominantly in high-frequency patterns, such as wing markings or beak curvatures. Because high frequencies carry significantly less total energy than low-frequency baselines, they are easily drowned out in conventional training. The framework isolates these nuances by sampling a high-frequency cutoff \(f_h < T_h\) and generating a high-pass filtered reconstruction \(h_i = \mathcal{F}^{-1}(H_{HP} \odot \mathcal{F}(x_i))\). In direct mirror symmetry to the low-frequency setup, high frequencies are anchored to the rightmost half of the latent space, contrasting across the terminal \(-d\) dimensions (the rightmost \(\frac{f_h}{T_h} \cdot \frac{D}{2}\) dimensions):

\[\mathcal{L}_{\text{high}} = (1-\lambda) \sum_{i \in \mathcal{B}} \mathcal{L}_{i|-d}^u + \lambda \sum_{i \in \mathcal{B}_{\mathcal{L}}} \mathcal{L}_{i|-d}^s\]

This structural decoupling ensures that high-pass contrastive gradients continuously stretch fine-grained inter-class margins without disrupting the smooth, overarching semantic manifolds stabilized in the left latent subspace.

4. Frequency Augmentation Classification Consistency: multi-view alignment and anti-overfitting

Although the low-frequency reconstruction \(l_i\), high-frequency reconstruction \(h_i\), and original image \(x_i\) exhibit radically divergent spectral characteristics, they depict the exact same semantic instance. For labeled samples, the framework imposes a joint binary cross-entropy (BCE) classification loss:

\[\mathcal{L}_{\text{cls}} = \mathcal{L}_{\text{BCE}}(p_{h_i}, y_i) + \mathcal{L}_{\text{BCE}}(p_{l_i}, y_i) + \mathcal{L}_{\text{BCE}}(p_{x_i}, y_i)\]

This objective compels the classifier to preserve correct posterior class probabilities even under degraded views (e.g., when color is completely absent in high-pass views, or when edges are blurred in low-pass views), forcing the network to extract robust, frequency-invariant representations while counteracting representation overfitting.

Loss & Training

The overall training objective aggregates the baseline GCD loss with the three frequency-aware self-supervised objectives:

\[\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{base}} + \alpha_{\text{low}} \mathcal{L}_{\text{low}} + \alpha_{\text{high}} \mathcal{L}_{\text{high}} + \alpha_{\text{cls}} \mathcal{L}_{\text{cls}}\]

A ViT-B/16 serves as the primary vision backbone. When integrated as a plug-and-play enhancement, FourSim freezes the initial 11 transformer blocks and fine-tunes only the final block on top of SimGCD, whereas FourEx fine-tunes the last three transformer blocks on top of SelEx. Balancing hyper-parameters are set empirically to \(\alpha_{\text{low}} = 0.1\), \(\alpha_{\text{high}} = 0.1\), and \(\alpha_{\text{cls}} = 0.1\).

Key Experimental Results

Main Results

Clustering accuracies across three fine-grained benchmarks (CUB-200, FGVC-Aircraft, Stanford-Cars) for All, Known, and Novel categories using a DINOv1 backbone are summarized below:

Dataset Metric SelEx Baseline FourEx (Ours) Gain
CUB-200 All / Known / Novel 76.4 / 72.4 / 78.4 80.2 / 75.1 / 82.7 +3.8 / +2.7 / +4.3
FGVC-Aircraft All / Known / Novel 61.2 / 67.7 / 58.0 65.9 / 67.9 / 64.9 +4.7 / +0.2 / +6.9
Stanford-Cars All / Known / Novel 56.9 / 76.9 / 47.3 59.2 / 76.9 / 50.6 +2.3 / 0.0 / +3.3
Average All / Known / Novel 64.8 / 72.3 / 61.2 68.4 / 73.3 / 66.1 +3.6 / +1.0 / +4.9

On the long-tailed, highly imbalanced Herbarium19 benchmark, FourEx achieves 44.5% All, 61.5% Known, and 35.3% Novel accuracy (+4.9%, +6.6%, and +4.0% over SelEx). On the small-scale, data-scarce Oxford-IIIT Pet dataset, FourEx attains 92.8% All and 93.4% Novel accuracy.

Ablation Study

Ablation of individual loss components across CUB-200, FGVC-Aircraft, and Stanford-Cars (using the 3-block tuned SelEx baseline on DINOv1):

Config \(\mathcal{L}_{\text{low}}\) \(\mathcal{L}_{\text{high}}\) \(\mathcal{L}_{\text{cls}}\) CUB-200 (All) Aircraft (All) Cars (All) Avg (All)
Baseline (SelEx) - - - 76.4 61.2 56.9 64.8
Low-pass only โœ“ - - 77.5 63.8 56.3 65.9
High-pass only - โœ“ - 78.8 63.9 56.8 66.5
Classification only - - โœ“ 78.0 62.8 57.1 66.0
Dual-frequency โœ“ โœ“ - 80.0 64.9 57.7 67.5
Low + Classification โœ“ - โœ“ 77.9 64.5 57.7 66.7
High + Classification - โœ“ โœ“ 79.7 63.2 57.9 66.9
Full Model โœ“ โœ“ โœ“ 80.2 65.9 59.2 68.4

Key Findings

  • High-frequency contrastive learning (\(\mathcal{L}_{\text{high}}\)) yields the sharpest single-module boost on novel category discovery (e.g., CUB-200 Novel jumping to 81.9%), corroborating that high-frequency textures provide critical discriminative boundaries for previously unseen species.
  • Combining low- and high-frequency filtering (\(\mathcal{L}_{\text{low}} + \mathcal{L}_{\text{high}}\)) produces strong complementary gains (+2.7% over baseline), verifying that spectral decoupling prevents the feature entanglement common in spatial contrastive learning.
  • In realistic open-world evaluations with unknown category counts (estimated as 231 on CUB and 230 on Cars via semi-supervised estimation), FourEx obtains 77.3% All and 77.9% Novel on CUB, outperforming competitive methods like NC-GCD (70.3% / 69.4%).

Highlights & Insights

  • SNR-Grounded Spectral Cutoff Derivation: Replaces empirical trial-and-error cutoff choices with an objective 15 dB video/image quality threshold, providing an automated, dataset-adaptive mechanism for frequency boundary selection.
  • Orthogonal Latent Subspace Partitioning: Divides latent vectors into dedicated left (low-pass) and right (high-pass) halves and modulates the active contrastive dimension slices according to sampled frequencies, mathematically decoupling abstract geometry from subtle textures.
  • Seamless Plug-and-Play Integration: Operates purely as a self-supervised representation enhancement agent without requiring modifications to standard ViT architectures, consistently improving baselines like SimGCD and SelEx.

Limitations & Future Work

  • Vulnerability in Color-Dominated Classification: Because ideal high-pass filtering eliminates low-frequency DC illumination and smooth color channels, the high-frequency branch provides diminished utility for fine-grained taxa distinguished solely by color patterns rather than structural geometry.
  • Sensitivity to Specular Reflection and Occlusion: Failure analysis indicates that high-frequency reconstructions can over-amplify specular reflections and partial occlusions, occasionally biasing clustering toward visually related sibling species.
  • Coarse Spectral Bipartition: The current design uses a discrete binary split (low-pass vs. high-pass); future extensions could explore multi-scale continuous wavelet decompositions to form spectral feature pyramids.
  • vs Standard GCD Baselines (Vaze et al., CVPR 2022 / SimGCD, ICCV 2023): Conventional GCD baselines rely entirely on spatial-domain contrastive objectives, which suffer from background noise and localized feature entanglement in fine-grained tasks. This paper resolves this by introducing orthogonal frequency-filtered views.
  • vs Hierarchical GCD Frameworks (InfoSieve, NeurIPS 2023 / SEAL, NeurIPS 2025): Explicit hierarchical approaches require expensive tree-search graphs or multi-level clustering overhead. In contrast, Fourier Self-Supervision naturally induces hierarchical structures through low-frequency abstraction without structural inference costs.

Rating

  • Novelty: โญโญโญโญโ˜† [Elegant integration of Fourier optics and latent subspace decoupling for fine-grained category discovery]
  • Experimental Thoroughness: โญโญโญโญโญ [Evaluated across three fine-grained benchmarks, long-tailed distributions, unknown category estimation, and coarse-grained datasets]
  • Writing Quality: โญโญโญโญโญ [Clear motivation grounded in the F-principle, rigorous mathematical formulation, and thorough ablations]
  • Value: โญโญโญโญโ˜† [Provides an effective, model-agnostic frequency decoupling paradigm for self-supervised fine-grained representation learning]