Skip to content

SRRA: Stable-Rank-Based Residual Adaptation for Generalizable Deepfake Detection

Conference: ECCV 2026
Paper: ECCV Official
PDF: Open Access
Code: https://github.com/LHK-CodeLab/SRRA
Area: AI Safety
Keywords: Deepfake Detection, Cross-Domain Generalization, Parameter-Efficient Fine-Tuning, Stable Rank, Singular Value Decomposition

TL;DR

Addressing the generalization bottleneck caused by rigid uniform residual rank assignment across layers in fine-tuning pretrained Vision Transformers, this paper proposes Stable-Rank-Based Residual Adaptation (SRRA), which adaptively assigns low ranks to shallow layers and higher ranks to deeper layers via Stable Rank, combined with residual energy and orthogonality constraints to achieve state-of-the-art cross-domain deepfake detection.

Background & Motivation

With rapid advancements in generative adversarial networks (GANs) and diffusion models, facial manipulation techniques can now synthesize photo-realistic manipulated facial images and videos. The proliferation of such deepfakes presents significant societal risks, including identity theft, disinformation dissemination, and fraud. Early deepfake detectors predominantly relied on convolutional neural networks (such as Xception and EfficientNet) along with handcrafted frequency or boundary artifact cues. However, convolutional architectures exhibit strong inductive biases toward local textures, making them prone to shortcut learning and overfitting to generator-specific artifacts. Consequently, when confronted with unseen manipulation methods or distribution shifts, their detection accuracy deteriorates substantially.

To enhance cross-domain generalization, recent studies have turned to Vision Transformers (ViTs, such as pretrained CLIP models) that possess global receptive fields and weaker inductive biases. Nonetheless, because deepfake training datasets are relatively small, full fine-tuning often corrupts general pretrained visual representations and induces severe overfitting. While parameter-efficient fine-tuning (PEFT) approaches such as standard LoRA or SVD-based methods (e.g., Effort) alleviate parameter overhead, they universally assign an identical, fixed residual rank \(r\) across all Transformer self-attention layers. This uniform allocation overlooks the intrinsic representation disparity and singular value spectral properties across different layer depths.

Spectral analysis reveals that shallow layers in pretrained ViTs display a distinctly low-rank structure; their pretrained representations already sufficiently encode low-level cues such as edges and textures, thus necessitating only minimal adaptation. Conversely, middle and deep layers exhibit much higher effective ranks to handle long-range semantic composition and high-level structural patterns, demanding greater residual capacity to model intricate manipulation traces. Core idea: adaptively allocate layer-wise residual fine-tuning ranks via Stable Rank—assigning lower ranks to shallow layers to preserve general priors and higher ranks to deep layers for complex forgery modeling—while regularizing residual energy and singular subspace orthogonality to suppress overfitting and prevent directional collapse.

Method

Overall Architecture

SRRA adaptively decomposes and fine-tunes the linear projection weight matrices within the self-attention blocks of a pretrained Vision Transformer (e.g., CLIP ViT-L/14). The entire framework consists of three coordinated stages: first, Stable Rank Decomposition (SRD) adaptively partitions the pretrained weight matrix into a frozen principal component and a trainable residual component based on its stable rank; second, during fine-tuning, the Residual Energy Constraint (REC) enforces Frobenius energy preservation and nuclear-norm penalization on the residual singular values; finally, Residual Subspace Orthogonalization (RSO) enforces column orthogonality on the residual singular bases to prevent feature collapse onto isolated forgery directions.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Pretrained Attention Weights W<br/>and Face Images"] --> B["Stable Rank Decomposition (SRD)<br/>SVD + Adaptive Rank Split via Stable Rank"]
    B --> C["Residual Fine-Tuning & Forward Pass<br/>Frozen Principal Wr + Trainable Residual ΔW"]
    C --> D["Residual Energy Constraint (REC)<br/>Global Energy Conservation + Nuclear-Norm Penalty"]
    D --> E["Residual Subspace Orthogonalization (RSO)<br/>Orthogonal Regularization on Residual Bases"]
    E --> F["Output: Classification & Multi-Loss Backprop<br/>Generalizable Prediction & Joint Optimization"]

Key Designs

1. Stable Rank Decomposition: Dynamic Layer-Wise Capacity Allocation via Stable Rank

Uniform rank assignment across heterogeneous Transformer layers leads to shallow layer over-parameterization (disrupting well-formed low-level priors) and deep layer under-parameterization (failing to capture high-level semantic artifacts). To resolve this, SRRA introduces Stable Rank Decomposition (SRD). For any square weight matrix \(W \in \mathbb{R}^{n \times n}\) in the self-attention mechanism, singular value decomposition factorizes it into \(W = U \Sigma V^\top\), where singular values are sorted in descending order \(\sigma_1 \ge \sigma_2 \ge \dots \ge \sigma_d > 0\). The matrix is split into a principal component \(W_r = U_r \Sigma_r V_r^\top\) formed by the top \(r\) singular triplets, and a residual component \(\Delta W = U_{res} \Sigma_{res} V_{res}^\top\) formed by the remaining \(r_{res}\) triplets. Rather than picking an arbitrary energy cutoff, SRRA leverages the Stable Rank:

\[\operatorname{Srank}(W) = \frac{\|W\|_F^2}{\|W\|_2^2} = \frac{\sum_{i=1}^d \sigma_i^2}{\sigma_1^2}\]

Using a global scaling hyper-parameter \(\tau > 0\), the residual rank is computed as \(r_{res} = \lceil \tau \cdot \operatorname{Srank}(W) \rceil\), with the principal rank fixed to \(r = n - r_{res}\). Throughout training, the principal component \(W_r\) remains frozen to safeguard pretrained representations, while only the low-rank residual component \(\Delta W\) is updated. This guarantees that shallow layers receive compact residual budgets, whereas deeper layers gain higher adaptation capacity for intricate forgery patterns.

2. Residual Energy Constraint: Dual Energy Regularization Against Excessive Perturbations

Although increasing residual rank in deeper layers bolsters representation capacity, excessive residual freedom risks fitting sample-specific noise in small-scale deepfake datasets, causing the model to drift from the pretrained manifold. To maintain numerical stability while releasing deep capacity, SRRA introduces the Residual Energy Constraint (REC), comprising two complementary objectives:

\[\mathcal{L}_{rec1} = \left| \|\hat{W}\|_F^2 - \|W\|_F^2 \right|, \quad \mathcal{L}_{rec2} = \sum_{i=1}^{r_{res}} \hat{\sigma}_i\]

where \(\hat{W}\) denotes the updated weight matrix and \(\hat{\sigma}_i\) represents the updated singular values of the residual component. The first term \(\mathcal{L}_{rec1}\) enforces global Frobenius energy consistency before and after adaptation, preserving overall representational scale. The second term \(\mathcal{L}_{rec2}\) imposes a nuclear-norm regularizer on the residual component, directly penalizing runaway inflation of residual singular values. Together, they prevent the updated weights from corrupting the pretrained spectral geometry.

3. Residual Subspace Orthogonalization: Decoupling Residual Bases to Avoid Directional Collapse

Expanding the residual rank introduces multiple trainable singular directions in deeper layers. Unconstrained optimization of these vectors can easily collapse them onto a few dominant forgery artifacts present in the training set, limiting out-of-domain detection. Unlike previous frameworks like Effort that enforce cross-orthogonality between principal and residual components (which constrains residual flexibility and pulls it toward the principal space), SRRA designs Residual Subspace Orthogonalization (RSO) strictly within the residual singular bases:

\[\mathcal{L}_{rso} = \|\hat{U}_{res}^\top \hat{U}_{res} - I\|_F + \|\hat{V}_{res}^\top \hat{V}_{res} - I\|_F\]

RSO penalizes non-orthogonal correlations among the columns of the left and right residual singular vector matrices \(\hat{U}_{res}\) and \(\hat{V}_{res}\). This drives the residual adapter to distribute its expressive power across diverse, decoupled directions, significantly lowering the risk of overfitting to generator-specific shortcuts.

Loss & Training

SRRA is trained end-to-end via a composite objective combining cross-entropy classification and layer-wise regularizers:

\[\mathcal{L}_{total} = \mathcal{L}_{cls} + \frac{1}{m}\sum_{i=1}^m \mathcal{L}_{rec1}^{(i)} + \lambda_1 \frac{1}{m}\sum_{i=1}^m \mathcal{L}_{rec2}^{(i)} + \lambda_2 \frac{1}{m}\sum_{i=1}^m \mathcal{L}_{rso}^{(i)}\]

where \(\mathcal{L}_{cls}\) is the binary cross-entropy loss, \(m\) denotes the total number of adapted self-attention weight matrices, and default loss weights are set to \(\lambda_1 = 0.1\), \(\lambda_2 = 0.1\), with the stable rank scaling factor \(\tau = 0.1\). The model is trained on a single NVIDIA GeForce RTX 3090 GPU using the Adam optimizer with a batch size of 16, a fixed learning rate of \(2 \times 10^{-4}\), and a weight decay of \(5 \times 10^{-4}\).

Key Experimental Results

Main Results

The model is trained solely on the lightly compressed FaceForensics++ (FF++ c23) dataset and directly evaluated on four unseen target benchmarks (Celeb-DF / CDF, DFDC, DFDCP, and DFD) using CLIP ViT-L/14 as the backbone. Frame-level and video-level AUC metrics are summarized below:

Evaluation Level Method Venue CDF DFDC DFDCP DFD Average AUC
Frame-Level Face X-ray CVPR'20 0.679 0.633 0.694 0.766 0.693
Frame-Level F3Net ECCV'20 0.735 0.702 0.735 0.798 0.743
Frame-Level SRM CVPR'21 0.755 0.704 0.741 0.812 0.752
Frame-Level RECCE CVPR'22 0.732 0.713 0.742 0.812 0.750
Frame-Level UCF ICCV'23 0.753 0.719 0.759 0.807 0.760
Frame-Level LoRA ICCV'23 0.838 0.717 – 0.817 –
Frame-Level LSDA CVPR'24 0.830 0.736 0.815 0.880 0.815
Frame-Level UDD AAAI'25 0.869 0.758 0.856 0.910 0.848
Frame-Level Effort ICML'25 0.901 0.798 – 0.923 –
Frame-Level SRRA (Ours) ECCV'26 0.911 0.856 0.890 0.947 0.901
Video-Level F3Net ECCV'20 0.789 0.718 0.749 0.844 0.775
Video-Level SBI CVPR'22 0.932 0.724 0.862 0.976 0.874
Video-Level UCF ICCV'23 0.837 0.742 0.770 0.867 0.804
Video-Level UDD AAAI'25 0.931 0.812 0.881 0.955 0.895
Video-Level StA CVPR'25 0.947 0.843 0.909 0.965 0.916
Video-Level Effort ICML'25 0.956 0.843 0.909 0.965 0.918
Video-Level SRRA (Ours) ECCV'26 0.968 0.889 0.923 0.980 0.940

On the most challenging DFDC benchmark, SRRA achieves a frame-level AUC of 0.856 (+5.8% over Effort's 0.798) and a video-level AUC of 0.889 (+4.6% over Effort's 0.843), demonstrating remarkable cross-domain robustness under heavy distribution shifts.

Ablation Study

1. Component Contribution (SRD, REC, and RSO)

Config SRD (Adaptive Stable Rank) REC (Residual Energy Constraint) RSO (Residual Orthogonalization) CDF DFDC DFDCP DFD Average AUC
1 (Baseline) × × × 0.840 0.801 0.800 0.888 0.832
2 ✓ × × 0.894 0.840 0.849 0.937 0.880
3 ✓ ✓ × 0.901 0.843 0.888 0.943 0.894
4 ✓ × ✓ 0.889 0.832 0.866 0.936 0.881
5 (Full Model) ✓ ✓ ✓ 0.911 0.856 0.890 0.947 0.901

2. Comparison Across Fine-Tuning Strategies and Residual Ranks

Fine-Tuning Strategy Residual Rank \(r_{res}\) CDF DFDCP DFD Average AUC
LoRA 1 0.857 0.851 0.942 0.883
LoRA 4 0.854 0.865 0.939 0.886
LoRA 16 0.865 0.855 0.941 0.887
LoRA 64 0.840 0.844 0.934 0.872
Effort 1 0.881 0.865 0.933 0.893
Effort 4 0.892 0.865 0.931 0.896
Effort 16 0.891 0.859 0.941 0.897
Effort 64 0.879 0.856 0.940 0.891
SRRA (Ours) Adaptive 0.911 0.890 0.947 0.916

Key Findings

  • SRD provides the foundational performance jump: Introducing adaptive stable-rank decomposition alone improves average AUC from 0.832 to 0.880 (+4.8%), proving that aligning fine-tuning budgets with depth-wise representational demands is critical to combat shortcut learning.
  • REC and RSO act synergistically: Applying RSO in isolation without energy bounds risks amplifying residual noise, yielding negligible gains; however, when constrained by REC's nuclear-norm regularization, RSO promotes orthogonal subspace exploration without instability, pushing the average AUC to 0.901.
  • Fixed ranks suffer from under- or over-allocation: Both LoRA and Effort degrade when \(r_{res}\) increases to 64 due to excessive perturbation in shallow layers. SRRA's dynamic allocation outperforms every tested fixed rank configuration across all test domains.

Highlights & Insights

  • Theoretical generalization metric repurposed as an adaptive PEFT allocator: Grounded in spectral generalization bounds, SRRA utilizes stable rank as an intrinsic measure of matrix complexity, eliminating heuristic manual grid-searches for layer-wise ranks.
  • Decoupled intra-residual orthogonalization: In contrast to previous works enforcing rigid cross-orthogonality between principal and residual components, SRRA restricts orthogonality regularization to the residual bases, preserving full adaptation flexibility while preventing directional collapse.
  • Model-agnostic and extensible: The principle of stable-rank-driven residual allocation is not restricted to deepfake detection; it offers an intuitive paradigm for cross-domain PEFT in multimodal foundation models and low-resource domain adaptation.

Limitations & Future Work

  • Unexplored extension to non-square projection matrices: The current implementation targets square linear matrices within ViT self-attention blocks. Formulations for non-square feed-forward layers (FFN) or cross-attention projections require further investigation.
  • Singular value computation overhead during training: Enforcing the nuclear-norm penalty on residual singular values introduces slight computational overhead during backward propagation compared to standard low-rank products, suggesting the potential utility of polynomial spectral approximations.
  • vs LoRA (ICLR'22): LoRA updates all layers using a globally identical rank without considering layer-depth semantics. SRRA adaptively matches rank to layer-wise complexity via stable rank and freezes SVD principal components to preserve foundational priors.
  • vs Effort (ICML'25): Effort partitions weights via SVD but retains a fixed rank across all layers and constrains the residual component to be orthogonal to the principal component. SRRA dynamically adjusts rank by layer and applies orthogonalization exclusively within residual singular vectors, providing richer discriminative modeling of complex forgery cues.

Rating

  • Novelty: ⭐⭐⭐⭐☆ (Insightful observation of layer-wise singular value distributions in ViTs, skillfully translating stable rank into adaptive fine-tuning and regularization)
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ (Comprehensive cross-dataset benchmarks at both frame and video levels, accompanied by thorough ablation and hyper-parameter sensitivity studies)
  • Writing Quality: ⭐⭐⭐⭐⭐ (Clean structure, solid mathematical formulations, and compelling empirical validation)
  • Value: ⭐⭐⭐⭐☆ (Provides substantial generalization gains in deepfake detection and offers valuable architectural insights for adaptive PEFT across foundation models)