Skip to content

LoRC: Detecting AI-Generated Images via Low-Rank Collapse in the Semantic-Residual Space

Conference: ECCV 2026
Paper: ECCV Official
Area: AIGC Detection
Keywords: AIGC Detection, Low-Rank Collapse, Semantic-Residual Subspace, Vision Foundation Models, Zero-Shot Generalization

TL;DR

Addressing the failure of semantic forensics caused by modern generators over-aligning with macroscopic semantics, this paper identifies an architecture-agnostic low-rank collapse signature in the orthogonal semantic-residual subspace induced by terminal decoding bottlenecks, proposing the LoRC framework with orthogonal decomposition, low-rank attention, and subspace separation to achieve 97.0% zero-shot accuracy across 39 unseen generators.

Background & Motivation

With the rapid evolution of diffusion models and autoregressive image generators, AI-generated images (AIGIs) have achieved remarkable fidelity, closely matching authentic photographs in macroscopic layout, compositional structure, and semantic coherence. Early forensic detectors heavily relied on low-level microscopic artifacts, high-frequency spectral traces, or upsampling discrepancies; however, these cues are fragile and easily obliterated by lossy post-processing such as JPEG compression or smooth diffusion rendering. While recent detectors have embraced vision foundation models (such as CLIP and DINO) to leverage rich prior representations, these foundation backbones are primarily pre-trained to capture macroscopic semantics—answering "what is depicted" rather than identifying generative synthesis flaws. In modern coarse-to-fine generation pipelines, global semantics are stabilized early, causing macroscopic semantic discrepancies between authentic and synthetic images to diminish substantially.

This asymmetric progress induces a critical tension: as discriminative power along the dominant semantic direction approaches exhaustion, decisive forensic cues are systematically pushed into the orthogonal semantic-residual subspace, where subtle non-semantic visual discrepancies such as texture naturalness and local physical consistency reside. Nevertheless, existing detectors either rely on fragile pixel-level statistical alignments or feed global semantic embeddings directly into linear classifiers, inevitably suppressing weak residual forensic signals under strong semantic interference. Whether a universal, cross-architectural forensic signature exists within this residual manifold remains a central open challenge for generalized zero-shot detection.

This work attacks the problem from the geometric manifold of modern generative architectures: across fundamentally distinct generative paradigms—including diffusion models, flow-matching frameworks, and autoregressive models—all pipelines share a terminal decoding or rendering stage (such as a VAE decoder or a discrete de-tokenizer) that maps latent tokens to pixel space, inevitably acting as a severe information bottleneck. Empirical singular value and principal component analyses on semantic-residual variations confirm that synthetic reconstructions undergo severe rank degeneracy and structural flattening. Core idea: orthogonally decouple visual representations into dominant semantic directions and semantic-residual subspaces, utilize a low-rank attention bottleneck to capture the architecture-agnostic low-rank collapse signature induced by terminal generative decoders, and enlarge the manifold margin via subspace separation loss to achieve robust zero-shot detection.

Method

Overall Architecture

The LoRC pipeline consists of three tightly coupled stages: semantic decomposition, low-rank attention modeling, and geometric subspace separation. Given an input image, patch tokens and a global [CLS] token are extracted using a frozen DINOv3 ViT-H+/16 backbone. The normalized frozen [CLS] token serves as an instance-specific semantic anchor, onto whose orthogonal complement all patch tokens are projected to isolate semantic residuals. The isolated residuals are then fed into a Low-Rank Attention Block, which enforces self-attention within a rank-r bottleneck to capture coherent low-rank forgery cues while suppressing unconstrained feature noise. Finally, the pooled low-rank residual features are concatenated with the frozen [CLS] token and classified via a linear head, optimized jointly with an auxiliary Subspace Separation Loss that maximizes the geometric distance between real and synthetic residual covariance descriptors.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image I"] --> B["DINOv3 Backbone<br/>Extract [CLS] Token & Patch Tokens"]
    B --> C["Semantic Decomposition<br/>Project onto Orthogonal Complement"]
    C --> D["Low-Rank Attention Block<br/>Rank-r Bottleneck Suppresses Noise"]
    D --> E["Feature Concatenation & Head<br/>Joint Classification with [CLS]"]
    E --> F["Real / Fake Decision"]
    C -.-> G["Subspace Separation Loss (SSL)<br/>Minimize Inner Product of Covariances"]

Key Designs

1. Semantic Decomposition: Projecting onto the Orthogonal Complement of the Semantic Anchor To address the issue where dominant semantic components overpower weak forensic residuals, this design explicitly decouples visual representations via geometric orthogonal projection. Given an input image \(I\), deep visual features comprising the global [CLS] token \(c \in \mathbb{R}^D\) and spatial patch tokens \(X \in \mathbb{R}^{N \times D}\) are extracted from a frozen vision foundation model. The frozen [CLS] token is normalized into a unit semantic anchor \(\hat{c} = c / \|c\|_2\), and the patch feature map is decomposed into its dominant semantic projection \(X_{sem}\) and orthogonal residual component \(X_{res}\): $\(X_{res} = X \left(I - \hat{c}\hat{c}^\top\right)\)$ This sample-specific projection removes label-dependent bias and eliminates collinear interference caused by macroscopic scene content, shifting real/fake discrimination from a shared semantic space into an intrinsic forensic residual manifold.

2. Low-Rank Attention Block: Capturing Degenerate Structural Geometry in a Rank-r Bottleneck Because vision foundation model embeddings are inherently high-dimensional (\(D \gg 1000\) in ViT-H), the subtle low-rank collapse signature is easily obscured by extraneous feature variance and high-dimensional noise. To amplify sensitivity to this structural degeneracy, a rank-constrained self-attention module is introduced, restricting the attention manifold to a compressed subspace \(r \ll D\) (with \(r=32\) by default). Residual tokens \(X_{res}\) are projected via \(W_Q, W_K, W_V \in \mathbb{R}^{D \times r}\) into Query, Key, and Value representations in \(\mathbb{R}^r\), performing self-attention within this compressed manifold: $\(A = \text{Softmax}\left(\frac{Q K^\top}{\sqrt{r}}\right) V\)$ The aggregated low-rank features \(A \in \mathbb{R}^{N \times r}\) are subsequently mapped back to the original feature dimension via an output projection \(W_O \in \mathbb{R}^{r \times D}\) yielding \(Y = A W_O\). Restricting the degree of freedom forces the attention mechanism to focus exclusively on dominant collapsed directions where the real/fake structural gap is maximized, naturally filtering out stochastic high-frequency perturbations.

3. Subspace Separation Loss: Enlarging the Covariance Margin Between Full-Rank and Collapsed Manifolds Standard binary cross-entropy (BCE) offers only scalar classification guidance and fails to explicitly constrain the geometric structure of the residual space. To enforce the geometric hypothesis that authentic residuals span a high-dimensional full-rank manifold while synthetic residuals collapse onto a low-rank subspace, an auxiliary Subspace Separation Loss (SSL) based on normalized covariance descriptors is incorporated. Within each mini-batch, patch residuals are partitioned by ground-truth labels into \(R_{real}\) and \(R_{fake} \in \mathbb{R}^{B N \times D}\), from which normalized covariance matrices are computed: $\(P = \frac{R^\top R}{\|R^\top R\|_F}\)$ Minimizing the Frobenius inner product between the descriptors decorrelates the principal directions of real and fake residual subspaces: $\(\mathcal{L}_{SS} = \langle P_{real}, P_{fake} \rangle = \text{Tr}\left(P_{real}^\top P_{fake}\right)\)$ The overall optimization objective is \(\mathcal{L} = \mathcal{L}_{BCE} + \lambda_{SS} \mathcal{L}_{SS}\) (with \(\lambda_{SS}=0.1\)), compelling the network to explicitly push synthetic representations into a distinct, degenerate subspace orthogonal to natural image variations.

Loss & Training

The architecture employs DINOv3 ViT-H+/16 as its visual backbone. To preserve pre-trained natural manifold priors while preventing catastrophic forgetting, the backbone parameters are kept frozen and adapted exclusively via LoRA (rank 16, scaling factor \(\alpha=16\)). Training strictly follows the DDA protocol using MS-COCO paired with its Stable Diffusion 2.1 VAE reconstructions, ensuring that semantic content is strictly matched so that the model learns only the structural collapse induced by the decoding bottleneck. Training uses the Adam optimizer with a learning rate of \(10^{-4}\), a batch size of 64, and 4 gradient accumulation steps.

Key Experimental Results

Main Results

LoRC was comprehensively evaluated across 7 representative benchmarks covering standard datasets, unconstrained in-the-wild scenarios, and cutting-edge generative suites. All testing adhered to the DDA protocol with JPEG 96 compression applied to synthetic images to eliminate format-induced shortcuts. Balanced Accuracy (B.Acc, %) is reported as the primary metric.

Benchmark Category Dataset NPR (CVPR'24) UnivFD (CVPR'23) DRCT (ICML'24) DDA (NeurIPS'25) LoRC (Ours)
Standard GenImage 51.5 64.1 84.7 91.7 97.9
Standard DRCT-2M 37.3 61.8 90.5 98.1 99.3
Standard Synthbuster 50.0 67.8 81.3 90.1 99.9
Standard AIGCDetectionBenchmark 53.1 72.5 81.4 87.8 96.3
In-the-Wild Chameleon 59.9 50.7 56.6 82.4 92.6
In-the-Wild WildRF 63.5 55.3 50.6 90.3 97.6
Recent Generators T2I-CoReBench (39 generators) 60.0 10.8 48.6 91.1 97.0
Overall Average All 7 Benchmarks Avg. 53.6 54.7 71.0 90.2 97.2 (+7.0)

On Synthbuster across 9 popular diffusion architectures, LoRC achieves near-perfect cross-generator generalization:

Method DALL·E 2 DALL·E 3 Firefly GLIDE Midjourney SD 1.4 SDXL Average Accuracy
NPR (CVPR'24) 51.1 49.3 46.5 48.5 52.8 51.8 52.8 50.0
UnivFD (CVPR'23) 83.5 47.4 89.9 53.3 52.5 69.9 68.0 67.8
DRCT (ICML'24) 77.2 86.6 84.1 82.6 73.7 86.6 71.3 81.3
DDA (NeurIPS'25) 86.3 90.0 91.9 76.5 93.5 92.7 93.5 90.1
LoRC (Ours) 99.6 99.9 99.6 99.8 100.0 100.0 100.0 99.9 (+9.8)

Ablation Study

To verify the contribution of each design component, a cumulative ablation study was performed across standard, in-the-wild, and zero-shot benchmarks:

Configuration Standard Benchmarks In-the-Wild Benchmarks T2I-CoReBench (Zero-Shot) Overall Average
Baseline (DINOv3 ViT-H + LoRA) 97.9 89.4 94.0 93.8
+ Semantic Decomposition 97.2 86.9 98.1 94.1
+ Low-Rank Attention 96.9 90.2 96.0 94.4
+ Subspace Separation Loss (Full Model) 98.1 95.2 97.0 96.8

Key Findings

  • Synergistic Modular Interactions: Semantic Decomposition provides a dramatic 4.1% boost on T2I-CoReBench (jumping from 94.0% to 98.1%), demonstrating that neutralizing dominant macroscopic semantics is essential for modern generators with strong semantic alignment. Low-Rank Attention recovers performance on in-the-wild benchmarks from 86.9% to 90.2% by filtering high-dimensional noise, while Subspace Separation Loss further expands this margin to 95.2%, leading to the highest overall average of 96.8%.
  • Hyperparameter Sensitivity Trade-offs: The Low-Rank Attention bottleneck achieves optimal performance at rank \(r=32\) (over-compression at \(r=8\) discards structural information, while \(r=64\) admits redundant feature fluctuations). The Subspace Separation Loss weight reaches its sweet spot at \(\lambda_{SS}=0.1\), and LoRA rank 16 provides the best trade-off between adaptation capacity and parameter efficiency.
  • Intrinsic Robustness to Pixel Perturbations: When subjected to severe post-processing corruptions—such as JPEG compression (quality factors varying from 100 down to 60), bilinear resizing scales (0.5× to 2.0×), and Gaussian blurring kernels (\(\sigma\) from 0 to 2.0)—LoRC demonstrates near-flat accuracy curves, heavily outperforming baselines that deteriorate under lossy perturbations.

Highlights & Insights

  • From Superficial Artifacts to Manifold Geometry: Rather than relying on fragile frequency spikes or upsampling artifacts, this paper presents the first systematic study identifying the universal "low-rank collapse" in semantic-residual space induced by terminal decoding bottlenecks.
  • Plug-and-Play Orthogonal Operator: By using the pre-trained [CLS] token of DINOv3 as an adaptive instance-specific anchor, semantic interference is eliminated in a completely unsupervised, parameter-free manner without requiring auxiliary semantic removal networks.
  • Exceptional Zero-Shot Generalization: Trained solely on MS-COCO paired with SD 2.1 VAE reconstructions, LoRC generalizes seamlessly to 39 diverse generative architectures—including FLUX.2, SD-3.5, Infinity-8B, Janus-Pro-7B, and commercial closed-source APIs (Seedream 4.5, Gemini 2.0 Flash, GPT-4o).

Limitations & Future Work

  • Author-Admitted Limitations: The method relies on the premise that synthetic images undergo a final decoding or rendering bottleneck (e.g., VAE or discrete de-tokenizer). Should future generative paradigms synthesize high-resolution images entirely in end-to-end continuous pixel space without downsampling bottlenecks, the low-rank collapse signature may be attenuated.
  • Potential Improvement Directions: The current formulation utilizes a single global [CLS] token as the semantic anchor, which may experience subtle semantic leakage in complex scenes with multiple disparate subjects. Future work could explore multi-center semantic anchors or hierarchical subspace decomposition to refine the residual space.
  • vs UnivFD [CVPR'23] / Simplicity Prevails [2026]: These approaches directly feed pre-trained foundation model embeddings into binary classifiers; however, as modern generators achieve high semantic realism, classifiers suffer from severe semantic interference. LoRC eliminates this bias by projecting features onto the orthogonal complement of the semantic anchor.
  • vs DDA [NeurIPS'25]: DDA requires complex dual pixel-frequency dataset alignment during training. In contrast, LoRC addresses the generative decoding bottleneck at the architectural manifold level, achieving higher cross-model generalization (+7.0% overall average) with a cleaner architectural design.
  • vs NS-Net [2025]: While NS-Net constructs a semantic null-space, it lacks explicit geometric modeling of residual structures. LoRC identifies the low-rank collapse within residual space and actively reinforces the geometric margin via a low-rank attention bottleneck and covariance subspace separation.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Elegant geometric perspective uncovering the universal low-rank collapse signature in the semantic-residual subspace.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluation across 7 benchmarks, 39 cutting-edge unseen generators, and extensive robustness evaluations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Lucid theoretical motivation, tightly coherent geometric formulations, and compelling empirical validations.
  • Value: ⭐⭐⭐⭐⭐ Provides an architecture-agnostic foundation for the next generation of generalizable AIGI detection.