HLRAD: High-dimensional Latent Representation for Unified Anomaly Detection¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Code: https://github.com/xxx
Area: Object Detection
Keywords: Unified Anomaly Detection, High-dimensional Latent Representation, Diffusion Transformer, Selective Gated Fusion, Dimension Expansion
TL;DR¶
HLRAD tackles multi-class unified unsupervised anomaly detection by challenging the conventional low-dimensional bottleneck paradigm, introducing a high-dimensional latent reconstruction framework equipped with a frozen Diffusion Transformer semantic branch and selective gated fusion to eliminate compression information loss and block identity shortcuts.
Background & Motivation¶
Unsupervised anomaly detection aims to learn the underlying distribution of normal samples from training data and accurately identify and segment arbitrary anomalies deviating from this distribution during inference. In industrial manufacturing inspection, conventional methods typically train dedicated models for individual object categories (class-separated). However, as category counts and defect patterns grow in modern factories, maintaining hundreds of distinct models incurs prohibitive computational and deployment costs. Unified anomaly detection resolves this dilemma by training a single unified model across all categories simultaneously. Nevertheless, under this unified regime, existing reconstruction-based methods face two critical vulnerabilities: the identity shortcut learning problem, where unconstrained autoencoders tend to learn an identity mapping that directly passes and reconstructs unseen anomalies during testing; and severe degradation in representation semantic fidelity, caused by compressing high-resolution spatial features into ultra-compact latent spaces.
Prior state-of-the-art frameworks, such as DiAD, HVQ-Trans, and PMAD, heavily rely on continuous or discrete tokenizers (e.g., VAEs or VQ codebooks) to compress pixel space into low-dimensional latent bottlenecks. While this compression was widely assumed to promote cross-category structural abstraction, it irreversibly discards fine-grained textural cues and subtle geometric boundaries, causing decoders to fail when reconstructing slender cracks or localized defects. Meanwhile, methods operating within the raw feature dimensions of pretrained encoders (e.g., ViTAD, Dinomaly) constrain the decoder strictly within fixed teacher dimensions (\(d_t=1024\)), leaving it vulnerable to the domain shift between natural pretraining datasets and downstream industrial components. This paper contends that latent space dimensional compression is not an inherent prerequisite for anomaly detection, but rather the primary information bottleneck degrading reconstruction fidelity. Provided that direct input-to-output shortcuts are effectively severed, expanding the latent space to a higher dimension grants the decoder significantly greater capacity to model complex, multi-modal normal distributions across diverse categories.
Building upon this insight, the authors introduce a counter-intuitive reconstruction paradigm based on latent dimension expansion rather than compression. Core idea: abandon low-dimensional latent bottlenecks in favor of a high-dimensional latent representation, utilizing a frozen Diffusion Transformer to extract rich semantic priors, adaptively merging high-dimensional semantics with local textures via selective gated fusion, and reconstructing features within an expanded latent space via an expansion MLP and wide decoder with stochastic perturbations to prevent shortcut learning.
Method¶
Overall Architecture¶
The overall architecture of HLRAD operates through a coordinated four-stage pipeline: multi-scale feature extraction with a high-dimensional semantic branch, selective gated fusion between semantic and local features, dimension-expanding projection with high-dimensional decoding, and reverse distillation anomaly scoring. Given an industrial inspection image, a frozen DINOv3-ViT-L encoder first extracts multi-layer localized patch tokens. In parallel, extracted features are projected into a frozen Diffusion Transformer (DiT) encoder to re-encode global structural semantics in a higher-dimensional space (\(d_s = 1152\)). The self-attention map from the final encoder layer then identifies structurally salient patches, guiding a Selective Gated Fusion (SGF) module to inject semantic features adaptively while preserving native fine textures elsewhere. The fused tokens are mapped into an expanded latent space (\(d_d = 2048\)) via a Dimension-expanding MLP regularized with triple dropout, where a shallow yet wide Transformer decoder reconstructs multi-layer teacher representations. Finally, pixel-wise cosine distances between teacher and decoded features yield dense anomaly localization heatmaps.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image & Multi-layer Extraction<br/>DINOv3-ViT-L multi-layer features & grouping"] --> B["DiT Latent Semantic Branch<br/>High-dimensional re-encoding & domain drift mitigation"]
A --> C["Saliency-guided Selective Gated Fusion<br/>Attention saliency guided channel-wise adaptive gating"]
B --> C
C --> D["Dimension-expanding Projection & Wide Decoding<br/>Triple dropout MLP + wide latent space reconstruction"]
D --> E["Reverse Distillation Cosine Metric<br/>Pixel-level and image-level anomaly scoring"]
Key Designs¶
1. DiT Latent Semantic Branch: High-dimensional Denoising Prior to Mitigate Industrial Domain Shift
Pretrained visual backbones suffer from systematic feature drift when applied to industrial components, as their training sets consist primarily of natural scenes, while single-layer features lack the capacity to express intricate semantic relations across multiple categories. HLRAD integrates a frozen 28-block Diffusion Transformer encoder (\(D_e=28\), hidden dimension \(d_s=1152\)) pretrained from Representation Autoencoders (RAE). Extracted patch tokens \(\bar{\mathbf{F}} \in \mathbb{R}^{N \times d_t}\) are mapped into the DiT high-dimensional latent space through a learnable linear layer \(\phi_{\text{in}}\), added to sinusoidal positional embeddings \(\mathbf{E}_{\text{pos}}\), and processed through the frozen DiT blocks modulated by learnable unconditional embeddings \(\mathbf{c} \in \mathbb{R}^{d_s}\) and Rotary Position Embeddings (RoPE):
A linear projection \(\phi_{\text{out}}\) then maps the final representations back to dimension \(d_t\), producing semantic features \(\mathbf{S} \in \mathbb{R}^{N \times d_t}\). Crucially, the DiT parameters remain completely frozen, training only \(\phi_{\text{in}}\), \(\phi_{\text{out}}\), and \(\mathbf{c}\). Operating at \(d_s > d_t\) equips the semantic branch with enhanced capacity to capture hierarchical structural abstractions. Furthermore, because DiT was trained as a generative denoising backbone, it inherently exhibits robustness against distributional perturbations, treating the out-of-distribution industrial domain gap as input noise and suppressing feature drift without requiring domain-specific fine-tuning.
2. Saliency-guided Selective Gated Fusion: Attention-driven Semantic Injection with Local Texture Preservation
High-dimensional semantic features \(\mathbf{S}\) capture global structural layouts, whereas the original features \(\bar{\mathbf{F}}\) retain the high-frequency local textures and precise spatial boundaries extracted by DINOv3. Uniformly substituting all features with semantic representations would obliterate subtle surface texture cues (e.g., hairline cracks or minor discoloration), while omitting semantic features entirely hinders the detection of global structural defects. To balance both requirements, the Selective Gated Fusion (SGF) module leverages multi-head self-attention maps \(\mathbf{A} \in \mathbb{R}^{N_h \times N \times N}\) from the encoder target layer to compute patch-wise semantic saliency scores \(\alpha_i = \frac{1}{N_h} \sum_{h=1}^{N_h} \mathbf{A}^{(h)}[i]\). High-attention locations consistently mark object silhouettes, structural components, and semantic boundaries.
During training, a gating ratio \(r\) is uniformly sampled from \([0, r_{\max}]\) (\(r_{\max}=0.5\) by default), and the \(k = \lfloor r \cdot N \rfloor\) patches with the highest saliency scores form the gating set \(\mathcal{K}\). For each selected token \(i \in \mathcal{K}\), channel-wise fusion gates \(\mathbf{g}_i \in [0, 1]^{d_t}\) are derived via a linear layer and Sigmoid activation:
Positions with lower saliency (such as plain backgrounds and homogeneous textures) bypass the fusion and retain pure textural representations. Stochastic sampling of the gating ratio serves as a structured regularizer, preventing the downstream decoder from developing co-dependencies on semantic features while continuously exposing it to semantically enriched normal representations.
3. Dimension-expanding Projection & Wide Decoding: Eliminating Bottlenecks and Breaking Identity Shortcuts
To unlock sufficient representational capacity for multi-class distributions without suffering from shortcut learning, HLRAD pairs a Dimension-expanding MLP with a shallow yet wide Transformer decoder. Fused features \(\hat{\mathbf{F}} \in \mathbb{R}^{N \times d_t}\) are projected into an expanded latent space (\(d_d=2048\), with an intermediate hidden dimension of \(4d_d=8192\)) using a two-layer MLP equipped with triple dropout regularization:
Applying dropout (\(p=0.3 \sim 0.4\)) pre-projection, post-activation, and post-projection stochastically disrupts feature continuity, completely severing direct identity mapping paths from input to output. Consequently, the wide decoder cannot memorize trivial pixel-to-pixel shortcuts and is compelled to reconstruct normal features using learned global prototypes. The subsequent decoder consists of \(D_d=8\) Transformer blocks operating natively at \(d_d=2048\) with 32 attention heads. A DropHead strategy (\(p_h=0.1\)) drops attention heads at random during training to prevent head co-adaptation. Finally, grouped linear projection layers with LayerNorm project the multi-layer decoder states back to \(d_t\) to align with the teacher features.
Loss & Training¶
Following the reverse distillation paradigm, eight evenly spaced layers \(\{4, 6, 8, 10, 12, 14, 16, 18\}\) from the frozen DINOv3 backbone are grouped into 2 feature aggregates. The training objective minimizes the global cosine distance between the projected decoder representations and the frozen teacher features:
A stop-gradient operator is applied to the teacher branch. The network is optimized using AdamW with AMSGrad, weight decay \(10^{-4}\), and a cosine learning rate scheduler decaying to \(2 \times 10^{-4}\) after a 100-step warmup. Training runs for 20k iterations across 8 NVIDIA A100 GPUs. During inference, pixel-wise cosine distances at each patch position directly constitute the anomaly heatmap, which is upsampled via bilinear interpolation to match the native input resolution (\(448 \times 448\)).
Key Experimental Results¶
Main Results¶
HLRAD is evaluated under the unified multi-class setting across MVTec-AD, VisA, and the large-scale multi-view benchmark Real-IAD, competing against state-of-the-art reconstruction-based and embedding-based methods.
| Dataset | Metric | HLRAD (Ours) | Prev. SOTA (Method) | Gain |
|---|---|---|---|---|
| MVTec-AD | Image-AUROC | 99.7 | 99.7 (INP-Former) | Tied |
| Image-AP | 99.9 | 99.9 (INP-Former) | Tied | |
| Image-F1max | 99.3 | 99.2 (INP-Former) | +0.1% | |
| Pixel-AUROC | 98.6 | 98.5 (INP-Former) | +0.1% | |
| Pixel-AP | 71.3 | 71.0 (INP-Former) | +0.3% | |
| Pixel-F1max | 70.7 | 69.7 (INP-Former) | +1.0% | |
| VisA | Image-AUROC | 99.0 | 98.9 (INP-Former) | +0.1% |
| Image-AP | 99.0 | 99.0 (INP-Former) | Tied | |
| Image-F1max | 96.6 | 96.6 (INP-Former) | Tied | |
| Pixel-AUROC | 99.1 | 98.9 (INP-Former) | +0.2% | |
| Pixel-AP | 55.5 | 53.2 (Dinomaly) | +2.3% | |
| Pixel-F1max | 57.5 | 55.7 (Dinomaly) | +1.8% | |
| Real-IAD | Image-AUROC | 90.1 | 90.5 (INP-Former) | -0.4% |
| Image-AP | 87.5 | 88.1 (INP-Former) | -0.6% | |
| Image-F1max | 81.1 | 81.5 (INP-Former) | -0.4% | |
| Pixel-AUROC | 99.0 | 99.0 (INP-Former) | Tied | |
| Pixel-AP | 47.7 | 47.5 (INP-Former) | +0.2% | |
| Pixel-F1max | 50.5 | 50.3 (INP-Former) | +0.2% |
Ablation Study¶
1. Component Contribution Analysis (MVTec-AD and VisA Benchmarks) Ablation evaluating the cumulative impact of the Semantic Encoder (Senc), Selective Gated Fusion (SGF), Dimension-expanding MLP (DM), and High-dimensional Decoder (HD):
| Configuration | Senc | SGF | DM | HD | MVTec P-AP | MVTec P-F1max | VisA P-AP | VisA P-F1max | Note |
|---|---|---|---|---|---|---|---|---|---|
| Baseline | - | - | - | - | 69.3 | 69.2 | 53.2 | 55.7 | Native reconstruction at native dimensions |
| + Senc | ✓ | - | - | - | 70.4 | 69.9 | 54.0 | 56.5 | Adds DiT semantic prior; strong initial boost |
| + Senc + DM | ✓ | - | ✓ | - | 71.0 | 70.2 | 54.7 | 56.7 | Triple dropout MLP severs identity shortcuts |
| + Senc + DM + HD | ✓ | - | ✓ | ✓ | 71.2 | 70.5 | 55.3 | 57.2 | Expands decoder workspace to \(d_d=2048\) |
| + Senc + SGF + DM | ✓ | ✓ | ✓ | - | 71.0 | 70.3 | 55.1 | 57.1 | SGF dynamically balances local & global cues |
| Full HLRAD | ✓ | ✓ | ✓ | ✓ | 71.3 | 70.7 | 55.5 | 57.5 | Synergistic combination achieves peak performance |
2. Tokenizer Representation Fidelity on Normal Training Samples (MVTec-AD) Direct comparison between the conventional continuous latent compression tokenizer (DiAD VAE) and HLRAD's dimension-expanding tokenizer under identical spatial resolutions:
| Reconstruction Metric | DiAD VAE Tokenizer (Compressed) | HLRAD Tokenizer (Expanded) | Relative Error Reduction |
|---|---|---|---|
| L1 Loss | 0.0633 | 0.0213 | -66.4% |
| L2 Loss | 0.0100 | 0.0010 | -90.0% |
| Cosine Distance | 0.0280 | 0.0023 | -91.8% |
Key Findings¶
- High-dimensional latent space significantly sharpens pixel-level localization: Across all benchmarks, HLRAD achieves its most substantial gains in fine-grained localization metrics. On VisA, Pixel-AP reaches 55.5% (+2.3% over the prior state of the art). Varying the decoder dimension \(d_d\) from 768 to 2048 confirms that expanding latent width directly drives localization precision, effectively eliminating the fuzzy boundary artifacts typical of low-dimensional reconstructions.
- Over 90% reduction in normal reconstruction error: Evaluating the reconstruction fidelity on MVTec-AD normal training sets demonstrates a 90.0% drop in L2 loss and a 91.8% drop in cosine distance compared to DiAD's fine-tuned VAE tokenizer. This validates the core thesis that low-dimensional compression was the primary bottleneck throttling normal pattern modeling.
- Frozen backbones outperform fine-tuning: Ablations on the semantic branch confirm that freezing the DiT encoder yields superior performance compared to end-to-end fine-tuning (Pixel-AP 71.3% vs. 70.9%). Training the semantic branch causes it to co-adapt with the decoder, dissolving the semantic bottleneck required to separate abnormal deviations from normal patterns.
Highlights & Insights¶
- Paradigm shift from compression to dimension expansion: Re-examines the longstanding assumption that anomaly detection autoencoders must compress features into low-dimensional bottlenecks, demonstrating that expansion combined with stochastic dropout achieves superior reconstruction fidelity while preventing shortcut learning.
- Saliency-conditioned semantic-texture decoupling: By using self-attention saliency maps to gate high-dimensional semantic injection only into complex structural regions, HLRAD preserves raw textural features on homogeneous surfaces, simultaneously resolving subtle texture defects and macro-structural anomalies.
- Generative denoising priors as domain-shift buffers: Capitalizes on frozen Diffusion Transformers pretrained on generative tasks, utilizing their inherent denoising capabilities to absorb out-of-distribution feature drift when adapting natural vision backbones to industrial imagery.
Limitations & Future Work¶
- Image-level discrimination under extreme category diversity: On Real-IAD (comprising 30 object classes across multi-view setups), HLRAD's image-level performance (I-AUROC 90.1, I-AP 87.5) falls slightly behind INP-Former (90.5, 88.1). When multi-class distributions become excessively diverse, modeling the normal distribution of every object in a single unified space becomes harder, occasionally under-representing rare normal angles.
- Computational footprint of wide Transformer decoding: Expanding the decoding space to \(d_d=2048\) with intermediate MLP widths of 8192 incurs higher peak memory and FLOPs during inference compared to compact distilled networks, presenting a trade-off for high-throughput industrial edge hardware.
- Future directions: Exploring dynamic class-conditioned prompt routing or mixture-of-experts in the high-dimensional latent space to maintain pixel-level precision while scaling image-level discrimination across hundreds of industrial categories.
Related Work & Insights¶
- vs Dinomaly / ViTAD: Both methods explore reverse distillation without pixel-level generative autoencoders, but remain confined to the raw feature dimensions of the teacher encoder (\(d_t=1024\)). HLRAD introduces a high-dimensional DiT semantic branch and expands the decoder to \(d_d=2048\), boosting Pixel-AP on VisA by 2.3% over Dinomaly.
- vs DiAD / HVQ-Trans: These methods compress features into VAE latent vectors or discrete VQ codebooks to enable generative editing, incurring over 90% reconstruction error on normal samples. HLRAD proves that stochastic perturbation during dimension expansion prevents shortcut learning without compromising feature fidelity.
- vs INP-Former: INP-Former mines internal normal prototypes within individual query images to alleviate cross-class confusion, excelling at global image classification. HLRAD focuses on latent topological capacity, achieving sharper defect boundary alignment and superior pixel-level localization (Pixel-F1max / AP).
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneers a dimension-expanding latent reconstruction paradigm that overhauls the traditional low-dimensional bottleneck constraint in unsupervised anomaly detection]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensively validated across MVTec-AD, VisA, and Real-IAD, accompanied by detailed tokenizer fidelity benchmarks and systematic architectural ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, rigorous logical progression, well-defined mathematical formulations, and thorough qualitative analysis]
- Value: ⭐⭐⭐⭐⭐ [Offers valuable practical and theoretical insights for unified multi-class industrial inspection and representation learning]