Skip to content

Beyond Normal References: Discriminative Few-Shot Anomaly Detection

Conference: NeurIPS2026
arXiv: 2605.23231
Code: https://github.com/mala-lab/IDEAL
Area: Object Detection (visual anomaly detection and localization)
Keywords: few-shot anomaly detection, normal variation suppression, intrinsic deviation, cross-attention, unseen anomalies

TL;DR

IDEAL uses a fixed set of few normal and masked anomalous references to suppress local normal variations, extract diverse intrinsic deviation directions, and score their projections, improving seen and unseen anomaly detection without target-domain retraining; under N1A1, it achieves 96.3 / 96.8 image-level / pixel-level AUROC on MVTecAD.

Background & Motivation

Few-shot anomaly detection typically uses a few normal images as references to determine whether a query departs from normal appearance. Specialist methods still fit a model for each target category, whereas generalist methods such as InCTRL and ResAD train once on auxiliary data and deploy across datasets using target-category normal references. Normal references describe normality but do not directly identify which departures matter: illumination, texture, and viewpoint differences can produce large feature distances without indicating a genuine defect.

Deployment may also provide one or two confirmed anomalous images, but anomalies are highly heterogeneous. A scratch reference does not represent holes, missing components, or contamination. Directly treating an anomalous reference as a matching template risks defining abnormality as resemblance to that particular image. The challenge is therefore not merely to add an anomalous nearest-neighbor bank, but to extract transferable discriminative evidence from very few anomalies without restricting detection to the defect types present in the references.

The paper formulates discriminative few-shot anomaly detection: auxiliary training data provide normal and anomalous samples with supervision, while target-domain inference uses fixed normal and anomalous references, with a one-time region annotation for anomalous references by default. References come from one randomly selected anomaly type and are reused for all test queries, which is closer to deployment than resampling references according to each query's anomaly type. Core Idea: rather than matching absolute anomalous appearance, suppress normal variations, learn diverse intrinsic deviation directions from anomalous-to-normal residuals, and detect anomalies through the alignment of query deviations with those directions.

Method

Overall Architecture

Inputs are a query image, few normal references, few anomalous references, and anomalous-reference masks; outputs are an anomaly map and an image-level anomaly score. After a frozen visual encoder extracts patch features, the Normal Variation Eraser (NVE) converts anomalous references and queries into denoised normal-relative residuals. The Intrinsic Deviation Encoder (IDE) extracts deviation directions only from the reference side, and Complementary Projection Scoring determines whether query deviations follow those directions.

The three stages are the Normal Variation Eraser, Intrinsic Deviation Encoder, and Complementary Projection Scoring. The first two do not consume query labels at inference: only training uses auxiliary-query region masks and image labels to supervise predictions and applies discriminability and orthogonality constraints to reference-side deviation directions. Target-domain deployment does not update model parameters, but still requires target references and their anomalous-region information.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    R["Normal/anomalous references<br/>Anomalous-reference masks"] --> F["Frozen encoder"]
    Q["Query image"] --> F
    F --> N["Normal Variation Eraser"]
    N -->|Reference deviations as values| I["Intrinsic Deviation Encoder"]
    F -->|Original anomalous features as keys| I
    R -->|Anomalous-reference masks| I
    N -->|Query deviations| P["Complementary Projection Scoring"]
    I -->|Intrinsic deviation directions| P
    P --> O["Anomaly map<br/>Top 1% aggregation"]
    S["Training only: auxiliary-query<br/>masks and image labels"] -.->|Focal / Dice / BCE| P
    I -.->|Training only: discriminability and orthogonality| I

Key Designs

1. Normal Variation Eraser: subtract normal content, then suppress variations that also occur in normal samples

The first step does not cluster anomalous references. Instead, for each patch, it finds the closest normal reference feature and subtracts that neighbor from the anomalous feature. Matching searches the normal reference feature bank without requiring strict positional alignment between images. Subtraction attempts to cancel the object's ordinary appearance and retain normal-relative departures, which offers a more transferable representation than absolute anomalous-image features.

Residuals can nevertheless contain normal appearance variations. Around each anomalous patch, NVE retrieves the 12 nearest normal features, centers those neighbors, and performs local PCA, retaining the leading 4 directions. This estimates a normal variation subspace around that local appearance rather than a single fixed set of principal components for the entire dataset. Illumination or texture differences among similar normal patches can be explained within this subspace.

With nearest normal neighbor \(f_{i,\min}^{n}\) and local orthonormal basis \(U_i\), the core operation is:

\[ f_{i,\mathtt{res}}^{a}=f_i^{a}-f_{i,\min}^{n},\qquad f_{i,\mathtt{den}}^{a}=f_{i,\mathtt{res}}^{a}-\alpha U_iU_i^{\top}f_{i,\mathtt{res}}^{a}. \]

The default \(\alpha=0.8\) suppresses the residual component aligned with the normal variation subspace; it neither uniformly shrinks the entire residual nor completely removes that subspace. Retaining 20% of the aligned component avoids entirely discarding genuine anomalies that partially share normal variation directions. The appendix's suppression-strength ablation also shows that complete removal is not necessarily best.

Queries undergo the same normal-neighbor retrieval and local suppression to obtain denoised normal-relative deviations. This symmetry matters: denoising references while leaving queries full of normal variations would introduce inconsistent noise into subsequent alignment. NVE does not perform target-domain gradient updates; its local PCA is computed from the current normal references.

2. Intrinsic Deviation Encoder: address regions with original anomalous features and convey discriminative content through denoised residuals

Denoised reference residuals still cover few seen anomalies, so directly applying nearest-neighbor matching can remain prone to overfitting. IDE uses 45 learnable vectors as cross-attention queries to extract 45 intrinsic deviation vectors from reference residuals. These vectors are conditioned on the input references rather than being immutable anomaly-category prototypes at deployment.

Attention keys come from original anomalous-reference features, while values come from NVE's denoised residuals. Their roles differ: original features preserve local semantics and help learnable queries address relevant regions, whereas values convey normal-relative differences without reintroducing object background into the deviation representation. The network therefore selects where to attend using original appearance and determines what to retrieve using residuals.

Not every patch of an anomalous reference is anomalous. Reference masks are downsampled to the feature grid, and attention logits strongly suppress normal-patch weights so aggregation focuses on annotated anomalous regions. Thus, the default method uses region masks as well as anomalous-image labels; inference does not require masks for target queries. Attention is followed by a feed-forward network, and the paper omits residual-path operations from its equation for brevity. That equation should not be mistaken for a complete implementation without residual connections.

Attention alone does not ensure that the 45 outputs represent different patterns. A discriminability constraint brings each annotated anomalous patch's denoised residual closer to its nearest intrinsic deviation vector. An orthogonality constraint penalizes squared cosine similarities between different vectors. The former anchors directions to actual reference anomalies; the latter reduces redundant directions so limited capacity covers more deviation patterns.

Orthogonality is encouraged by a loss, not guaranteed through exact orthogonalization. Directions also need not correspond to human-defined classes such as scratches or holes: they are discriminative directions learned through auxiliary training. Coverage of new defects must be tested experimentally; calling them intrinsic does not establish unconditional generalization.

3. Complementary Projection Scoring: assess anomalous alignment while retaining normal-matching evidence

Each denoised query residual is projected onto IDE's intrinsic deviation directions, and the directional components are summed. The paper defines this projection as:

\[ \tilde f_{i,\mathtt{den}}^{q}=\sum_{m=1}^{M} \frac{\langle f_{i,\mathtt{den}}^{q},t_m\rangle}{\langle t_m,t_m\rangle}t_m. \]

Rather than requiring a query to resemble an anomalous reference image, this asks whether its normal-relative deviation is preserved by learned anomalous directions. An unseen defect can respond strongly despite a different appearance if it departs from normality along some shared discriminative directions, whereas unrelated normal variations are suppressed. Since directions are only regularized toward orthogonality, this is the paper's sum of directional projections, not necessarily the exact orthogonal projection onto a subspace with an arbitrary nonorthogonal basis.

Scoring is not based solely on projection magnitude. It equally combines directional consistency between the denoised residual and its projected representation with the distance from the original query feature to its nearest normal feature:

\[ A_i^{q}=\frac{1}{2}\left(1-\mathbf d_{\mathrm{cos}}(f_{i,\mathtt{den}}^{q},\tilde f_{i,\mathtt{den}}^{q})+\mathbf d_{\mathrm{cos}}(f_i^{q},f_{i,\min}^{n})\right). \]

This retains the combination in the paper's Equation (9); \(\mathbf d_{\mathrm{cos}}\) is interpreted as cosine distance in the nearest-neighbor and scoring context. The cached text does not explicitly specify its scaling, clipping, or zero-vector handling, so no exact numerical range or missing implementation detail is inferred from the expression. Mechanistically, the first term supplies anomalous-direction alignment and the second supplies normal mismatch. Their complementarity avoids relying solely on a small but directionally aligned residual.

Patch scores are upsampled to the original image resolution for localization. The mean of the highest 1% of patch scores becomes the image-level score. This local high-score aggregation lets a small defect influence the image decision while depending less on one accidental peak than a single-patch maximum.

A Worked Example

Consider industrial surface inspection with 1 fixed normal reference and 1 anomalous reference with a scratch mask, while the query contains a different defect type absent from the reference. This is an explanatory example, not a quantitative measurement from an experimental figure.

The normal reference first supplies a patch feature bank. Each scratch-reference patch subtracts its nearest normal feature and suppresses normal variations using 4 PCA directions estimated from 12 normal neighbors. IDE excludes background through the scratch mask and generates 45 deviation directions from the remaining denoised residuals.

The query likewise subtracts normal neighbors and undergoes denoising. A genuine defect can activate Complementary Projection Scoring if it shares discriminative deviations with those directions, even without resembling the scratch in absolute feature space. Normal texture fluctuations primarily aligned with the local normal subspace are suppressed. The highest 1% of heatmap patches then jointly determine the image-level score.

This describes why cross-anomaly-type transfer can work, not why every new anomaly must be covered. A new defect whose deviation lies largely outside the learned directions, or unrepresentative normal references, can still cause failure.

Loss & Training

Training uses auxiliary-data episodes containing a query, normal references, and anomalous references from a single anomaly type. Appendix C.3 defaults to randomly selecting 600 normal/anomalous samples from the auxiliary dataset's test split as training queries. This is supervised auxiliary-source training, not training with target test-query labels. MVTecAD is the source for evaluating other datasets, while VisA is used as the source when evaluating MVTecAD.

Pixel-level Focal and Dice losses supervise query heatmaps, image-level BCE supervises the top-1% aggregated score, and reference-deviation discriminability and orthogonality losses are added:

\[ \mathcal L_{\mathrm{train}}=\mathcal L_{\mathrm{Focal}}+\mathcal L_{\mathrm{Dice}}+\mathcal L_{\mathrm{BCE}}+\mathcal L_{\mathrm{Dual}}. \]

The discriminability weight is \(\lambda_1=1.0\), and the orthogonality weight is \(\lambda_2=0.8\). The default freezes ViT-S/14 and updates only IDE parameters; inputs are 448ร—448, AdamW starts at learning rate 0.001, training lasts 20 epochs, and batch size is 16. The appendix additionally specifies 2 warm-up epochs, weight decay 0.0001, 384-dimensional cross-attention with 12 heads, and a feed-forward hidden dimension of 1536.

Normal shots are tested at 1, 2, 4, and 8; anomalous shots at 1 and 4, with additional analyses at 0 and 2. Although the formulation motivates more normal than anomalous references, experiments include N1A1. These are experimental grids with different budgets, not experiments that all strictly require more normal references than anomalous ones.

The generalization analysis assumes bounded episode loss, independent and identically distributed source episodes, finite capacity, and related conditions. It relates target risk to source empirical risk, capacity terms, sourceโ€“target distribution discrepancy, and joint optimal risk. Target terms simplify only with additional assumptions that target episodes follow a weighted source mixture and a low-joint-risk model exists. This is not an unconditional risk guarantee for arbitrary industrial or medical target domains.

Key Experimental Results

Main Results

Evaluation covers 5 industrial datasets and 3 medical imaging datasets, with results averaged over 3 independent runs. The table selects N1A1 results from the paper's Table 1. All values are image-level / pixel-level AUROC (%); the final column reports IDEAL's percentage-point gains over NAGL.

Dataset NAGL, N1A1 IDEAL, N1A1 Gain (image / pixel)
MVTecAD 95.1 / 96.1 96.3 / 96.8 +1.2 / +0.7
VisA 88.5 / 97.5 90.8 / 97.8 +2.3 / +0.3
AITEX 71.6 / 75.5 75.8 / 82.7 +4.2 / +7.2
MPDD 77.1 / 96.5 78.5 / 97.9 +1.4 / +1.4
BTAD 92.0 / 95.8 93.4 / 97.0 +1.4 / +1.2
BraTS 67.7 / 93.7 74.9 / 94.8 +7.2 / +1.1

General uses the full test set, including anomalies both matching and differing from the reference type. Hard excludes test samples of the reference anomaly type and evaluates only unseen anomaly types. Their test compositions differ, so a decline is not a paired comparison on the same queries.

In the paper's Table 2, N1A1 MVTecAD Hard results are 95.8 / 96.4 for IDEAL and 93.2 / 94.0 for NAGL; MPDD results are 74.5 / 97.2 and 63.7 / 91.8. Relative to General, MPDD image AUROC declines by 4.0 and 13.4 percentage points, respectively. This indicates less dependence on reference-anomaly appearance for IDEAL, although unseen types remain challenging.

The N1A1 advantage over normal-only N1 cannot be attributed entirely to architecture: the former also receives an anomalous image and region information. DRA/AHL additionally use full-shot normal data and target-specific training, while other methods differ in backbones and input resolutions. NAGL comparisons with matched reference budgets and fixed references, together with internal ablations, provide stronger evidence for specific mechanisms.

Ablation Study

The table selects the left side of the paper's Table 3, again reporting image-level / pixel-level AUROC (%). The first row uses anomalous references only and should not be described as having the same inputs as the subsequent dual-reference rows.

Config MVTecAD VisA BraTS
A1 only, direct matching 72.4 / 88.0 61.3 / 85.6 39.7 / 83.3
N1A1, direct matching 90.1 / 92.2 77.8 / 92.7 57.6 / 87.5
N1A1 + NVE, denoised-deviation matching 93.9 / 94.3 82.7 / 95.3 63.6 / 92.0
N1A1 + IDE, non-denoised residuals 94.5 / 96.0 87.3 / 96.5 70.1 / 92.9
N1A1 + NVE + IDE, full model 96.3 / 96.8 90.8 / 97.8 74.9 / 94.8

On VisA, the full model gains 3.5 / 1.3 percentage points over the IDE row without NVE, and 8.1 / 2.5 over the NVE-matching row. These support complementarity between denoising and direction learning, but differences between scoring mechanisms should not be assigned entirely to one isolated operator.

Key Findings

  • More normal references still help: under N4A1, IDEAL reaches 98.2 / 97.5 on MVTecAD and 87.6 / 98.5 on MPDD. N1A1 results should not represent every reference budget.
  • Mask quality is a practical condition. With pseudo-masks formed from the highest 1% of anomalous-to-normal feature distances in Table 6, IDEAL falls from 90.8 / 97.8 to 89.8 / 97.6 on VisA. This supports reduced annotation requirements, not the claim that the default needs no anomalous-region information.
  • Appendix E's deviation-vector count curve approaches saturation around 45, and excessive orthogonality loss separates correlated anomaly patterns too aggressively. Neither direction count nor diversity regularization should be maximized indiscriminately.
  • The VisA efficiency experiment reports 23.83M parameters, 0.5 hours of training, and 21.9 FPS for IDEAL. Different backbones and full-pipeline timing conventions prevent interpreting every efficiency ratio as a controlled same-model operator speedup.

Highlights & Insights

  • Normal residuals are not inherently clean anomaly representations. Estimating how normal features can vary through local normal neighbors and selectively suppressing those directions provides more explicit noise handling than merely enlarging the anomalous-reference bank.
  • The key/value split separates appearance-based addressing from deviation representation. Original anomalous features help identify region semantics, while residual values avoid treating the retrieved region's absolute appearance as the essence of abnormality.
  • Fixed single-type references and Hard evaluation directly test limited coverage. They better assess reliance on anomalous templates than averages reported only on a full test set containing seen anomalies.

Limitations & Future Work

  • The default requires target-domain anomalous references and region masks, and auxiliary training is supervised. It is neither purely unsupervised, zero-shot, nor normal-only; pseudo-mask experiments validate only one replacement strategy.
  • Local PCA depends on coverage from very few normal references, and real anomalies overlapping normal variation directions may also be suppressed. Reference contamination, illumination shifts, and varying defect sizes deserve further failure analysis.
  • Soft orthogonality does not necessarily yield a strictly orthogonal basis. Comparing the numerical behavior of summed directional projections with explicit orthogonalization or general subspace projection is a proposed analysis, not an experiment already completed by the authors.
  • The authors validate only image anomaly datasets, leaving transfer to tabular and time-series data untested. Medical imaging results also do not substitute for clinical deployment validation.
  • vs InCTRL / ResAD: These primarily use normal references and residuals for generalist detection. IDEAL adds anomalous references, denoises residuals, and converts them into discriminative deviation directions. Gains come with additional reference information and must be interpreted within that setting.
  • vs NAGL: Both use normal and anomalous references, but this paper emphasizes fixed references, single-anomaly-type coverage, and direction learning beyond anomalous-appearance matching. Its evaluation constructs fixed references for methods requiring anomalous examples; the NAGL values should not be treated as results under NAGL's original dynamic-reference protocol.
  • vs RegAD / PromptAD: These fit target-category models, whereas IDEAL trains on auxiliary data without target retraining. Generalist deployment reduces adaptation costs but does not eliminate reference collection and anomaly annotation costs.

Rating

  • Novelty: 4/5 โ€” A clear combination of local normal variation suppression and discriminative direction learning goes beyond direct dual-reference matching.
  • Experimental Thoroughness: 4/5 โ€” Eight datasets, Hard evaluation, and multiple ablations provide support, with reference sampling and supervision differences requiring care.
  • Writing Quality: 4/5 โ€” The problem and modules are connected clearly, while cosine notation and projection implementation details could be more explicit.
  • Value: 4/5 โ€” Useful for cross-domain visual detection with a few confirmed anomalies, rather than settings with no anomalous evidence at all.