Skip to content

Towards Generalizable 3D Anomaly Detection via Relational Inconsistency Modeling

Conference: NeurIPS2026
arXiv: 2609.35059
Code: https://github.com/VisualScienceLab-KHU/GRIM
Area: 3D Vision / Anomaly Detection
Keywords: point cloud anomaly detection, relational inconsistency, geometric graphs, within-sample clustering, pseudo-anomaly supervision

TL;DR

GRIM learns a defect criterion from controlled local relational violations in normal point clouds through edge-aware graph refinement and within-sample cluster-deviation modeling, achieving 97.4/94.5 object-/point-level AUROC on Anomaly-ShapeNet and 83.6/89.6 when transferred directly to Real3D-AD.

Background & Motivation

Industrial point cloud defects need not be local regions that are intrinsically unfamiliar: a surface patch can look plausible in isolation while being incompatible with the orientation, curvature, or connectivity of surrounding surfaces. Existing methods generally learn normality from defect-free samples. Reconstruction methods score reconstruction errors, whereas feature-matching methods score distances to normal references; the former may also reconstruct real defects, while the latter may flag rare but valid structures. Neither directly teaches the model to distinguish normal variation from changes that disrupt structural relationships.

This problem becomes particularly pronounced under unified multi-category training. One model per category narrows the normal distribution but increases deployment and maintenance costs; a shared model must accommodate more diverse, overlapping normal structures. Across domains, categories, scanning conditions, and visible geometry change further, making memorization of source-domain normal shapes less transferable. The paper therefore targets geometric inconsistency that depends less on object identity, rather than enumerating all possible defect appearances.

Pseudo-anomalies are not a replacement for a database of real defects here. They are controlled interventions: modify a small part of a normal object, preserve its surroundings, and examine compatibility with neighboring and structurally similar regions. Core idea: replace โ€œdistance from normalityโ€ with โ€œviolation of geometric relationships,โ€ using pseudo-anomaly supervision to jointly learn neighborhood-aware representations and relative deviation within structural peer groups.

Method

Overall Architecture

GRIM takes only a point cloud as input and produces a point-level anomaly map and an object-level anomaly score. All real training data are normal samples; real defect labels are not used. A pseudo-anomaly generator additionally supplies binary masks of modified points. Unified training pools all categories without category labels or category-specific parameters.

FPS first selects centers and collects neighborhoods, and a frozen PointMAE extracts local features. Edge-aware Graph Refinement (EGR) then incorporates relationships to neighboring structures. Cluster-Deviation Modeling (CDM) finds structurally similar peers within the current sample and concatenates a relative-deviation embedding with refined features before a shared classifier. Inference reuses EGR, CDM, and the classifier, but neither generates pseudo-anomalies nor uses supervision masks.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Training: real normal cloud"] --> B["Pseudo-anomaly<br/>Relational Supervision"]
    B -->|normal and pseudo-anomalous samples| C["FPS grouping<br/>Frozen PointMAE"]
    T["Inference: test cloud"] --> C
    C --> D["Edge-aware Graph<br/>Refinement EGR"]
    D --> E["Cluster-Deviation<br/>Modeling CDM"]
    E --> F["Shared classifier<br/>Point map and object score"]
    B -.->|training only: masks and cluster-balanced BCE| F
    B -.->|training only: anomaly labels and deviation loss| E

Key Designs

1. Pseudo-anomaly Relational Supervision: learning structural violations through local interventions

The generator randomly selects a seed point and forms a local patch from its neighbors, typically modifying approximately 1% of the object's points. PCA estimates the patch normal, which is oriented outward relative to the global point cloud center. Bulge displaces points outward along the normal, while sink moves them inward in the opposite direction. Displacement is strongest near the patch center and decreases smoothly toward the boundary, with a scale sampled uniformly from 0.01โ€“0.03. The altered patch retains object context while disrupting local surface continuity and geometric relationships with its neighborhood.

Hole does not simply delete points. It moves central points toward the boundary and slightly inward, while mildly expanding the boundary ring to create a cavity-like deformation. Modified points receive anomaly labels; unchanged points remain normal. A subsequent random rotation transforms the entire point cloud without changing its mask. The three primitives provide outward, inward, and local surface-support-depletion supervision, not a complete defect taxonomy. The objective is to learn transferable relational violations rather than exhaust the appearance of test defects.

2. Edge-aware Graph Refinement EGR: evaluating local shapes in their neighborhood context

FPS selects 2048 representative centers, each with 128 neighboring points. Each group is centered before encoding with a frozen PointMAE pretrained on ModelNet40. Groups become graph nodes, and their centroids define a directed 32-NN graph. Local encoding describes the shape of a patch, while edges restore relative spatial and surface relationships that local centering alone does not adequately express.

Each group supplies a centroid, coordinate-wise standard deviation, surface normal, and curvature. The normal is the eigenvector corresponding to the smallest covariance eigenvalue, and curvature is that eigenvalue divided by the sum of all eigenvalues. For a directed edge from group \(j\) to group \(i\), the initial eight-dimensional relational feature is:

\[ \mathbf{e}_{ji}=\left[(\boldsymbol{\mu}_i-\boldsymbol{\mu}_j)\,\|\,(\boldsymbol{\sigma}_i-\boldsymbol{\sigma}_j)\,\|\,(1-|\mathbf{n}_i^\top\mathbf{n}_j|)\,\|\,|\kappa_i-\kappa_j|\right]. \]

These terms encode relative displacement, differences in point spread, normal inconsistency, and curvature differences. Taking the absolute normal inner product avoids mistaking a normal's sign ambiguity for a geometric anomaly. The graph therefore explicitly incorporates geometry rather than connecting regions solely by feature similarity.

At each layer, both endpoint features and the current edge feature are concatenated to refine the edge representation. Attention is computed from the destination-node query and the refined edge, normalized across incoming neighbors. Projected source-node features are attention-weighted and summed into a neighborhood message; both node and edge representations receive residual updates. The implementation uses two refinement layers, four attention heads, and an attention dimension of 384. The same local surface can consequently acquire different features under different neighborhood relationships, making the criterion more contextual than isolated patch appearance.

3. Cluster-Deviation Modeling CDM: comparison with structural peers in the current object, not a normal reference bank

Beyond neighborhood relationships, the model must distinguish incompatibility with peers from membership in a different valid structural type. CDM \(\ell_2\)-normalizes the refined node features of one sample, initializes clustering with k-means++, and runs eight iterations to form 64 clusters. Each center is the normalized mean of its member features, and each node is compared with its assigned center. Every test sample undergoes the same clustering at inference; no cross-sample normal memory bank is stored, and centers are not retrieved from training data.

The cosine distance between a node and its assigned center is:

\[ d_g=1-\hat{\mathbf{x}}_g^\top\bar{\mathbf{z}}_g. \]

Distances are subsequently minโ€“max normalized within each cluster using the 0.1โ€“0.9 quantiles, encoded into a 128-dimensional sinusoidal embedding with temperature 10000 and scale 10, and linearly projected. The deviation embedding is concatenated with EGR-refined features and fed to a shared MLP with hidden dimensions [256, 128] and dropout 0.1. The classifier predicts the final anomaly score, which is not simply the cosine distance. It therefore receives both structural content and relative peer-deviation evidence. The paper does not sufficiently specify clipping or degenerate-cluster handling for quantile normalization here; those implementation details are not reconstructed.

After node scoring, each original point inherits the score of its nearest sampled center, followed by Gaussian smoothing. The maximum of the smoothed point map becomes the object-level score. No per-sample minโ€“max normalization is applied to node or point scores before taking that maximum, avoiding forced high-score regions in every normal object. For evaluation only, object scores are minโ€“max normalized across evaluation samples within each category after aggregation. This evaluation procedure is distinct from training the model without category labels.

A Worked Example

Consider a test object with a locally depressed surface. This is a mechanism illustration, not an additional experiment. The cloud becomes 2048 groups. Groups near the depression may still have familiar local surface features, but their normals, curvature, and positional relationships to surrounding groups change. EGR incorporates this incompatibility into their representations.

CDM then forms 64 structural peer clusters within the object. If depression-related nodes remain assigned to clusters dominated by normal surfaces, their deviation from the center provides additional evidence. The classifier combines refined features and deviation embeddings, nearest-center mapping and smoothing recover the point map, and the maximum point score indicates object abnormality. If a large, internally coherent anomaly instead forms its own cluster, deviation may become small; this is an explicitly acknowledged limitation.

Loss & Training

A group label is the maximum of its member-point labels: one pseudo-anomalous point suffices to label the group anomalous. Classification loss averages BCE within each cluster and then across clusters, preventing abundant normal structural clusters from dominating gradients. A separate deviation loss applies only to pseudo-anomalous nodes. It penalizes cosine similarity to the assigned center above a margin \(\delta\), discouraging anomalous nodes from remaining close to their peers.

\[ \mathcal{L}_{\mathrm{dev}}=\frac{1}{|\mathcal{G}_1|}\sum_{g\in\mathcal{G}_1}\max(\hat{\mathbf{x}}_g^\top\bar{\mathbf{z}}_g-\delta,0),\qquad \mathcal{L}=\mathcal{L}_{\mathrm{BCE}}+1.5\mathcal{L}_{\mathrm{dev}}. \]

\(\mathcal{G}_1\) denotes the anomalous-node set. The cache defines the margin but does not provide its numerical value, which is not guessed. Deviation loss shapes the representation, whereas the embedding supplies relative-position information to the classifier. These are distinct roles; CDM is not merely clustering followed by distance-based detection.

Each normal training sample generates four pseudo-anomalous variants, with bulge/sink/hole probabilities of 0.4/0.4/0.2. PointMAE remains frozen. The remaining model is trained with AdamW for 100 epochs at batch size 1, with learning rate and weight decay both \(10^{-4}\) and StepLR decay. Training takes approximately 12 hours on one RTX 3090. Cross-domain evaluation directly applies the source model without fine-tuning or target-domain adaptation.

Key Experimental Results

Main Results

Anomaly-ShapeNet contains 40 categories and 1600 samples, with four normal training samples per category. Real3D-AD has 12 categories and likewise four normal training samples per category; training captures cover 360ยฐ, whereas test scans are single-view. The following selection uses Tables 1โ€“2. A denotes Anomaly-ShapeNet and R denotes Real3D-AD; values are category-mean AUROC (%), higher is better.

Training โ†’ test Method / setting O-AUROC P-AUROC
A โ†’ A MC3D-AD, unified model 84.2 75.9
A โ†’ A GRIM, unified model 97.4 94.5
R โ†’ R MC3D-AD, unified model 78.2 76.8
R โ†’ R GRIM, unified model 87.2 91.3
A โ†’ R MC3D-AD, no target-domain adaptation 55.4 38.3
A โ†’ R GRIM, no target-domain adaptation 83.6 89.6
R โ†’ A MC3D-AD, no target-domain adaptation 78.3 48.8
R โ†’ A GRIM, no target-domain adaptation 91.7 78.9

Against the previous best method for each metric, GRIM gains 7.4 percentage points over PASDF in object-level AUROC and 4.7 over PO3AD in point-level AUROC on A. On R, the corresponding gains are 7.0 over PASDF and 3.5 over Reg2Inv. Most of those baselines train separately per category, a protocol difference that should not be overlooked. Appendix Tables 11โ€“12 additionally retrain PO3AD and PASDF under unified training and compare cross-domain results, supporting the same overall conclusion.

Source numerical conflict: ยง4.2 reports gains over MC3D-AD on R of 8.9/14.4, but Table 1 lists 78.2/76.8 versus 87.2/91.3, implying 9.0/14.5 percentage points. The original table values and the conflicting prose claim are both retained rather than silently reconciled. The reported A gains of 13.2/18.6 are consistent with the table.

For A โ†’ R, GRIM drops only 3.6/1.7 relative to its own model trained on R. For R โ†’ A, the drops are 5.7/15.6, revealing substantial directional asymmetry in cross-domain localization. These are AUROC values, not fixed-threshold FPR/FNR. The introductory Figure 2 separately evaluates errors at fixed TPR/TNR; this table does not replace that threshold-based evidence.

Ablation Study

The following selection combines Table 3 and Appendix Table 10. Table 3 compares module combinations; Table 10 preserves CDM clustering while separating deviation loss and deviation embedding. All four metric columns report AUROC (%).

Config A โ†’ A object A โ†’ A point A โ†’ R object A โ†’ R point
Backbone + classifier, no EGR/CDM 89.5 58.1 56.9 42.8
EGR only 94.3 86.4 59.4 53.0
CDM only 96.6 79.8 62.3 64.9
EGR + full CDM 97.4 94.5 83.6 89.6
Full model without deviation loss 96.6 92.1 79.3 84.1
Full model without deviation embedding 94.9 88.7 68.5 66.8
Clustering retained, both loss and embedding removed 94.3 86.4 59.4 53.0

EGR alone increases A โ†’ A point-level AUROC by 28.3. CDM alone increases A โ†’ R point-level AUROC by 22.1 over the backbone. Their combination substantially outperforms either module in cross-domain evaluation. Removing deviation embedding from the full model reduces cross-domain point-level AUROC by 22.8, versus 5.5 when removing deviation loss, highlighting the importance of explicitly exposing relative deviation to the classifier.

Table 4 also shows that increasing the number of pseudo-anomaly types does not monotonically improve every metric. The following subset illustrates this point without treating the three primitives as an exhaustive defect inventory.

Training pseudo-anomalies A โ†’ A object A โ†’ A point A โ†’ R object A โ†’ R point
sink 98.5 95.9 81.2 81.8
bulge + hole 98.6 96.0 80.3 87.9
sink + hole 99.2 98.1 80.3 86.7
bulge + sink + hole 97.4 94.5 83.6 89.6

Key Findings

  • All three primitives perform best in the evaluated A โ†’ R setting, but sink + hole reaches 99.2/98.1 on A โ†’ A, above the full configuration. Complementary supervision and the best in-domain result are not equivalent.
  • Appendix Table 9 reports 87.1/89.2 for hole-only training on R, whose real test defects are bulge/sink rather than hole. Its proximity to the three-primitive result of 87.2/91.3 offers limited but direct evidence for transfer across defect types.
  • Figure 5 reports mean center distances of 0.73 for anomalous features and 0.48 for all features. โ€œAllโ€ includes anomalies and is not a normal-only control. Broken/bending/scratch in Appendix Table 8 are not explicitly synthesized, whereas concavity is geometrically close to sink and should not be treated as entirely unseen.
  • Appendix Table 5 increases cluster count from 16 to 32 to 64, yielding A โ†’ R point-level AUROC of 73.3, 79.0, and 89.6. This supports only the tested range, not an unrestricted more-clusters-is-better claim. Table 6 gives cross-domain point-level values of 88.4/86.7/89.6 for deviation weights 0.5/1.0/1.5, also a non-monotonic sequence.
  • Appendix Table 7 measures average inference time over A's 1312 test samples on an RTX 3090: 0.35 s versus MC3D-AD's 0.84 s, with 34.0M versus 34.2M parameters. These are hardware- and dataset-specific measurements, not a general real-time guarantee.

Highlights & Insights

  • Supervision targets relational violations rather than appearance templates. Synthetic and real defects need not match exactly for the model to learn geometric compatibility cues reusable across categories.
  • EGR and CDM provide distinct reference frames. EGR evaluates spatial-neighborhood compatibility, while CDM compares structural peers within the current object; joint ablations reveal their complementary value more clearly than either module alone.
  • Deviation is both a training constraint and an input feature. Ablations with unchanged clustering avoid attributing gains to a new clustering algorithm and show separate contributions from representation shaping and deviation injection.

Limitations & Future Work

  • The authors explicitly note that CDM assumes anomalies are relatively sparse. Large coherent anomalies may form their own cluster and have small center distances. Providing EGR features to the classifier does not itself establish robustness; reducing cluster count remains an unverified direction.
  • Geometry alone misses cues such as color changes. The authors propose RGB-D integration; an extension should verify whether multimodal relationships remain transferable instead of merely adding an appearance branch.
  • Two datasets and three local primitives do not cover the full range of industrial defects. Reliability of normals, outward orientation, and peer clusters on thin structures, complex topology, and large deformations requires dedicated testing.
  • AUROC supports improved score ranking, not necessarily low false alarms at deployed thresholds. Further evaluation should include shared thresholds, cross-category calibration, false positives at fixed recall, and multiple random seeds with uncertainty intervals.
  • vs MC3D-AD: Both use a unified model, but MC3D-AD learns normality through geometric reconstruction. GRIM directly supervises relational violations using pseudo-anomalies, with a larger localization advantage under domain shift than in-domain.
  • vs PO3AD / R3D-AD: These approaches use pseudo-anomalies to learn point offsets or reconstruction, respectively. GRIM instead uses synthetic changes to train a relational criterion rather than maximize synthetic appearance coverage. Sharing pseudo-anomalies does not imply sharing the learning objective.
  • vs normal feature-bank methods: CDM centers come from the current point cloud, not a training normal-feature bank. A transferable direction is to construct within-sample peers in other structured anomaly tasks while explicitly testing failures when anomalies form coherent groups of their own.

Rating

  • Novelty: 4/5. The relational-violation perspective and joint EGR/CDM design are distinctive, although graph message passing and k-means are not new algorithms.
  • Experimental Thoroughness: 4/5. Bidirectional transfer, component separation, defect-type tests, and unified retraining provide substantial evidence; large-anomaly stress tests and repeated-run statistics remain absent.
  • Writing Quality: 4/5. Mechanisms and limitations are clear, but Real3D-AD gain claims conflict arithmetically with the table, and some implementation parameters are insufficiently specified.
  • Value: 4/5. The framework offers a reusable direction for unified point cloud detection with few normal samples; industrial generalization still requires broader defect and scanning-condition validation.