Skip to content

ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/CLendering/ReFP-AD
Area: Object Detection
Keywords: Anomaly Detection, Energy-Based Models, Rectified Flow, Preconditioning, Langevin Dynamics

TL;DR

Addressing the core bottleneck where geometric anisotropy in high-dimensional foundation model token spaces destabilizes finite-step MCMC sampling in Energy-Based Models (EBMs), ReFP-AD introduces an optimal transport rectified flow preconditioning mechanism that maps uncompressed 1536-D features into an isotropic latent space, enabling stable preconditioned SGLD training and setting a new state-of-the-art for unified EBM anomaly detection.

Background & Motivation

Unsupervised visual anomaly detection aims to identify and localize rare, unseen defects during inference while training exclusively on defect-free normal samples. With growing demands across industrial quality inspection and autonomous systems, unified anomaly detection requires a single shared model to simultaneously generalize across diverse object categories, shapes, and surface textures. While modern self-supervised Vision Transformers (such as DINOv2) provide semantically rich and dense token embeddings, prevailing approaches predominantly rely on nearest-neighbor memory banks, patch similarities, or subspace reconstruction residuals. Although effective, these discriminative heuristics lack explicit, principled generative density modeling.

Energy-Based Models (EBMs) offer a compelling theoretical formulation for modeling complex multi-modal distributions by assigning a scalar energy value to each input and defining an unnormalized density, placing normal samples within low-energy basins and anomalies in high-energy regions without restrictive architectural assumptions. In practice, however, maximum-likelihood EBM training requires approximating negative-phase gradients through finite-step Markov Chain Monte Carlo (MCMC) sampling, typically via Stochastic Gradient Langevin Dynamics (SGLD) with persistent contrastive divergence (PCD). Standard Langevin dynamics inherently assumes diffusion under a Euclidean metric with isotropic noise. Unfortunately, foundation token representations (such as 1536-D embeddings from DINOv2-G) are severely anisotropic with extensive cross-dimensional correlations and extreme condition numbers. Under finite sampling budgets, this metric mismatch triggers severe oscillatory "zig-zag" trajectories and slow mixing, destabilizing the negative phase. Consequently, prior EBMs (such as MPDR) were forced to collapse visual representations into low-dimensional autoencoder bottlenecks (e.g., 272-D CNN features), sacrificing the fine-grained semantic granularity essential for precise anomaly localization.

This work identifies that the training instability of token-space EBMs is fundamentally a geometric mismatch rather than a deficiency in model capacity or energy parameterization. Since Euclidean Langevin dynamics requires well-conditioned, isotropic coordinates, the representation geometry must be proactively reshaped before energy modeling begins. The core idea is to employ an optimal transport (OT)-coupled rectified flow as a geometric preconditioner that transports high-dimensional, anisotropic visual tokens into a well-conditioned isotropic latent space, allowing an unconstrained EBM to be stably trained in full dimensions via preconditioned SGLD, while anomaly scoring directly leverages the restoring force gradient norm on the learned energy landscape.

Method

Overall Architecture

The ReFP-AD pipeline comprises five interconnected stages: unified token extraction and standardization, optimal transport rectified flow preconditioning, SGLD-oriented manifold validation with dynamic early stopping, unconstrained latent EBM training with preconditioned SGLD (pSGLD), and energy landscape gradient norm scoring.

First, input images are processed by a frozen DINOv2 ViT-G/14 backbone to extract multi-layer patch tokens, which are normalized using per-category Z-score standardization. Next, an entropic Sinkhorn optimal transport coupling guides a continuous rectified flow velocity field, transporting the anisotropic tokens into an isotropic Gaussian latent space. During flow training, spectral conditioning, off-diagonal cross-correlation, and heavy-tail ratios are evaluated within a random orthogonal subspace under topological rank preservation constraints to select the optimal MCMC-ready checkpoint. Subsequently, a residual MLP energy function is trained in this latent coordinate system using PCD and preconditioned SGLD with stratified replay buffer sampling. Finally, during inference, the norm of the energy function's restoring gradient with respect to latent coordinates yields pixel-level defect localization maps and image-level anomaly scores.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image & Multi-Layer Token Extraction<br/>DINOv2 ViT-G/14 Mean Pooling & Standardization"] --> B["Optimal Transport Rectified Flow Preconditioning<br/>Sinkhorn OT Coupling & Continuous Velocity Transport"]
    B --> C["SGLD-Oriented Manifold Validation & Early Stopping<br/>Random Subspace Conditioning & Rank Preservation"]
    C --> D["Preconditioned Stochastic Gradient Langevin Dynamics<br/>PCD Stratified Buffer & Adaptive Diagonal pSGLD"]
    D --> E["Energy Landscape Gradient Norm Anomaly Scoring<br/>Latent Restoring Force Norm & Top-1% Pixel Aggregation"]
    E --> F["Output Pixel Anomaly Segmentation Map & Image Score"]

Key Designs

1. Optimal Transport Rectified Flow Preconditioning: Reshaping Sampling Geometry via Continuous Transport

To address the severe anisotropy and ill-conditioning in raw foundation token representations that destabilize Euclidean Langevin diffusion, this module learns a bijective transport map that warps the empirical token manifold into an isotropic Gaussian reference prior \(p_{\text{prior}}(u) = \mathcal{N}(0, \tau^2 I)\), where the temperature \(\tau = 0.1\) governs latent scale and contraction-expansion balance. The continuous transport is defined by the ODE \(\dot{x}_t = v_\theta(t, x_t)\). To prevent trajectory crossing and minimize transport energy, standardized tokens \(z \in \mathbb{R}^D\) (\(D=1536\)) and Gaussian samples \(u\) are paired through an entropic optimal transport coupling \(\pi(z, u)\) computed via log-domain Sinkhorn iterations. Along linear interpolations \(x_t = (1-t)z + tu\), the velocity network is optimized with the rectified flow objective: $$ \mathcal{L}{\text{RF}} = \mathbb{E} \left[ |v_\theta(t, x_t) - (u - z)|_2^2 \right] $$ The velocity field weights are tracked via an Exponential Moving Average (EMA). During downstream stages and inference, exact token transport is computed deterministically by integrating the ODE from }[0, 1], (z, u) \sim \pi\(t=0\) to \(t=1\) using a 10-step 4th-order Runge-Kutta (RK4) numerical solver. This completely decouples representation geometric conditioning from density estimation without requiring any dimensionality compression.

2. SGLD-Oriented Manifold Validation & Early Stopping: Diagnostic-Guided Checkpoint Selection

Because low flow-matching loss does not necessarily imply well-behaved finite-step Langevin mixing, and unchecked flow training can distort the natural manifold topology, this design introduces a geometry-aware validation framework. Direct covariance inversion in 1536 dimensions is numerically ill-posed; therefore, transported tokens \(u\) are projected onto a fixed random orthogonal subspace of dimension \(k=128\), preserving second-order moments in expectation. Three MCMC diagnostic terms are computed on the projected covariance \(C\): the anisotropy condition ratio \(\mathcal{P}_\kappa = \log(\lambda_{\max}(C)/\lambda_{\min}(C))\), the mean squared off-diagonal correlation penalty \(\mathcal{P}_{\text{corr}} = \text{mean}(R_{\text{off}}^2)\), and the heavy-tail ratio \(\mathcal{P}_{\text{tail}} = q_{0.99}(\|u\|_2) / \text{median}(\|u\|_2)\). These terms form the composite SGLD-fitness criterion: $$ \mathcal{F}(u) = \mathcal{P}\kappa + \frac{1}{2} \mathcal{P} $$ To safeguard against degenerate manifold collapse, the framework tracks the Spearman rank correlation }} + \frac{1}{5} \mathcal{P}_{\text{tail}\(\rho_t\) between pairwise token distances in the original space \(z\) and latent space \(u\). Early stopping selects the checkpoint that minimizes \(\mathcal{F}(u_t)\) subject to the strict topological constraint \(\rho_t \ge 0.6\), ensuring optimal MCMC convergence properties while preserving data manifold fidelity.

3. Preconditioned Stochastic Gradient Langevin Dynamics: Adaptive Negative Sampling on Multi-Modal Landscapes

To address rapid curvature variations and mode collapse across heterogeneous categories in the unified setting, the EBM employs a persistent contrastive divergence (PCD) replay buffer combined with adaptive diagonal preconditioning. The buffer maintains 400,000 states with a 5% stratified re-initialization rate from the data manifold to maintain multi-class balance. The negative phase updates sample particles over \(K\) steps using preconditioned SGLD (pSGLD): $$ u_{k+1} = u_k - \frac{\eta}{2} M_k \nabla_u E_\phi(u_k) + \sigma \sqrt{\eta M_k} \, \xi_k, \quad \xi_k \sim \mathcal{N}(0, I) $$ Here, the diagonal preconditioner \(M_k = (\sqrt{v_k} + \epsilon)^{-1}\) tracks running squared gradients via an RMSProp-style update \(v_k = \beta v_{k-1} + (1-\beta)(\nabla_u E_\phi(u_k))^2\). The unconstrained energy function \(E_\phi(u)\) is a 3-layer residual MLP trained with energy magnitude regularization and positive-gradient smoothness penalties: $$ \mathcal{L}{\text{EBM}} = \mathbb{E}}[E_\phi(u)] - \mathbb{E{u^-}[E\phi(u)] + \alpha \mathbb{E}{u^\pm}[E\phi(u)^2] + \lambda \mathbb{E}{u^+}[|\nabla_u E\phi(u)|_2^2] $$ with hyperparameters set to \(\alpha = 0.1\) and \(\lambda = 10.0\). The preconditioner \(M_k\) rescales gradient updates along ill-conditioned directions, allowing the negative chains to mix across multi-modal energy basins in as few as 20–40 steps.

4. Energy Landscape Gradient Norm Anomaly Scoring: Deterministic Local Restoring Force Metric

Addressing the susceptibility of normalizing flows and reconstruction networks to assign high likelihoods or low reconstruction errors to out-of-distribution inputs, ReFP-AD treats the learned energy landscape as an explicit density discriminator. Normal tokens settle into smooth, low-energy valleys where local gradients are near zero, whereas anomalous patches reside on steep energy slopes where the landscape exerts a strong restoring force pulling the token toward the normal manifold. The patch-level anomaly score is formulated as the raw energy gradient norm: $$ S_{\text{patch}}(u) = |\nabla_u E_\phi(u)|_2 $$ At test time, the patch scores are reshaped onto the 2D spatial grid, bilinearly interpolated to input resolution, and smoothed with a Gaussian filter (\(\sigma=4.0\)). The image-level anomaly score is calculated as the mean of the top 1% pixel scores to eliminate localized noise spikes. Because the latent space has been preconditioned into a smooth geometry, this gradient norm serves as a highly discriminative, stable metric without requiring iterative test-time optimization or heuristic baselines.

Loss & Training

Training is executed in two decoupled stages: 1. Rectified Flow Preconditioning: The 8-layer MLP velocity field is trained on normal tokens from all categories using Adam (learning rate \(5 \times 10^{-5}\), batch size 8192, up to 150 epochs) with an EMA decay of 0.999. Training early-stops dynamically according to the SGLD-fitness metric \(\mathcal{F}(u)\) with a patience of 20 epochs. 2. EBM Optimization: Using the frozen flow checkpoint, tokens are mapped to \(u\)-space via 10-step RK4 integration. The 3-layer residual MLP energy function (hidden dimension 1024) is trained for 15 epochs using Adam (initial learning rate \(10^{-5}\), weight decay \(10^{-5}\), decayed by \(\gamma=0.4\) at epochs 8 and 12). Negative sampling uses \(K=60\) pSGLD steps with step size \(\eta=10^{-3}\), momentum \(\beta=0.99\), and noise standard deviation \(\sigma=0.1\).

Key Experimental Results

Main Results

Under a strict unified protocol, a single shared model is trained and evaluated across all categories of MVTec-AD (15 categories) and VisA (12 categories). ReFP-AD is benchmarked against feature-based baselines, normalizing flows, and state-of-the-art EBMs.

Method Category Model MVTec-AD I-AUROC (%) MVTec-AD P-AUROC (%) VisA I-AUROC (%) VisA P-AUROC (%)
Baseline Methods PaDiM 84.2 89.5 86.8 97.0
Baseline Methods MKD 81.9 84.9 74.2 93.9
Baseline Methods DRAEM 88.1 87.2 85.5 90.5
Normalizing Flows FastFlow 91.8 96.0 77.2 95.1
Normalizing Flows CFLOW 89.0 94.0 88.0 95.9
Normalizing Flows HGAD 98.4 97.9 97.1 98.9
Energy-Based Models EBM (Genc et al.) 72.0 70.7 - -
Energy-Based Models MPDR† (Stabilized) 96.0 96.7 86.5 96.5
Our Method ReFP-AD (Ours) 98.6 97.9 97.3 99.0

On the heterogeneous VisA benchmark, ReFP-AD outperforms the tuned EBM baseline MPDR† by +10.8% in Image AUROC (97.3% vs. 86.5%) and reaches 99.0% in Pixel AUROC. Furthermore, ReFP-AD surpasses the leading normalizing flow HGAD on both benchmarks, achieving 98.6% Image AUROC on MVTec-AD and 97.3% on VisA.

Ablation Study

Systematic ablations were conducted to isolate the impact of geometric preconditioning, sampling steps, backbone scale, and SGLD dynamics.

1. Geometric Preconditioning and Key Component Removal (VisA & MVTec-AD)

Experimental Configuration Mechanism Variation MVTec-AD I-AUROC MVTec-AD P-AUROC VisA I-AUROC VisA P-AUROC
Full Model (ReFP-AD) Full OT Flow Preconditioning + pSGLD 98.6 97.9 97.3 99.0
w/o Flow Preconditioning Raw 1536-D anisotropic tokens directly into EBM 91.1 (-7.5) 95.8 (-2.1) 87.0 (-10.3) 96.6 (-2.4)
w/o SGLD Preconditioning Standard SGLD without diagonal scaling - - 75.3 (-22.0) 94.3 (-4.7)
Per-Category Modeling Dedicated model per object class 99.2 (+0.6) 98.0 (+0.1) 97.7 (+0.4) 99.1 (+0.1)

2. Impact of SGLD Negative Steps \(K\) on MCMC Convergence (Unified MVTec-AD)

SGLD Steps (\(K\)) Image AUROC (%) Pixel AUROC (%) Dynamic Behavior
\(K = 10\) 54.7 67.8 Chains fail to mix; negative samples collapse, degrading training
\(K = 20\) 98.2 94.1 Preconditioned manifold enables rapid mixing; performance saturates early
\(K = 40\) 98.5 97.9 Pixel-level localization fully stabilizes
\(K = 80\) 98.6 97.9 Negligible marginal gain over \(K=40\)

3. Robustness Across Foundation Backbone Capacity

Feature Extractor Dimension (\(D\)) MVTec-AD I-AUROC (%) MVTec-AD P-AUROC (%) VisA I-AUROC (%) VisA P-AUROC (%)
DINOv2 ViT-B/14 768 97.7 97.6 95.7 98.8
DINOv2 ViT-L/14 1024 97.6 97.6 96.7 98.9
DINOv2 ViT-G/14 1536 98.6 97.9 97.3 99.0

Key Findings

  • Rectified flow preconditioning is strictly required for high-dimensional EBM stability: Omitting the OT flow preconditioning causes Image AUROC to plummet by 10.3% on VisA and 7.5% on MVTec-AD, confirming that severe anisotropy and cross-correlations in raw foundation token spaces prevent finite-step MCMC convergence.
  • Adaptive diagonal scaling (pSGLD) is essential for local multi-modal curvature: Removing the diagonal preconditioner in SGLD causes VisA Image AUROC to collapse from 97.3% to 75.3%, demonstrating that global transport alone is insufficient and local curvature adaptation is vital across multi-class energy landscapes.
  • Rapid MCMC mixing achieved in finite budgets: Owing to isotropic coordinate transformation, the EBM reaches near-optimal performance (98.2% Image AUROC) within only 20 SGLD steps, overcoming the traditional requirement for hundreds of Langevin iterations in high-dimensional spaces.
  • Minimal gap between unified and per-category models: The unified model trails dedicated per-category counterparts by merely 0.6% on MVTec-AD and 0.4% on VisA, proving that preconditioned latent spaces provide ample expressive capacity to represent multi-modal normal manifolds simultaneously.

Highlights & Insights

  • Diagnosing EBM instability as a geometric metric mismatch: Rather than constraining model capacity or compressing representations via autoencoder bottlenecks, ReFP-AD reframes the challenge as an issue of representation geometry, preserving full 1536-D semantic fidelity while restoring Euclidean Langevin validity.
  • MCMC-oriented manifold validation criteria: By formulating the SGLD-fitness score combining spectral conditioning, cross-correlations, and tail ratios alongside topological rank preservation, the framework avoids selecting overfitted flow checkpoints that fail during downstream MCMC sampling.
  • General paradigm for energy modeling in foundation feature spaces: The two-stage philosophy—condition representation geometry first, then perform unconstrained energy modeling—offers a reusable blueprint for training EBMs across generative modeling, representation learning, and out-of-distribution detection.

Limitations & Future Work

  • Inference latency overhead: Computing anomaly scores requires integrating a 10-step RK4 ODE solver followed by a backward pass through the energy MLP to evaluate gradient norms. While both networks are compact MLPs, the combined ODE and gradient evaluation incurs higher latency than feedforward classifiers or k-NN retrieval, rendering it more suitable for high-precision offline inspection.
  • Dependence on category-level standardization: Initial input preprocessing requires per-category feature statistics, which assumes test images possess known class labels, limiting pure zero-shot open-set deployments where category identity is unavailable.
  • Future research avenues: Distilling the preconditioned energy landscape into a single-step feedforward evaluation network to drastically reduce inference latency, and developing internal, adaptive normalizers to eliminate the need for category-conditioned Z-scoring.
  • vs MPDR (Yoon et al., NeurIPS 2023): MPDR addresses EBM instability by compressing features into a 272-D CNN autoencoder bottleneck and regularizing sampling via reconstruction fidelity, which struggles to converge in unified multi-class settings; ReFP-AD models uncompressed 1536-D tokens directly, gaining +10.8% Image AUROC on VisA.
  • vs HGAD (Yao et al., ECCV 2024): HGAD models densities via hierarchical Gaussian mixture normalizing flows; ReFP-AD uses flows strictly as geometric preconditioners and derives scores from EBM restoring gradient norms, avoiding high-likelihood false positives on out-of-distribution defects.
  • vs SimpleNet / PatchCore: While synthetic-noise classifiers and memory banks provide fast inference, they lack explicit density estimation foundations; ReFP-AD demonstrates that geometrically conditioned generative density models can achieve superior detection accuracy across unified industrial benchmarks.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates the instability of high-dimensional token EBMs as a geometric metric problem and introduces an elegant OT-coupled flow preconditioning paradigm.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous unified evaluation on MVTec-AD and VisA with extensive ablations on sampling steps, backbone scaling, and sampling dynamics, complemented by qualitative defect maps.
  • Writing Quality: ⭐⭐⭐⭐⭐ Strong theoretical motivation with a clean, logical narrative progressing from geometric diagnosis to empirical verification.
  • Value: ⭐⭐⭐⭐⭐ Provides a principled, scalable strategy for training unconstrained energy-based models on foundation token representations.