Skip to content

Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift

Conference: ECCV 2026
Paper: ECCV Official
Area: Object Detection
Keywords: Backdoor Detection, Object Detection, Pre-NMS Distribution Shift, Jensen-Shannon Divergence, Scene-Level Attacks

TL;DR

This paper introduces DistScan, which reveals for the first time that backdoor injection systematically distorts an object detector's pre-NMS prediction class distribution on clean inputs away from the training class prior, achieving trigger-free, weight-free, and re-training-free black-box backdoor detection via Jensen-Shannon divergence.

Background & Motivation

Object detection models are widely deployed in safety-critical applications such as autonomous driving, intelligent surveillance, and medical image analysis, where detection outputs directly dictate downstream decisions and system safety. However, by poisoning a small fraction of training samples to forge a stealthy association between a trigger pattern and malicious behaviors, backdoor attacks can preserve near-normal performance on benign images while triggering catastrophic failuresโ€”such as object disappearance, misclassification, or spurious object insertionโ€”whenever the trigger appears. This poses severe security threats to the distribution and deployment of third-party trained detectors.

Existing backdoor defenses for object detection are largely adapted from image classification or hinge on architecture-specific heuristics, suffering from crippling limitations against modern threat models. On the one hand, trigger inversion approaches such as ODSCAN rely on strong assumptions regarding trigger locality and explicit victim-to-target category transitions, rendering search spaces intractable when trigger geometry, style, and size are unrestricted. On the other hand, inconsistency-based defenses like MIA are strictly restricted to two-stage detectors and assume that the backdoor is isolated to a single sub-module. Crucially, neither approach can counter destructive "scene-level attacks" such as Detector Collapse (which induces scene-wide false alarm overload via SPONGE or complete blindness via BLINDING), where simultaneous corruption across branches entirely erases inter-module discrepancies.

Under the realistic requirements of requiring no trigger knowledge, imposing no architectural constraints, and defending against scene-level attacks, an effective detection signal must be intrinsic to standard inference and readily observable on clean inputs alone. The authors discover that regardless of one-stage or two-stage paradigms, an object detector's raw pre-NMS intermediate predictions inherently reflect the class-frequency prior of its training set; backdoor poisoning unavoidably warps this internalized prior, yielding a persistent class distribution shift even on benign images. The core idea is to aggregate intermediate pre-NMS class predictions on a small clean validation set, quantify their divergence from the ground-truth training class frequencies using Jensen-Shannon divergence, and flag backdoored models without trigger information or model weight access.

Method

Overall Architecture

DistScan operates through three interconnected stages: architecture-aware validation set construction tailored to one-stage and two-stage detectors, confidence-filtered pre-NMS prediction distribution extraction, and threshold-based decision via Jensen-Shannon (JS) divergence. The defender requires only black-box inference access to pre-NMS candidate outputs, a small set of clean images, and class-wise instance counts from the training set.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Model Under Inspection & Clean Samples"] --> B["Architecture-Aware Validation Set Construction<br/>Full images for one-stage / Cropped objects for two-stage"]
    B --> C["Confidence-Filtered Pre-NMS Distribution Extraction<br/>Filter low-confidence predictions & accumulate class counts"]
    C --> D["Distribution Normalization & Reference Grounding<br/>Normalize empirical distribution & training reference frequencies"]
    D --> E["Shift-Based Detection via JS Divergence<br/>Compute D(f) and classify against threshold"]

Key Designs

1. Architecture-Aware Validation Set Construction: Aligning Classification Head Exposure to Eliminate Confounders

Object detection benchmarks inherently suffer from severe class imbalance and extreme foreground-background ratios, which can inject spurious distribution shifts into validation predictions if sampling is unconstrained. Because one-stage and two-stage detectors expose their classification heads to fundamentally different inputs during training, a one-size-fits-all validation set compromises distribution fidelity. In one-stage models like YOLO, classification heads are optimized over full scenes via dense anchor predictions; DistScan thus uses randomly sampled complete clean images matching this global training distribution. Conversely, in two-stage models like Faster R-CNN, the second-stage classification head operates exclusively on RoI-pooled foreground proposals; DistScan therefore constructs the validation set using cropped single-object images (Equal Size) resized to median dimensions with balanced per-category counts. This tailored proxy suppresses background clutter and guarantees that observed distribution shifts stem from the model's corrupted class prior rather than image composition artifacts.

2. Confidence-Filtered Pre-NMS Distribution Extraction: Suppressing Low-Confidence Noise to Accumulate Class Priors

Before NMS filtering, a detector outputs a dense candidate set \(P_i = \{(\hat{c}_j, \hat{b}_j, \hat{s}_j)\}_{j=1}^{K_i}\) for each image \(x_i\). If all raw proposals are retained without filtering, two-stage detectorsโ€”which allocate an equal number of proposals across categories before suppressionโ€”collapse into a nearly uniform distribution that drowns out discriminative shifts; conversely, over-filtering discards genuine class responses. DistScan introduces a gentle confidence threshold \(\delta = 0.0005\), retaining only predictions with score \(\hat{s}_j \ge \delta\), and builds a class-wise count vector:

\[\mathbf{b}_i(f_\theta) = [b_i^1(f_\theta), b_i^2(f_\theta), \dots, b_i^C(f_\theta)]\]

Summing counts across validation images yields accumulator \(\mathbf{S}(f_\theta) = \sum_{i=1}^{|\mathcal{X}|} \mathbf{b}_i(f_\theta)\), which is normalized via the \(L_1\) norm to produce empirical pre-NMS prediction class distribution \(\hat{\mathbf{B}}(f_\theta) = \mathbf{S}(f_\theta) / \|\mathbf{S}(f_\theta)\|_1\).

3. Shift-Based Detection via JS Divergence: Symmetric and Bounded Distribution Shift Quantification

To establish a principled, model-agnostic baseline, DistScan normalizes the training set's annotated class instance frequencies \(\mathbf{R} = [r_1, r_2, \dots, r_C]\) into reference distribution \(\hat{\mathbf{R}} = \mathbf{R} / \|\mathbf{R}\|_1\). To prevent zero-probability tail classes from causing numerical instability or infinite penalties under standard Kullback-Leibler (KL) divergence, DistScan employs the symmetric and bounded Jensen-Shannon (JS) divergence:

\[D(f_\theta) = \mathrm{JS}(\hat{\mathbf{B}}(f_\theta) \parallel \hat{\mathbf{R}})\]

A model is flagged as backdoored if divergence \(D(f_\theta) > \tau\). Because benign models closely adhere to the training data manifold upon convergence while backdoor poisoning inevitably distorts the decision landscape across classes, the divergence metric yields a wide, robust separation margin between clean and backdoored models.

Key Experimental Results

Main Results

Evaluation was conducted on PASCAL VOC and MS-COCO across YOLOv5 and Faster R-CNN architectures under three scene-level attack scenarios: SPONGE (widespread false positive insertion), BLINDING (scene-wide object disappearance), and GMA (targeted misclassification). The evaluation pool comprised 288 independently trained detection models (18 backdoored and 18 benign models per configuration).

Scenario Dataset Architecture DistScan Acc.(%) DistScan TPR(%) DistScan FPR(%) MIA Acc.(%) ODSCAN Acc.(%)
SPONGE VOC YOLOv5 100.00 100.00 0.00 - 50.00
SPONGE VOC Faster R-CNN 97.22 94.44 0.00 50.00 50.00
SPONGE COCO YOLOv5 100.00 100.00 0.00 - 50.00
SPONGE COCO Faster R-CNN 91.67 83.33 0.00 50.00 50.00
BLINDING VOC YOLOv5 100.00 100.00 0.00 - 50.00
BLINDING VOC Faster R-CNN 100.00 100.00 0.00 88.89 50.00
BLINDING COCO YOLOv5 97.22 100.00 5.56 - 50.00
BLINDING COCO Faster R-CNN 94.44 88.89 0.00 77.78 50.00
GMA VOC YOLOv5 100.00 100.00 0.00 - 50.00
GMA VOC Faster R-CNN 91.67 83.33 0.00 86.11 50.00
GMA COCO YOLOv5 91.67 83.33 0.00 - 50.00
GMA COCO Faster R-CNN 100.00 100.00 0.00 58.33 50.00

Ablation Study

1. Effect of Validation Set Construction Strategy on Detection Accuracy (%)

Dataset Strategy YOLOv5 Mean Acc. Faster R-CNN Mean Acc.
PASCAL VOC Equal Number (Fixed per-class instance count) 100.00 67.59
PASCAL VOC Equal Size (Cropped single-object images) 88.89 96.30
PASCAL VOC Random (Unconstrained random full images) 100.00 67.59
MS-COCO Equal Number (Fixed per-class instance count) 93.52 77.78
MS-COCO Equal Size (Cropped single-object images) 95.37 95.37
MS-COCO Random (Unconstrained random full images) 96.30 68.52

2. Sensitivity Analysis of Confidence Filtering Threshold \(\delta\) (Mean Detection Accuracy %)

Threshold \(\delta\) VOC YOLOv5 VOC Faster R-CNN COCO YOLOv5 COCO Faster R-CNN
0.0000 (No filtering) 100.00 50.00 96.30 50.00
0.0005 (Recommended) 100.00 96.30 96.30 95.37
0.0010 96.30 96.30 96.30 67.59
0.0100 91.67 75.93 97.22 75.93
0.0500 90.74 71.30 95.37 78.70

Key Findings

  • Inversion baseline ODSCAN completely collapses under scene-level attacks, yielding a flat 50.00% accuracy (equivalent to random guessing). Its core premise of localized triggers inducing specific victim-to-target transitions is invalid against unlocalized, category-agnostic attacks.
  • Module inconsistency baseline MIA fails on one-stage detectors entirely and degrades to an AUROC of 0.00 on SPONGE with Faster R-CNN because scene-level attacks simultaneously compromise both RPN and classification heads.
  • DistScan achieves an outstanding overall average detection accuracy of 96.99% across all 288 models, outperforming applicable baselines by 27.32 percentage points on average while maintaining near-perfect AUROC (averaging 0.98+).
  • Alignment between validation data and training head exposure is decisive: YOLOv5 excels with full images, whereas Faster R-CNN demands cropped foreground objects. Setting \(\delta=0\) in Faster R-CNN causes detection accuracy to crash to 50.00% due to uninformative, uniformly distributed low-scoring proposals.

Highlights & Insights

  • Novel Defensive Perspective: Shifting focus away from trigger inversion optimization and internal module discrepancies to raw, pre-NMS class distribution statistics unlocks a universal, model-agnostic backdoor detection primitive.
  • Extreme Operational Efficiency: DistScan requires no gradient computation, no back-propagation, and no model parameter access; it executes purely as an inference-time forward-pass auditing tool over a small clean set.
  • Deep Empirical Grounding: PCA projections demonstrate that benign models strictly cluster around the training data class prior, whereas backdoor injection induces persistent, multifaceted distribution shifts on clean inputs that reflect specific attack semantics (e.g., target-class inflation vs. multi-class suppression).

Limitations & Future Work

  • Reliance on Reference Training Class Frequencies: DistScan assumes the defender possesses ground-truth class frequency statistics from the training set; in fully blind proprietary scenarios, reference distributions must be approximated via external proxy datasets.
  • Slight Margin Compression on Granular Datasets: On dense datasets like MS-COCO with 80 fine-grained categories, minor semantic overlap across classes slightly narrows the separation margin; integrating intermediate feature-level distributions could provide stronger multi-layer verification.
  • Adaptive Evasion Attacks: Potential future adversaries could attempt to regularize pre-NMS class distributions on clean training samples to mimic benign reference frequencies, introducing an interesting cat-and-mouse defense dynamic.
  • vs ODSCAN (IEEE S&P 2024): ODSCAN optimizes trigger patterns under solid-color patch and localized target assumptions, stalling when triggers are non-localized or structured; DistScan sidesteps trigger synthesis entirely by capturing global statistical footprints on clean images.
  • vs MIA (ICPR 2024): MIA exploits discrepancies between RPN proposals and classification scores, restricting it to two-stage models and failing when attacks simultaneously corrupt both stages; DistScan operates on pre-NMS outputs common to both paradigms and reliably defends against dual-branch scene-level attacks.

Rating

  • Novelty: โญโญโญโญโญ Pioneering observation and systematic utilization of pre-NMS prediction class distribution shift under clean inputs for backdoor defense.
  • Experimental Thoroughness: โญโญโญโญโญ Rigorous validation across 288 models spanning two detector paradigms, two benchmark datasets, and three destructive scene-level attacks.
  • Writing Quality: โญโญโญโญโญ Flawless narrative structure, clear mathematical formulation, insightful failure analysis of baselines, and compelling visualizations.
  • Value: โญโญโญโญโญ Highly practical, lightweight, and deployable black-box auditing tool for safeguarding object detection supply chains.