DETR is Secretly a Multispectral Detector: Zero-Parameter Adaptation via Semantic Alignment¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/UserXiangYang/ZPA-MDETR
Area: Object Detection
Keywords: Multispectral Object Detection, DETR, Zero-Parameter Adaptation, Semantic Alignment, Asymmetric Cross-Modal Interaction
TL;DR¶
Challenging the parameter-expanding convention in multispectral detection, this paper introduces ZPA-MDETR, which preserves the unimodal DETR architecture and weights without adding learnable fusion parameters, achieving state-of-the-art performance via asymmetric query-memory organization and parameter-free spatial-channel semantic alignment.
Background & Motivation¶
Multispectral object detection (MOD) aims to boost perceptual robustness under adverse conditions such as low illumination, night, and heavy fog by integrating complementary signals from visible (RGB) and infrared (IR) sensors. While RGB images capture rich texture, fine geometry, and distinct color information, their quality degrades sharply in low-light environments. Thermal infrared sensors detect object-emitted thermal radiation and operate independently of external lighting, but they typically suffer from low resolution, noise, and lack of fine appearance textures. To effectively fuse these complementary cues, existing CNN-based and DETR-based detectors predominantly follow a parameter-expanding paradigm, designing specialized dual-stream interaction modules, cross-scale fusion encoders, or attention-based fusion decoders.
However, this parameter-expanding paradigm introduces two fundamental practical bottlenecks. First, newly introduced fusion modules alter the unimodal network topology, preventing them from benefiting from rich unimodal pre-trained representations and requiring random initialization. Second, adding customized fusion layers inevitably bloats model parameter scale and memory footprint, creating a severe deployment obstacle for resource-constrained edge computing platforms in autonomous driving and aerial robotics.
This paper challenges the conventional belief that multimodal fusion inherently requires structural expansion. By re-examining the DETR architecture, the authors observe that its native decoding mechanism is fundamentally grounded in a query-memory interaction paradigm, where object queries retrieve and aggregate target representations from the feature memory across a shared semantic space. The core idea is to strictly preserve the network topology and parameter scale of unimodal DETR, realizing multimodal collaboration purely through asymmetric input organizationโassigning one modality as memory and the other to initialize object queriesโpaired with parameter-free spatial-channel semantic alignment regularizers during training (ZPA-MDETR).
Method¶
Overall Architecture¶
ZPA-MDETR completely retains the architecture, parameter count, and operator graph of single-modal DETR (built upon real-time RT-DETR), without introducing any extra learnable parameters. The overall pipeline concatenates RGB and channel-replicated 3-channel IR inputs along the batch dimension, passing them through a shared encoder to extract multi-scale feature representations. During the decoding stage, an asymmetric input organization strategy designates one modality as the feature memory while using the other modality to generate candidate object queries and reference points. Finally, cross-modal interaction is carried out via standard deformable cross-attention layers. During training, parameter-free spatial and channel semantic alignment regularizers constrain the shared encoder to project both modalities into a unified semantic space; these constraints are discarded during inference at zero deployment cost.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Paired RGB-IR Inputs<br/>IR replicated to 3 channels"] --> AIO["Asymmetric Input Organization<br/>Batch concatenation into shared encoder"]
AIO --> MemorySplit["Multi-Scale Feature Splitting<br/>Obtain RGB memory Mrgb and IR memory Mir"]
MemorySplit --> SA["Parameter-Free Semantic Alignment<br/>Spatial cosine alignment + channel decorrelation"]
MemorySplit --> QuerySelect["Asymmetric Query Selection<br/>Sample Top-K queries and reference points from query modality"]
QuerySelect --> DCA["Asymmetric Deformable Cross-Attention<br/>Sample and aggregate memory context using query offsets"]
SA -.->|Regularize unified space| DCA
DCA --> DecOut["6-Layer Cascaded Decoder Outputs<br/>Bounding box regression and classification"]
Key Designs¶
1. Asymmetric Input Organization: Preserving Single-Modal Topology and Decoupled Interaction Addressing the topological inconsistency and initialization hurdles caused by customized dual-branch designs, this paper proposes an asymmetric input organization strategy. Given an RGB image \(I_{rgb} \in \mathbb{R}^{3 \times H \times W}\) and a single-channel IR image \(I_{ir} \in \mathbb{R}^{1 \times H \times W}\), the IR image is first replicated along the channel dimension into \(I'_{ir} \in \mathbb{R}^{3 \times H \times W}\). Both are then concatenated along the batch dimension to form a unified tensor \(I_{unified} \in \mathbb{R}^{2 \times 3 \times H \times W}\). A single unimodal pre-trained encoder processes this tensor directly, avoiding any architectural alterations. The extracted features are split along the batch dimension and flattened into modality-specific sequence memories \(M_{rgb}, M_{ir} \in \mathbb{R}^{T \times C}\) (where \(C=256\)). The system explicitly breaks modal symmetry: one modality is routed through the native query selection module to yield initial object queries \(Q \in \mathbb{R}^{K \times C}\) and reference boxes \(P \in \mathbb{R}^{K \times 4}\) (\(K=500\), comprising 300 score-selected queries and 200 denoising queries), while the other modality serves as the global contextual memory. This converts cross-modal fusion into the native query-memory key-value retrieval of the DETR decoder.
2. Asymmetric Deformable Cross-Attention: Sparse Cross-Modal Target Probing To aggregate cross-modal information efficiently, the framework directly inherits standard deformable cross-attention. Taking RGB as the query modality and IR as the memory modality, at each layer of the 6-layer cascaded decoder, the self-attention-enhanced query \(\tilde{Q}_{rgb}\) dynamically predicts sampling point offsets \(\Delta p_{rgb}^{(hns)}\) and attention weights \(A_{rgb}^{(hns)}\) targeting the IR memory via linear projections. The offsets are scaled by the width and height of reference box \(P_{rgb}\) to maintain scale adaptivity, after which features are sampled from multi-scale IR feature maps \(M_{ir}^{(n)}\) and aggregated: $$ Q_{fus} = \sum_{h=1}^{N_h} W_h \left( \sum_{n=1}^{N_l} \sum_{s=1}^{N_s} A_{rgb}^{(hns)} \cdot W'h M}^{(n)} \left( \Phi_n(\hat{p{rgb}) + \Delta p \right) \right) $$ where }^{(hns)\(N_h=8, N_l=5, N_s=3\), and \(\Phi_n(\hat{p}_{rgb})\) maps normalized coordinates onto the \(n\)-th feature level. In this manner, queries derived from one modality act as spatial probes to retrieve complementary target cues from the other modality's global context, combining computational sparsity with rich multimodal synergy.
3. Parameter-Free Semantic Alignment: Bridging Spatial Discrepancy and Removing Channel Redundancy While DETR decoders are structurally modality-agnostic, severe distribution divergence between RGB and IR representations can induce cross-modal semantic confusion. To resolve this without adding parameters, two training-only regularizers are imposed on the encoder outputs. Spatial semantic alignment enforces consistency across modalities at each identical spatial position \(t\) by maximizing cosine similarity: $$ \mathcal{L}{spatial} = \frac{1}{T} \sum|}^{T} \left( 1 - \frac{M_{rgb}^{(t)} \cdot M_{ir}^{(t)}}{|M_{rgb}^{(t)2 |M|}^{(t)2} \right) $$ To prevent representation collapse and suppress redundant information, channel semantic alignment normalizes memories across the channel dimension and optimizes the cross-correlation matrix \(C \in \mathbb{R}^{C \times C}\): $$ \mathcal{L}^2 $$ with } = \sum_{i} (C_{ii} - 1)^2 + \lambda_{off} \sum_{i} \sum_{j \neq i} C_{ij\(\lambda_{off}=0.005\). The diagonal objective aligns identical semantic channels across modalities, whereas the off-diagonal penalty pushes distinct channels toward decorrelation and orthogonality, ensuring compact, diverse, and well-aligned latent representations.
4. Reliability-Guided Query-Memory Role Selection Analyzing the asymmetric roles across diverse scenarios reveals that the optimal assignment is strongly correlated with relative unimodal detection reliability. Because memory serves as the key-value foundation for cross-attention, the modality offering cleaner, more reliable evidence should be designated as memory, while the complementary modality serves as the query probe. On low-light or severely imbalanced benchmarks (FLIR, LLVIP, RGBTDronePerson), IR features are substantially more robust and are best assigned as memory. Conversely, on well-illuminated and balanced benchmarks (M3FD), RGB features retain richer semantic details and are better assigned as memory.
Loss & Training¶
The total training loss combines the native detection loss with the parameter-free regularizers: $$ \mathcal{L}{total} = \mathcal{L}} + \mathcal{L{spatial} + \mathcal{L} $$ where \(\mathcal{L}_{det}\) denotes the standard RT-DETR detection loss (incorporating Focal Loss, L1 bounding box loss, and GIoU loss). All models are trained end-to-end on a single NVIDIA GeForce RTX 3090 GPU using a COCO-pretrained ResNet50 backbone. Input resolutions are \(640 \times 640\) for FLIR, M3FD, and RGBTDronePerson, and \(1024 \times 1024\) for LLVIP.
Key Experimental Results¶
Main Results¶
On four widely used multispectral detection benchmarks (FLIR, M3FD, LLVIP, RGBTDronePerson), ZPA-MDETR maintains identical parameter count to unimodal RT-DETR (40.88M) while outperforming both CNN-based and DETR-based state-of-the-art detectors.
| Dataset | Modality | Backbone / Method | mAP (%) | mAP75 (%) | mAP50 (%) | Params (M) | FPS (RTX 3090) |
|---|---|---|---|---|---|---|---|
| FLIR | RGB-IR | DeformCAT [TMM'25] | 46.93 | 43.65 | 86.48 | 120.11 | 27.86 |
| FLIR | RGB-IR | GM-DETR [CVPRW'24] | 45.80 | 42.60 | 83.90 | 76.77 | 23.51 |
| FLIR | RGB-IR | DAMSDet [ECCV'24] | 49.29 | 48.10 | 86.65 | 78.91 | 8.55 |
| FLIR | RGB-IR | ZPA-MDETR (Ours) | 49.82 | 48.32 | 87.05 | 40.88 | 30.20 |
| M3FD | RGB-IR | DeformCAT [TMM'25] | 46.49 | 49.78 | 73.28 | 120.11 | - |
| M3FD | RGB-IR | GM-DETR [CVPRW'24] | 45.72 | 47.15 | 71.25 | 76.77 | - |
| M3FD | RGB-IR | MM-DETR [arXiv'25] | - | - | 73.39 | - | - |
| M3FD | RGB-IR | ZPA-MDETRโ (Ours) | 50.58 | 53.18 | 77.26 | 40.88 | - |
| LLVIP | RGB-IR | DeformCAT [TMM'25] | 66.13 | 77.07 | 97.82 | 120.11 | - |
| LLVIP | RGB-IR | MS-DETR [TITS'24] | 66.10 | 76.30 | 97.90 | 40.88+ | - |
| LLVIP | RGB-IR | ZPA-MDETR (Ours) | 69.14 | 78.79 | 98.11 | 40.88 | - |
| RGBTDronePerson | RGB-IR | COXNet [TCSVT'26] | - | - | 50.04 | - | - |
| RGBTDronePerson | RGB-IR | GM-DETR [CVPRW'24] | 20.63 | 9.37 | 54.97 | 76.77 | - |
| RGBTDronePerson | RGB-IR | ZPA-MDETR (Ours) | 21.01 | 10.75 | 57.13 | 40.88 | - |
(Note: โ indicates the configuration where queries are derived from IR and memory from RGB.)
Ablation Study¶
Ablation on FLIR investigates the contributions of Asymmetric Input Organization (AIO) and Semantic Alignment (SA, spatial and channel):
| Configuration / Components | AIO | Spatial SA | Channel SA | mAP (%) | mAP75 (%) | mAP50 (%) | Params (M) |
|---|---|---|---|---|---|---|---|
| Unimodal baseline (RGB-only) | - | - | - | 33.88 | 27.61 | 69.47 | 40.88 |
| IR-Query / RGB-Memory + AIO | โ | - | - | 47.62 | 44.70 | 85.82 | 40.88 |
| IR-Query / RGB-Memory + AIO + Spa. | โ | โ | - | 47.95 | 44.86 | 86.51 | 40.88 |
| IR-Query / RGB-Memory + AIO + Cha. | โ | - | โ | 47.80 | 44.77 | 86.53 | 40.88 |
| IR-Query / RGB-Memory Full model | โ | โ | โ | 48.16 | 44.89 | 86.65 | 40.88 |
| Unimodal baseline (IR-only) | - | - | - | 45.19 | 41.84 | 81.71 | 40.88 |
| RGB-Query / IR-Memory + AIO | โ | - | - | 49.21 | 47.53 | 86.58 | 40.88 |
| RGB-Query / IR-Memory + AIO + Spa. | โ | โ | - | 49.48 | 47.77 | 86.66 | 40.88 |
| RGB-Query / IR-Memory + AIO + Cha. | โ | - | โ | 49.58 | 47.94 | 86.52 | 40.88 |
| RGB-Query / IR-Memory Full model | โ | โ | โ | 49.82 | 48.32 | 87.05 | 40.88 |
Key Findings¶
- AIO is the primary driver of multimodal gains: Introducing AIO alone elevates mAP from 45.19% to 49.21% (+4.02%) over the IR baseline, and by +13.74% over the RGB baseline, proving that the DETR query-memory mechanism inherently integrates cross-modal information.
- Spatial and channel alignment provide complementary regularization: Spatial SA (+0.27%) and channel SA (+0.37%) individually improve performance, while their combination delivers a +0.61% gain. t-SNE visualizations confirm that semantic alignment bridges the modality gap, collapsing distinct modality clusters into unified semantic representations.
- Reliability-guided role selection matches empirical performance: Assigning the cleaner, more robust modality as memory and the noisier modality as queries yields optimal detection, corroborating that memory should act as reliable evidence while queries serve as exploratory probes.
Highlights & Insights¶
- Counter-intuitive zero-parameter paradigm: Disproves the common assumption that multimodal detection requires complex parameter-heavy fusion modules, showing that single-modal DETRs inherently possess multimodal capability.
- Cost-free training-time alignment: Unifies spatial cosine similarity and Barlow-Twins-like cross-channel decorrelation as training regularizers, bridging representation gaps without inference latency or parameter expansion.
- Superior deployment trade-off: Delivers 49.82% mAP on FLIR at 30.20 FPS with only 40.88M parameters, achieving 3.5ร faster inference and half the parameters of DAMSDet (78.91M, 8.55 FPS).
Limitations & Future Work¶
- Higher computational complexity (FLOPs): While parameter scale is identical to unimodal DETR, batch concatenation doubles encoder throughput, increasing FLOPs from 67.84G to 129.03G and reducing FPS from 50.37 to 30.20 compared to unimodal processing.
- Static image-level role assignment: The query-memory modality role is statically fixed per dataset, lacking the flexibility to dynamically reassign roles when illumination fluctuates across local image regions.
- Future directions: Investigating instance-level adaptive role selection and dynamic token pruning in the shared encoder to reduce multi-stream FLOPs.
Related Work & Insights¶
- vs DAMSDet [ECCV 2024]: DAMSDet incorporates competitive query selection and specialized multispectral deformable cross-attention modules, requiring 78.91M parameters and yielding only 8.55 FPS. ZPA-MDETR retains the original RT-DETR topology without extra parameters, achieving higher accuracy at 30.20 FPS.
- vs GM-DETR [CVPRW 2024]: GM-DETR relies on dedicated cross-scale fusion encoders and modality interaction units (76.77M parameters). ZPA-MDETR demonstrates that asymmetric input organization and parameter-free semantic alignment allow a standard unimodal detector to implicitly achieve superior cross-modal fusion.
Rating¶
- Novelty: โญโญโญโญโญ Elegant and counter-intuitive; proves single-modal DETR is secretly a multispectral detector.
- Experimental Thoroughness: โญโญโญโญโญ Evaluated across four diverse benchmarks and unaligned datasets with comprehensive ablations and visualizations.
- Writing Quality: โญโญโญโญโญ Rigorous motivation, sharp empirical analysis, and well-structured presentation.
- Value: โญโญโญโญโญ Practical and cost-effective, providing an efficient template for edge-device multispectral perception.