Skip to content

SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors

Conference: ECCV 2026
Paper: ECCV Official
Code: https://medem23.github.io/SM
Area: Object Detection
Keywords: multi-object tracking, end-to-end tracking, self-generating detection priors, query interference, attention mass

TL;DR

To tackle query interference between object detection and temporal association in end-to-end transformer trackers, SelfMOTR generates internal 4D anchor priors and confidence-modulated proposal queries via a lightweight detection forward pass, achieving state-of-the-art detector-free tracking performance on DanceTrack and Bird Flock Tracking.

Background & Motivation

Transformer-based end-to-end multi-object tracking (MOT) architectures, pioneered by MOTR, unify new object discovery (detection) and trajectory maintenance (association) into a single differentiable framework by propagating queries temporally across video frames. This paradigm eliminates the hand-crafted heuristics, complex post-processing, and multi-stage tuning characteristic of traditional tracking-by-detection approaches. However, end-to-end trackers suffer from two long-standing bottlenecks: their raw detection accuracy consistently lags behind dedicated standalone detectors trained on the same data, and joint detectionโ€“association optimization induces severe supervision imbalance, where track queries dominate the majority of ground-truth positive assignments while generic detect queries remain comparatively under-trained.

Existing approaches attempt to mitigate this conflict either by incorporating external pretrained detectors (such as YOLOX in MOTRv2 and MOTRv3) to inject bounding-box priors or soft distillation targets, or by maintaining complex auxiliary shadow queries and competitive label assignments as in CO-MOT. While external detectors ease the decoding burden and improve accuracy, they sacrifice the self-contained nature of end-to-end tracking, dramatically increase parameter counts, and introduce feature-space misalignment between heterogeneous networks. Crucially, empirical investigation reveals that when track queries are omitted at inference time, MOTR exhibits a dramatic rebound in detection performance (+6.3 mAP). This confirms that the true performance bottleneck is not a lack of representational capacity, but direct query interference within the shared transformer decoder.

Decoder attention dynamics further reveal that under joint decoding, generic detect queries suffer from severe attention polarization: a substantial portion becomes overwhelmingly track-dominated during intermediate contextualization (mirroring attention sink behaviors observed in large language models), which suppresses independent target discovery, while others become overly localized and lack essential contextual awareness. Rather than importing external priors from an auxiliary model, this work internalizes prior generation. Core idea: leverage the model's own shared encoder features and a lightweight detection-only forward pass to internally self-generate 4D spatial priors, converting them into confidence-modulated proposal queries that are natively parameter-aligned with the tracking decoder, thereby decoupling proposal discovery from association in a fully detector-free architecture.

Method

Overall Architecture

SelfMOTR decouples each video frame's processing into two sequential passes while fully sharing parameters across the backbone, encoder, and transformer decoder. At frame \(t\), the backbone extracts multi-scale visual features. First, a lightweight detection-only pass processes learnable detect queries without track query interference to produce self-generated detection hypotheses. High-confidence detections are then transformed into 4D anchor proposals and modulated by confidence encodings. Second, these proposal queries are concatenated with propagated track queries from the previous frame's Query Interaction Module (QIM) and fed into the shared decoder to perform coordinated bounding box regression, classification, and identity association.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Current Frame Image & Backbone Features"] --> B["Self-Generating Detection Priors<br/>Detection-only pass generates 4D anchors"]
    B --> C["Confidence-Modulated Proposal Queries<br/>4D anchor geometry + periodic confidence encoding"]
    C --> D["Track Query Concatenation<br/>Combined with propagated track queries from QIM"]
    D --> E["Shared Decoder Joint Optimization<br/>Decoupled spatial discovery and association"]
    E --> F["Current Frame Trajectories & New Detections"]

Key Designs

1. Self-Generating Detection Priors: Decoupling Discovery from Association Under joint decoding, detect queries interact with high-affinity track queries before forming coherent spatial hypotheses, causing discovery cues to be prematurely suppressed by existing trajectories. SelfMOTR introduces an internal detection-only forward pass that operates solely on current-frame encoder features using \(N_{\mathrm{det}}\) learnable detect queries. This pass yields initial bounding predictions \(\hat{\mathbf{y}}_k = (\hat{\mathbf{b}}^{\mathrm{det}}_k, \hat{c}^{\mathrm{det}}_k)\), where \(\hat{\mathbf{b}}^{\mathrm{det}}_k \in \mathbb{R}^4\) denotes a 4D anchor coordinate and \(\hat{c}^{\mathrm{det}}_k \in [0, 1]\) represents classification confidence. Shielded from historical track competition, the network's latent detection capacity is fully activated, producing clean, uncorrupted spatial candidates.

2. Confidence-Modulated Proposal Queries: Natively Aligned Query Tokens To translate spatial detections into decoder-compatible tokens, predictions are filtered using a conservative confidence threshold \(c_{\mathrm{prop}}\) (defaulting to 0.05 to maintain high recall), creating a candidate set \(\mathcal{Q} = \{(\hat{\mathbf{b}}^{\mathrm{det}}_k, \hat{c}^{\mathrm{det}}_k) : \hat{c}^{\mathrm{det}}_k > c_{\mathrm{prop}}\}\). Each candidate is converted into an explicit proposal query token \(\mathbf{z}^{\mathrm{prop}}_k\). Its spatial position directly uses the predicted 4D anchor box, \(\mathbf{z}^{\mathrm{prop}}_{\mathrm{pos},k} = \hat{\mathbf{b}}^{\mathrm{det}}_k\), while its content embedding is modulated by fusing a learned shared proposal query \(\mathbf{q}^{\mathrm{prop}}\) with a sine-cosine positional encoding (PE) of the predicted detection confidence: $\(\mathbf{z}^{\mathrm{prop}}_{\mathrm{content},k} = \mathbf{q}^{\mathrm{prop}} + \mathrm{PE}(\hat{c}^{\mathrm{det}}_k)\)$ To guarantee robustness against potential detection misses, a small set of fixed learnable proposal anchors (\(N_{\mathrm{learned}}=10\)) is appended. This construction provides the decoder with accurate geometric initialization and confidence awareness while ensuring full architectural and parametric compatibility without external network dependencies.

3. Shared Decoder Joint Optimization: Balancing Attention Mass and Supervision In the subsequent tracking pass, the confidence-modulated proposal queries \(\mathcal{Q}^{\mathrm{prop}}_t\) are concatenated with propagated track queries \(\mathcal{Q}^{\text{track}}_{t-1}\) and processed jointly by the weight-shared decoder. Because new object hypotheses are pre-localized, the decoder is relieved of blind spatial exploration. Using the proposed diagnostic metric Track Attention Mass (\(\tilde{M}_i\)), the authors demonstrate that while baseline MOTR queries degenerate into extreme bimodal polarization (excessive track sinkage or complete track blindness with degraded Shannon entropy), SelfMOTR maintains an optimal, balanced attention allocation centered around 0.4 with high Shannon entropy (\(\approx 0.85\)). Both passes are trained end-to-end under a joint objective: $\(\mathcal{L} = \mathcal{L}_{\text{MOTR}} + \lambda_{\text{prop}} \mathcal{L}_{\text{prop}}\)$ where \(\mathcal{L}_{\text{prop}}\) combines Focal loss, \(\ell_1\) box regression, and GIoU loss weighted by \(\lambda_{\text{prop}} = 0.5\).

Loss & Training

The network is optimized end-to-end across two unified stages per frame. The detection pass is supervised against current-frame ground-truth boxes via Hungarian bipartite matching. The tracking pass is supervised across video clips (\(N_{\text{clip}}=5\)) using standard Collective Average Loss (CAL) from MOTR, encompassing multi-frame classification and box regression. The model is trained using AdamW with weight decay \(1 \times 10^{-4}\), setting the backbone learning rate to \(2 \times 10^{-5}\) and transformer modules to \(2 \times 10^{-4}\). On DanceTrack, joint training with CrowdHuman static image synthesis is performed for 10 epochs, decaying the learning rate by a factor of 10 at epoch 8, while incorporating QIMv2 and query denoising.

Key Experimental Results

Main Results

On the challenging DanceTrack test benchmarkโ€”characterized by severe non-linear motion and uniform appearanceโ€”SelfMOTR achieves top-tier results among all end-to-end trackers, matching or outperforming prior methods without requiring any external detection models.

Tracker Family Method HOTA โ†‘ DetA โ†‘ AssA โ†‘ IDF1 โ†‘ MOTA โ†‘
Non End-to-End ByteTrack 47.4 71.0 32.1 53.9 89.6
Non End-to-End OC-SORT 55.1 80.3 38.3 54.6 92.0
Non End-to-End MOTRv2 (w/ YOLOX-X) 69.9 83.0 59.0 71.7 91.9
End-to-End TransTrack 45.5 75.9 27.5 45.2 88.4
End-to-End MOTR (Baseline) 54.2 73.5 40.2 51.5 79.7
End-to-End MeMOTR 68.5 80.5 58.4 71.2 89.9
End-to-End SambaMOTR 67.2 78.8 57.5 70.0 88.1
End-to-End MOTRv3 68.3 โ€“ โ€“ 70.1 91.7
End-to-End CO-MOT 69.4 82.1 58.9 71.9 91.2
End-to-End SelfMOTR (Ours) 69.2 80.9 59.3 72.5 89.9

Across dynamic multi-animal tracking benchmarks, SelfMOTR establishes new state-of-the-art benchmarks: on Bird Flock Tracking (BFT), it leads all evaluated methods with 71.1 HOTA, and on AnimalTrack it attains 45.5 HOTA and 53.7 IDF1 (+14.5 HOTA over TrackFormer).

Ablation Study

1. Detectionโ€“Association Conflict Ablation (DanceTrack Validation Set)

Setup Track Queries Active AP โ†‘ AP50 โ†‘ AP75 โ†‘ FPS (Tesla L40S) โ†‘
MOTR Baseline โœ— 66.3 85.5 70.7 โ€“
MOTR Baseline โœ“ 60.0 (-6.3) 78.1 (-7.4) 63.7 (-7.0) 24.8
SelfMOTR (Ours) โœ— 71.2 90.6 76.6 โ€“
SelfMOTR (Ours) โœ“ 70.9 (-0.3) 89.4 (-1.2) 76.3 (-0.3) 20.7

2. Shared Decoder vs. Dual Decoder Design (DanceTrack Test Set)

Decoder Configuration HOTA โ†‘ DetA โ†‘ AssA โ†‘ IDF1 โ†‘ MOTA โ†‘
Dual Decoder (Separate Parameters) 66.2 80.6 54.5 69.1 90.0
Shared Decoder (SelfMOTR) 69.2 80.9 59.3 72.5 89.9

3. Comparison of Detection Prior Injection Mechanisms (DanceTrack Validation Set)

Prior Injection Strategy HOTA โ†‘ DetA โ†‘ AssA โ†‘ IDF1 โ†‘ MOTA โ†‘
MOTR Baseline 51.2 68.8 38.4 49.1 74.4
+ Detector Pretraining (Full Transfer) 51.5 (+0.3) 69.0 (+0.2) 38.7 (+0.3) 49.4 (+0.3) 73.8 (-0.6)
+ Query Pretraining (Frozen Queries) 52.7 (+1.5) 68.0 (-0.8) 41.1 (+2.7) 51.9 (+2.8) 72.5 (-1.9)
+ Teacher Distillation 52.1 (+0.9) 72.0 (+3.2) 37.9 (-0.5) 50.8 (+1.7) 80.1 (+5.7)
+ Anchor Proposal (4D Box Transfer) 58.0 (+6.8) 71.2 (+2.4) 47.5 (+9.1) 59.5 (+10.4) 79.2 (+4.8)

Key Findings

  • Elimination of Query Conflict: While baseline MOTR suffers a severe \(-6.3\) AP penalty when track queries are enabled, SelfMOTR experiences a negligible drop of only \(-0.3\) AP, verifying that decoupling candidate discovery from association resolves query interference.
  • Superiority of Decoder Weight Sharing: Sharing weights across the detection and tracking passes improves AssA by \(+4.8\) and IDF1 by \(+3.4\) over an unshared dual-decoder counterpart, demonstrating that shared representations align discovery and association spaces.
  • Early Depth Saturation: Analysis across decoder depths reveals that the detection pass saturates rapidly after 3-4 layers, allowing practitioners to minimize computational overhead while extracting full performance benefits.

Highlights & Insights

  • Self-Contained Transformer Tracking: Dispenses with heavy external detector backbones (such as YOLOX's 99M parameters in MOTRv2), delivering equal or better association metrics with a compact 42M parameter footprint.
  • Track Attention Mass Diagnostic: Adapting attention sink theory from large language models provides a rigorous quantitative framework for diagnosing query imbalance and degradation in multi-task transformer decoders.
  • Strong Generalization on Deformable Targets: Outperforms all competitors on bird flocking and wildlife benchmarks where appearance cues are unreliable, highlighting the value of geometry-aligned internal proposals for complex motion modeling.

Limitations & Future Work

  • Detection Upper Bound: Raw proposal AP remains lower than models coupled with massive dedicated detectors (71.2 vs. 79.7 for YOLOX in MOTRv2), leaving a minor detection accuracy gap.
  • Sequential Decoder Latency: Executing two decoder passes incurs an inference penalty of roughly 4ms per frame (20.7 vs. 24.8 FPS). Future exploration could investigate single-pass attention masking or early-exit decoders.
  • vs. MOTR / TrackFormer: Prior end-to-end architectures mix detection and track queries within a single decoding stage, causing attention sink collapse; SelfMOTR resolves this internally via decoupled dual-pass proposal generation.
  • vs. MOTRv2 / MOTRv3: External detector coupling inflates model complexity (141M vs. 42M params) and causes cross-architecture feature mismatch; SelfMOTR relies entirely on self-generated priors while surpassing MOTRv2 in association metrics.
  • vs. CO-MOT: CO-MOT relies on intricate shadow sets and competitive label matching logic; SelfMOTR provides a cleaner, geometry-based anchor initialization that directly addresses decoder attention imbalance.

Rating

  • Novelty: โญโญโญโญโ˜† Elegantly uncovers and exploits latent transformer detection capacity, replacing external detectors with native self-proposals.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluations across human, avian, and wildlife datasets with convincing ablation and attention diagnostics.
  • Writing Quality: โญโญโญโญโญ Well-structured, clear analytical insights, and rigorous empirical validation.
  • Value: โญโญโญโญโญ Provides a standard, detector-free blueprint for future unified end-to-end tracking architectures.