Skip to content

SOVTrack: Open-Vocabulary Multi-Object Tracking with Self-Supervised Pseudo Labeling and Feature Distillation

Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Code: https://github.com/zekunqian/SOVTrack
Area: Video Understanding
Keywords: Open-Vocabulary Tracking, Multi-Object Tracking, Self-Supervised Learning, Knowledge Distillation, SAM2

TL;DR

Addressing poor generalization on unseen categories and insufficient temporal modeling in Open-Vocabulary Multi-Object Tracking (OVMOT), SOVTrack introduces detection-guided bidirectional cycle consistency to extract high-quality pseudo-labels with difficulty scores, coupled with a decoupled dual-branch architecture that adversarially distills multi-cue SAM2 representations, achieving state-of-the-art tracking on TAO using only 10K raw video frames.

Background & Motivation

Multi-Object Tracking (MOT) is transitioning from tracking closed-set categories like pedestrians and vehicles to detecting and tracking arbitrary objects across unconstrained real-world environments, giving rise to Open-Vocabulary Multi-Object Tracking (OVMOT). However, prevailing OVMOT methodologies continue to adhere to conventional training paradigms: they train appearance-based target association models exclusively under base-class supervision. Consequently, during inference, these trackers suffer severe performance degradation when encountering unseen or rare novel categories. While open-vocabulary object detection and segmentation have flourished by borrowing generalizable representations from large vision-language models, the tracking domain has struggled to trigger a comparable breakthrough in temporal target association.

Recent pioneering attempts such as MASA utilize single-image Segment Anything (SAM) models with geometric augmentations to construct paired pseudo-labels from static images. Nonetheless, this paradigm suffers from two major deficiencies. First, synthetic image transformations entirely omit genuine temporal dynamics, continuous motion trajectories, and realistic object occlusions found in physical video sequences. Furthermore, filtering pseudo-labels solely by static segmentation thresholds lacks temporal consistency validation, easily propagating tracking drift while failing to gauge sample difficulty. Second, supervising only target-level contrastive loss wastes the high-order temporal and contextual representations embedded in foundation models, delivering modest performance gains even after processing 500K synthetic image pairs.

To resolve these tensions, SOVTrack shifts from single images to raw video clips, tapping into the native spatiotemporal tracking capabilities of SAM2. Because SAM2 operates as a prompt-guided pixel-level video segmentation model, naively converting its auto-segmentation outputs into bounding boxes yields fragmented masks, cross-object errors, and severe false positives. Core idea: employ detector-guided bidirectional forward-backward VOS validation to filter pseudo-labels and quantify per-frame tracking difficulty, then deploy a decoupled dual-branch architecture that combines appearance contrastive learning with multi-cue adversarial distillation to transfer SAM2's universal temporal representations into an object-level tracker using minimal video data.

Method

Overall Architecture

SOVTrack comprises two synergistic components: Dual-Direction Pseudo Labeling (DDPL) for self-supervised data mining and a Dual-Branch Association architecture for model training. During offline data preparation, DDPL utilizes the tracker's detector to identify candidate boxes, initiates forward SAM2 VOS tracking across \(K\) frames, extracts bounding boxes from the final frame to run backward VOS, and validates trajectory reliability via cycle IoU closure while recording per-frame overlap as difficulty scores. During online training, the dual-branch framework shares a ResNet-50 backbone: the contrastive branch focuses on discriminative appearance similarity learning supervised by validated pseudo-labels, while the adversarial branch projects detector classification, localization, and implicit association features into SAM2's pointer space via Multi-Cue Adversarial Distillation (MCAD). Finally, an adaptive fusion mechanism dynamically balances both branches based on the predicted frame difficulty.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Raw Unlabeled Video Sequence"] --> B["Detector-Guided Initialization<br/>Extract Salient Object Box Prompts"]
    B --> C["Bidirectional Cycle Consistency Validation<br/>Forward-Backward VOS Trajectory Verification"]
    C -->|Cycle IoU > Threshold| D["Pseudo-label Generation & Difficulty Assessment"]
    D --> E["Dual-Branch Association Architecture<br/>Shared ResNet-50 Backbone"]
    E --> F["Contrastive Branch: Discriminative Appearance Learning"]
    E --> G["Adversarial Branch: Multi-Cue Adversarial Distillation (MCAD)"]
    F --> H["Adaptive Feature Fusion<br/>Confidence-Guided Dynamic Weighting"]
    G --> H
    H --> I["Robust Open-Vocabulary Association Feature"]

Key Designs

1. Detector-Guided Bidirectional Cycle Consistency: Unsupervised Initialization and Drift Pruning Directly applying SAM2's auto-segmentation to video frames produces fragmented, noisy masks that blur object boundaries under complex lighting or fast motion. SOVTrack instead uses the tracker's pretrained detector to identify salient object proposals \(\mathbf{P}^t = \{\mathbf{b}_i^t \mid \mathbf{d}_i^t \in \mathcal{D}^t, s_i^t > \theta_{\text{det}}\}\) at frame \(t\). SAM2 performs forward VOS tracking across \(K\) subsequent frames to produce trajectory \(\mathcal{T}^{\text{fwd}}\). The bounding box from the terminal mask at \(t+K\) then serves as the prompt for backward VOS, yielding \(\mathcal{T}^{\text{bwd}}\). The cycle consistency is evaluated via the spatial IoU at the starting frame: $\(\mathcal{S}_{\text{cycle}}^{i,t} = \text{IoU}\left(\mathbf{b}_i^t, \mathbf{b}(m_i^{t,\text{bwd}})\right)\)$ Trajectories exceeding the threshold \(\theta_{\text{cycle}}\) are preserved as reliable pseudo-labels, effectively eliminating false positive detector proposals and wandering segmentation traces without human intervention.

2. Fine-Grained Per-Frame Tracking Difficulty Assessment: Dynamic Priors for Multi-Modal Fusion While cycle consistency guarantees global trajectory reliability, object states fluctuate across intermediate frames due to partial occlusions, deformation, and lighting variations. Rather than treating all frames uniformly, SOVTrack computes the spatial overlap between forward and backward bounding boxes at each intermediate time step \(t+k\): $\(\mathcal{C}_i^{t+k} = \text{IoU}\left(\mathbf{b}(m_i^{t+k,\text{fwd}}), \mathbf{b}(m_i^{t+k,\text{bwd}})\right)\)$ This score serves as a continuous difficulty metric: a high score indicates clear, stable visibility where standard appearance matching suffices, whereas a low score flags localized deformation or occlusion where deeper semantic context is indispensable.

3. Multi-Cue Adversarial Distillation (MCAD): Transferring Foundation Representations to Object Trackers To capture the rich multi-cue object pointer representations generated by SAM2's cross-attention module \(\mathbf{h}_{\text{SAM2}}^{i,t+k} = \text{CrossAttn}(\mathbf{f}_{\text{img}}^{t+k}, \mathbf{f}_{\text{prompt},i}, \mathbf{f}_{\text{memory}}^{i,t+k})\), the adversarial branch extracts classification cues \(\mathbf{f}_{\text{cls}}\), localization cues \(\mathbf{f}_{\text{loc}}\), and implicit association cues \(\mathbf{f}_{\text{assoc}}\) from the detector heads. After linear projection into a unified dimension, these cues are merged by an MLP fusion block into \(\mathbf{f}_{\text{fused}}\). Instead of relying solely on pointwise L2 regression (\(\mathcal{L}_{\text{dist}}\)), which tends to over-smooth high-dimensional manifolds, a 3-layer MLP discriminator \(D\) is introduced to distinguish generated features from authentic SAM2 pointers. Minimax adversarial training with \(\mathcal{L}_{\text{adv}} = -\mathbb{E}[\log D(\mathbf{f}_{\text{fused}})]\) forces the tracker's internal features to match the distribution of the foundation model.

4. Difficulty-Aware Adaptive Feature Fusion: Harmonizing Discriminative and Generalized Embeddings The contrastive branch excels at learning fine-grained visual appearance \(\mathbf{f}_{\text{app}}\) on familiar targets, whereas the adversarial branch produces distilled features \(\mathbf{f}_{\text{fused}}\) with superior semantic robustness across unseen classes. An auxiliary prediction layer estimates appearance reliability \(\hat{\mathcal{R}}_{\text{reliability}}^{i,t+k} = \text{sigmoid}(\text{FC}_{\text{pred}}([\mathbf{f}_{\text{cls}}^{\text{proj}}; \mathbf{f}_{\text{loc}}^{\text{proj}}; \mathbf{f}_{\text{assoc}}^{\text{proj}}]))\), supervised by the DDPL confidence score \(\mathcal{C}_i^{t+k}\) via mean squared error \(\mathcal{L}_{\text{reliability}}\). The final association vector dynamically interpolates between the two representations: $\(\mathbf{f}_{\text{final}}^{i,t+k} = \hat{\mathcal{R}}_{\text{reliability}}^{i,t+k}\mathbf{f}_{\text{app}}^{i,t+k} + (1 - \hat{\mathcal{R}}_{\text{reliability}}^{i,t+k})\mathbf{f}_{\text{fused}}^{i,t+k}\)$ A contrastive objective \(\mathcal{L}_{\text{fusion}}\) is subsequently applied to \(\mathbf{f}_{\text{final}}\) to enforce global metric clustering across the entire feature space.

Loss & Training

The framework adopts a two-stage training scheme. Stage 1 trains the detector on LVIS base classes for 20 epochs following standard OVMOT setups, and freezes it to generate DDPL pseudo-labels. Stage 2 jointly optimizes the dual-branch association network on 10K raw video frames across 20 epochs using 4 RTX 3090 GPUs. The overall objective function is: $\(\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{app}} + \mathcal{L}_{\text{MCAD}} + \mathcal{L}_{\text{reliability}} + \mathcal{L}_{\text{fusion}} + \mathcal{L}_{\text{loc}}\)$ where \(\mathcal{L}_{\text{loc}}\) applies Smooth L1 loss using SAM2's high-precision segmentation bounding boxes to fine-tune the tracker's box regression head, while the discriminator is updated alternately via standard binary cross-entropy.

Key Experimental Results

Main Results

Evaluation is conducted on the standard TAO open-vocabulary multi-object tracking benchmark under the LVIS base/novel split, using Tracking Every Thing Accuracy (TETA) along with Localization Accuracy (LocA), Association Accuracy (AssocA), and Classification Accuracy (ClsA). All competing models share a ResNet-50 backbone.

Dataset / Split Method Base TETA Base LocA Base AssocA Novel TETA Novel LocA Novel AssocA
TAO Val OVTrack (CVPR'23) 35.5 49.3 36.9 27.8 48.8 33.6
TAO Val MASA (CVPR'24) 36.9 55.1 36.4 30.0 54.2 34.6
TAO Val SLAck (ECCV'24) 37.2 55.0 37.6 31.1 54.3 37.8
TAO Val OVTR (ICLR'25) 36.6 52.2 37.6 31.4 54.4 34.5
TAO Val TRACT (ICCV'25) 38.5 55.0 39.0 31.3 52.7 37.8
TAO Val SOVTrack (Ours) 40.1 58.1 42.3 35.7 59.0 42.4
TAO Test OVTrack (CVPR'23) 32.6 45.6 35.4 24.1 41.8 28.7
TAO Test SLAck (ECCV'24) 34.7 52.5 35.6 27.1 49.1 30.0
TAO Test TRACT (ICCV'25) 36.2 52.3 39.1 27.3 48.2 30.7
TAO Test SOVTrack (Ours) 39.2 57.3 43.1 28.9 50.5 33.0

In terms of computational efficiency, SOVTrack runs at 12.1 FPS during inference on a single RTX 3090 (close to MASA's 13.4 FPS). During offline preparation, DDPL requires only 3.6 GPU-hours of foundation model computation, representing a \(63\times\) speedup over MASA's full-image segmentation (227.8 GPU-hours), while requiring only 2% of the training data (10K video frames vs. 500K image pairs).

Ablation Study

Ablation experiments on the TAO validation set systematically isolate the contributions of the pseudo-labeling components, architecture branches, and training objectives:

Configuration / Variant Base TETA Base AssocA Novel TETA Novel AssocA Note
SOVTrack (Full Model) 40.1 42.3 35.7 42.4 Complete dual-branch + DDPL framework
w/o Cycle Consistency Filtering 38.1 38.4 32.2 37.6 Noisy pseudo-labels severely degrade novel association (-4.8% AssocA)
SAM Auto-segmentation as Prompts 38.8 40.2 32.5 37.3 Fragmented initial masks drop novel TETA by 3.2%
w/o Contrastive Branch 38.8 38.9 34.4 40.8 Base category discriminative tracking suffers (-1.3% Base TETA)
w/o Adversarial Branch 39.3 40.9 33.0 37.8 Loss of foundation distillation hurts novel generalization (-2.7% Novel TETA)
w/o Adaptive Feature Fusion 38.8 39.9 33.2 39.7 Unweighted combination degrades novel performance (-2.5% Novel TETA)
w/o Adversarial Distillation (L2 only) 39.3 40.7 34.2 40.1 Absence of distribution alignment weakens feature transfer
w/o Localization Supervision \(\mathcal{L}_{\text{loc}}\) 38.9 41.6 32.8 40.9 Removing SAM2 box supervision drops novel LocA by 3.7%

Key Findings

  • Elimination of the Generalization Gap in Association: SOVTrack attains 42.4% AssocA on novel classes, outperforming SLAck by 4.6% and matching its base category performance (42.3%). This proves that distilling SAM2's native spatiotemporal representations allows novel objects to be tracked as reliably as base classes.
  • Cycle Consistency is Essential for Self-Supervised Quality: Disabling bidirectional temporal validation causes novel TETA to drop by 3.5% (35.7% \(\to\) 32.2%) and AssocA to drop by 4.8%, highlighting that temporal cycle closure, rather than sheer data volume, is what prevents drift in unsupervised video learning.
  • Symmetric Branch Complementarity: The contrastive branch primarily bolsters base class discrimination (+1.3% Base TETA), whereas the adversarial distillation branch drives novel class generalization (+2.7% Novel TETA). Guided by the frame difficulty score, adaptive fusion achieves optimal dynamic trade-offs across diverse tracking conditions.

Highlights & Insights

  • Formulating Video Object Tracking as Closed-Loop Hypothesis Verification: By verifying forward and backward VOS trajectories via bounding box overlap, the method filters spurious detections and error drift at near-zero extra computational cost.
  • Multi-Cue Adversarial Distillation Beyond L2 Regression: Transferring fused semantic, localization, and temporal cues through adversarial domain matching avoids the over-smoothing typical of Mean Squared Error, tightly aligning the tracker with SAM2's object pointer manifold.
  • High Data and Training Efficiency: Achieving state-of-the-art results with only 10K raw video frames and 3.6 GPU-hours of SAM2 processing demonstrates that natural temporal continuity provides far more informative tracking supervision than hundreds of thousands of static synthetic images.

Limitations & Future Work

  • Dependence on Detector Initialization: If a novel or small object is completely missed by the pre-trained detector in the prompt frame, the bidirectional cycle verification loop cannot be triggered.
  • Temporal Window Horizon: Bidirectional tracking is restricted to local video windows of length \(K\). Long-term track re-identification and explicit trajectory lifecycle management across minutes-long videos remain open avenues for future exploration.
  • vs. MASA (CVPR 2024): MASA relies on 500K synthetic image pairs generated by augmenting static SAM segmentations, missing real-world motion dynamics and relying solely on single-image contrastive loss. SOVTrack uses raw videos with bidirectional validation and multi-cue adversarial distillation, surpassing MASA while cutting training data by 98%.
  • vs. OVTrack / SLAck (CVPR 2023 / ECCV 2024): Prior works either rely purely on base-class annotations or heuristic cue concatenation. SOVTrack is the first to distill the internal spatiotemporal representations of a video foundation model (SAM2) into a lightweight object tracker, demonstrating an effective paradigm for foundation-model-driven video tracking.

Rating

  • Novelty: โญโญโญโญโ˜† [Principled bidirectional VOS cycle consistency and multi-cue adversarial distillation from SAM2]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive comparisons on TAO benchmarks, full modular ablations, and detailed efficiency breakdowns]
  • Writing Quality: โญโญโญโญโญ [Clear motivation, structured explanations, intuitive diagrams, and restrained mathematical formulation]
  • Value: โญโญโญโญโญ [Sets a new benchmark for data-efficient, foundation-model-distilled open-vocabulary video tracking]