Skip to content

Motion-aware Sparse Pipeline for Lightweight Object Tracking

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/TsingWei/MaST
Area: Video Understanding
Keywords: Visual Object Tracking, Token Sparsification, Motion Prior, Lightweight Vision Transformer, Sparse Prediction Head

TL;DR

Addressing the issues of delayed pruning caused by noisy early attention and computation waste in dense prediction heads, MaST injects a Gaussian motion prior into first-layer token scoring and deploys a score-first, regress-once sparse head, setting new SOTA for lightweight tracking at nearly double the edge inference speed.

Background & Motivation

One-stream Vision Transformer trackers jointly model the interaction between template and search regions through self-attention, achieving remarkable tracking precision across standard benchmarks. However, the quadratic computational complexity of self-attention over long token sequences imposes heavy computational burdens, severely hindering deployment on resource-constrained edge platforms such as UAVs and mobile robots. While existing lightweight architectures and dynamic adaptive computation reduce parameter counts or average frame latencies, they either sacrifice discriminative tracking accuracy or introduce variable per-frame runtimes that complicate real-time hardware scheduling.

Leveraging the discrete nature of Vision Transformer patch sequences to prune redundant tokens offers an intuitive path to acceleration. Nonetheless, existing token sparsification methods suffer from two systematic bottlenecks. First, early-layer cross-attention maps are notoriously diffuse and noisy without localized semantic convergence, forcing existing models to defer pruning conservatively to intermediate layers (e.g., progressively at layers 4, 7, and 11) and leaving computationally intensive early stages fully dense. Second, downstream prediction heads predominantly rely on stacked convolutional layers that expect a regular 2D feature grid; this forces sparse tokens to be padded and reshaped back into dense feature maps, incurring redundant operations on empty spatial positions and risking catastrophic bounding box drift if target-center tokens are pruned.

The core tension lies in the fact that raw appearance-based cross-attention in shallow layers cannot reliably identify crucial target tokens, while dense prediction heads substantially negate the compute savings achieved by sparse backbones. Diagnostic experiments demonstrate that if ground-truth spatial guidance is provided, one-shot pruning at Layer 1 with 33% retention maintains tracking accuracy almost without degradation. Exploiting the inherent temporal motion continuity across adjacent frames in tracking sequences, this paper injects prior trajectory cues directly into early token selection and redesigns the prediction head. Core idea: inject a Gaussian motion prior from the previous frame into early cross-attention scores to achieve reliable one-shot token sparsification at encoder Layer 1, paired with a natively sparse score-first, regress-once prediction head that directly decodes bounding boxes from unstructured tokens without dense reshaping.

Method

Overall Architecture

The MaST pipeline operates across three primary stages: input patch embedding with joint template-search encoding, motion-aware shallow token sparsification, and target decoding from unstructured sparse tokens. Given a template image \(Z \in \mathbb{R}^{3 \times H_z \times W_z}\) and a search image \(X \in \mathbb{R}^{3 \times H_x \times W_x}\), both images are divided into non-overlapping patches and mapped to linear token embeddings. Once fed into the Transformer encoder, the sequence is immediately filtered at Layer 1 by the general sparsification module. All subsequent encoder layers compute exclusively over this compressed token subset, and a fully sparse head decodes the final bounding box.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Template and Search Patch Sequence"] --> B["Motion-Aware One-Shot Sparsification<br/>Fuses Shallow Attention and Gaussian Motion Prior"]
    B --> C["Sparse Transformer Encoder Layers<br/>Processes Only Top-K Retained Tokens"]
    C --> D["Score-First Regress-Once Sparse Head<br/>Score Branch Evaluates Retained Anchors"]
    D --> E["Single-Instance Bounding Box Regression<br/>Direct Decode from Highest-Confidence Anchor"]

Key Designs

1. Motion-Aware One-Shot Sparsification: Overcoming Early Attention Diffusion

Addressing the severe noise and diffuse distributions of raw early cross-attention maps that prevent accurate target-background separation, this module explicitly incorporates temporal motion continuity from prior frames. Inside the encoder, search tokens are formulated as queries \(Q_x \in \mathbb{R}^{P_x \times d}\) and template tokens as keys \(K_z \in \mathbb{R}^{P_z \times d}\). To minimize computation, the center patch token of the template \(k_c \in \mathbb{R}^d\) serves as a compact exemplar, and its scaled dot-product against search token \(q_i\) yields the raw appearance importance score \(s_i\). Simultaneously, the predicted box from the preceding frame \(b_{t-1} = (x_{t-1}, y_{t-1}, w_{t-1}, h_{t-1})\) defines a 2D spatial Gaussian motion window:

\[G_t(u, v) = \exp \left( -\frac{(u - x_{t-1})^2}{2\sigma_x^2} - \frac{(v - y_{t-1})^2}{2\sigma_y^2} \right)\]

where \((u, v)\) denotes the patch coordinate, and standard deviations are set to \(\sigma_x = \gamma w_{t-1}, \sigma_y = \gamma h_{t-1}\) with scaling factor \(\gamma = 0.5\). The refined importance score combines cross-attention and spatial proximity via element-wise multiplication:

\[w_i = G_t(u_i, v_i) \cdot s_i\]

This formulation suppresses background noise and concentrates retention on tokens within the probable kinematic trajectory. Consequently, the encoder performs reliable one-shot selection of the top-\(K\) search tokens \(\mathcal{L}_K \in \mathbb{R}^{N_K \times d}\) directly at Layer 1, avoiding the overhead of multi-stage progressive pruning.

2. Score-First Regress-Once Sparse Head: Eliminating Dense Grid Reshaping

Addressing the bottleneck where conventional convolutional heads require padding and reshaping sparse tokens back to a dense 2D grid, this module uses lightweight MLPs acting directly on the unstructured token set \(\mathcal{L}_K = \{\mathbf{f}_k\}_{k=1}^{N_K}\) alongside their stored 2D anchor coordinates \(p_k = (u_k, v_k)\). The prediction logic decouples into a coarse-scoring phase followed by a single-point regression step.

The score branch \(g_s\) first computes a scalar confidence \(s_k \in \mathbb{R}\) across all retained tokens, identifying the primary target token via \(k^* = \arg\max_k s_k\). Next, the regression branch \(g_r\) executes only once on the selected token feature \(\mathbf{f}_{k^*}\) to estimate displacement parameters \(\Delta_{k^*}\). The final box is directly calculated relative to its stored spatial anchor: \(\hat{\mathbf{b}} = \mathrm{Decode}(\mathbf{p}_{k^*}, \Delta_{k^*})\). Because scoring scales linearly with the reduced token count \(N_K\) and regression runs in constant time \(\mathcal{O}(1)\), the head avoids dense feature padding and eliminates redundant computation on background locations.

Loss & Training

MaST is trained end-to-end using an AdamW optimizer with a weight decay of \(10^{-4}\). The learning rate is initialized at \(4 \times 10^{-5}\) for the backbone and \(4 \times 10^{-4}\) for the head and sparsification parameters. Training runs for 300 epochs with a batch size of 128 on a single NVIDIA RTX 3090 GPU, with a 10ร— learning rate drop after epoch 240. The token retention rate is warmed up progressively during early training to ensure smooth representation learning. The total training objective combines classification loss and regression losses:

\[L_{\text{head}} = L_{\text{cls}}(\{s_k\}) + \lambda_{\ell_1} L_{\ell_1}(\hat{\mathbf{b}}_{k^{\text{gt}}}, \mathbf{b}) + \lambda_{\text{GIoU}} L_{\text{GIoU}}(\hat{\mathbf{b}}_{k^{\text{gt}}}, \mathbf{b})\]

where \(k^{\text{gt}} = \arg\min_k \|\mathbf{p}_k - \mathbf{c}\|\) identifies the sparse anchor closest to the ground-truth center \(\mathbf{c}\), and the loss weights are fixed to \(\lambda_{\ell_1} = 5\) and \(\lambda_{\text{GIoU}} = 2\).

Key Experimental Results

Main Results

On LaSOT, TrackingNet, and GOT-10k, MaST model variants (nano, tiny, small) establish superior speedโ€“accuracy trade-offs against competitive lightweight trackers on edge hardware:

Model Venue/Year MACs (G) LaSOT AUC LaSOT PNorm TrackingNet SUC TrackingNet PNorm GOT-10k AO RPi 5 (FPS) Jetson Nano (FPS)
MaST-nano Ours 0.585 58.6 67.7 77.2 82.5 61.6 30.1 230
AsymTrack-T AAAI 2025 0.708 60.8 68.7 76.2 80.9 62.3 18.1 99
FEAR-XS ECCV 2022 0.532 53.5 - - - 61.9 19.5 146
LightTrack CVPR 2021 0.483 53.8 - 72.5 77.8 61.1 13.4 124
MaST-tiny Ours 0.836 63.8 72.2 80.1 85.3 66.6 22.6 152
AsymTrack-S AAAI 2025 0.806 62.8 71.2 77.9 82.2 65.5 15.6 88
FARTrackpico ICLR 2026 1.080 58.6 67.1 75.6 81.3 62.8 17.2 134
HiT-Small ICCV 2023 1.130 60.5 68.3 77.7 81.9 62.6 21.5 106
MaST-small Ours 1.820 65.8 74.7 82.3 87.1 70.0 7.5 98
FERMT ECCV 2024 2.310 65.1 74.6 80.8 80.9 69.6 7.9 84
FARTracktiny ICLR 2026 2.650 63.2 71.6 80.7 85.6 70.6 6.9 87
AsymTrack-B AAAI 2025 1.810 64.7 73.0 80.0 84.5 67.7 6.2 57

Ablation Study

The following tables evaluate token sparsification strategies and prediction head designs based on a ViT-Tiny backbone evaluated on LaSOT and edge devices:

1. Token Sparsification Criteria Comparison (LaSOT & Computational Efficiency)

Sparsification Strategy LaSOT AUC LaSOT PNorm Encoder MACs (M) RPi 5 (FPS) Orin Nano (FPS)
None (Dense Baseline) 64.0 74.2 1752 9.1 94
Attention Only 60.5 71.9 824 22.9 157
Auxiliary Prediction Network 59.4 71.4 874 21.5 138
Motion Window Only 62.6 72.2 824 23.5 169
Attention + Motion (Ours) 63.8 73.6 824 22.6 152

2. Prediction Head Architecture Comparison (LaSOT & Complexity)

Prediction Head Paradigm Dense AUC Sparse AUC Dense MACs (G) Sparse MACs (G) RPi 5 (FPS)
Loc Tokens (MixFormerV2) Global Localization 59.0 57.9 1.76 0.845 23.9
Transformer Decoder Global Localization 61.4 60.1 1.88 0.833 22.7
3ร—3 Conv-stacked (OSTrack) Anchor Dense Regression 65.0 64.1 2.39 1.463 13.8
MLP Dense (LoRAT) Anchor Dense Regression 64.0 63.8 1.81 0.872 21.3
MLP Sparse (Ours) Score-First Regress-Once 64.0 63.8 1.75 0.836 23.2

Key Findings

  • Motion prior rescues aggressive early pruning: Pruning tokens using solely shallow attention causes LaSOT AUC to plummet from 64.0 to 60.5. Integrating the Gaussian motion window recovers performance to 63.8 AUC at Layer 1โ€”nearly matching the dense upper bound while yielding a 2.5ร— speedup on edge hardware.
  • Layer 1 sparsification maximizes computational return: Moving pruning from Layer 6 up to Layer 1 causes an insignificant 0.14 AUC drop (63.82 vs 63.96), yet doubles Raspberry Pi frame rates from 10.1 to 22.6 FPS, proving the immense benefit of early compute reduction.
  • Score-first regress-once avoids dense restructuring penalties: While traditional 3ร—3 convolutional heads bottleneck throughput at 13.8 FPS due to zero-padding and grid reshaping, the natively sparse MLP head delivers 23.2 FPS with zero accuracy loss (63.8 AUC).

Highlights & Insights

  • Translating temporal continuity into spatial compute pruning: Repurposing classical output-level penalty windows into early-stage token selection addresses the long-standing limitation of diffuse shallow attention at minimal computational cost.
  • Natively sparse head prevents compute leakage: By retaining original patch grid coordinates as sparse anchor tokens, the regression stage is reduced to constant-time lookup \(\mathcal{O}(1)\), eliminating dense feature reshaping.
  • Hardware-friendly deterministic latency: Unlike dynamic token dropping methods with variable frame-to-frame overhead, MaST enforces fixed top-\(K\) token retention budgets, ensuring stable frame scheduling on edge devices.

Limitations & Future Work

  • Limitations acknowledged by authors: When fed high-resolution inputs (e.g., 384ร—384 search images), patch embedding and Layer 1 self-attention still incur significant computational overhead before sparsification occurs.
  • Assumptions and edge cases: The Gaussian motion window assumes smooth object trajectories; abrupt camera shaking or prolonged target occlusions may introduce an inaccurate spatial bias toward historical locations.
  • Potential improvements: Pre-embedding token downsampling and adaptive motion window scaling via Kalman filtering could further boost robustness under rapid movements.
  • vs OSTrack (ECCV 2022): OSTrack performs progressive multi-stage pruning across intermediate layers and retains a dense convolutional head. MaST moves pruning directly to Layer 1 via motion guidance and introduces a sparse prediction head, achieving nearly double the edge throughput.
  • vs AsymTrack (AAAI 2025): AsymTrack compresses template computations using asymmetric two-stream architectures, but still processes search tokens densely. MaST achieves superior hardware frame rates by pruning search tokens at Layer 1 and executing single-instance regression.
  • vs LoRAT (ECCV 2024): LoRAT introduces MLP heads to accelerate dense tracking. MaST adapts MLP processing to non-grid token sets with a score-first, regress-once pipeline, validating native sparse box regression without grid restoration.

Rating

  • Novelty: โญโญโญโญโ˜† Combines temporal Gaussian priors with an unstructured sparse prediction head to resolve early-layer token pruning bottlenecks.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive benchmarks across LaSOT, TrackingNet, GOT-10k, UAV123, NFS, and VastTrack, backed by extensive real-device latency tests on Raspberry Pi 5, CPU, and Jetson Orin Nano.
  • Writing Quality: โญโญโญโญโญ Well-structured narrative with intuitive diagnostic experiments (Fig. 2) motivating the core design choices.
  • Value: โญโญโญโญโญ Delivers a practical, Pareto-optimal sparse pipeline blueprint for deploying Vision Transformer trackers on edge devices.