MobileSAM2: Lightweight Segment Anything in Images and Videos via Hypergraphical Knowledge Distillation¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: VLM Efficiency / Model Compression / Segmentation
Keywords: Video Segmentation, Knowledge Distillation, Hypergraph, Lightweight Model, Embodied AI
TL;DR¶
To tackle the excessive parameter scale and compute cost of SAM2 on resource-constrained edge devices, MobileSAM2 introduces Hypergraphical Knowledge Distillation (HyperKD), which explicitly captures high-order temporal associations and multi-granularity visual hierarchies via hypergraphs and pairs with progressive architecture contraction, yielding state-of-the-art efficiency and segmentation quality using only ~10% training data.
Background & Motivation¶
Large vision foundation models, notably the Segment Anything Model (SAM) and its video counterpart SAM2, have fundamentally reshaped zero-shot promptable visual segmentation. Trained on the unprecedented SA-V dataset containing over 100K diverse videos and 40M fine-grained mask annotations, SAM2 acquires robust temporal correlation tracking across video sequences alongside comprehensive multi-granularity conceptual understanding spanning objects, parts, and subparts. Despite these breakthroughs, SAM2 suffers from high memory consumption and latencyโits Hiera-L image encoder alone exceeds 212M parameters, easily inducing out-of-memory (OOM) failures on mobile GPUs and barring real-time deployment on edge platforms such as smartphones, drones, and autonomous robots.
This severe computational burden sharply clashes with the practical demands of edge-based spatial intelligence. Conventional knowledge distillation methods typically rely on point-wise intermediate feature alignment (e.g., FitNet) or pairwise instance affinity matrices. Such formulations inherently fail to represent high-order, multi-entity visual relations that persist across consecutive frames, nor can they effectively inherit the hierarchical multi-granularity abstractions embedded in SAM2. Under severely limited training data and compute budgets (e.g., training with only ~10% of SA-V on a single GPU), transferring both temporal continuity and multi-level granularity without catastrophic performance degradation remains an open challenge.
This work addresses the problem by introducing hypergraph structures capable of naturally modeling multi-way, high-order associations across frames and visual concept levels. Core Idea: Propose Hypergraphical Knowledge Distillation (HyperKD) to explicitly extract patch-wise and instance-wise temporal hypergraphs alongside multi-granularity concept hypergraphs from SAM2 to guide a student encoder, and search an optimal lightweight Hiera family via progressive model contraction to achieve fast, high-fidelity edge segmentation.
Method¶
Overall Architecture¶
MobileSAM2 substitutes SAM2's heavyweight Hiera-L image encoder with a compact, searched lightweight Hiera backbone while keeping the pre-trained memory module (memory attention, memory encoder, memory bank), prompt encoder, and mask decoder frozen (their combined footprint remains under 12M parameters). Distillation training is carried out on unlabeled video sequences using roughly 10% of the SA-V dataset with pre-trained SAM2-L as the teacher. The overall training pipeline integrates point-wise feature mimicry, Temporal HyperKD, and Granularity HyperKD, supervised through hypergraph incidence matrix alignment.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Unlabeled Video Sequences & Prompts"] --> B["Stage 1: Temporal HyperKD<br/>Extract Cross-Frame Patch and Instance Hypergraphs"]
B --> C["Stage 2: Granularity HyperKD<br/>Pool Features by Multi-Level Masks & Build Concept Hyperedges"]
C --> D["Stage 3: Progressive Model Contraction & Search<br/>Identify Pareto-Optimal Lightweight Hiera Backbones"]
D --> E["MobileSAM2 Edge-Friendly Deployment & Inference"]
Key Designs¶
1. Temporal HyperKD: Explicitly Capturing High-Order Cross-Frame Temporal Consistency
Standard feature distillation optimizes localized activations independently, losing contextual inter-frame dynamics. Temporal HyperKD constructs two complementary temporal hypergraphs from the teacher embedding \(z_t^I\): a Patch-wise Temporal Hypergraph \(G_t^{Pa} = (V_t^{Pa}, E_t^{Pa})\) and an Instance-wise Temporal Hypergraph \(G_t^{Ins} = (V_t^{Ins}, E_t^{Ins})\). Patch vertices are sampled directly from spatial grid embeddings, while instance vertices are formed via average pooling guided by high-confidence mask predictions from the teacher. Hyperedges connecting multiple related vertices across frames are constructed via an \(\epsilon\)-ball threshold in feature distance: $\(E = \{\text{ball}(v, \epsilon) \mid v \in V\}\)$ $\(\text{ball}(v, \epsilon) = \{u \mid \text{dist}(x_u^I, x_v^I) < \epsilon, u \in V\}\)$ The incidence matrix \(H(u, v)\) is populated using cosine distance: \(H(u, v) = 1 - \frac{\langle x_u^I, x_v^I \rangle}{\|x_u^I\|\|x_v^I\|}\) when vertex \(u\) falls inside the hyperedge centered at \(v\), and 0 otherwise. The student network builds incidence matrices \(H_s^{Pa}\) and \(H_s^{Ins}\) following the same metric. Distillation minimizes the L1 discrepancy between teacher and student incidence topologies: $\(\mathcal{L}_{THKD} = \|H_t^{Pa} - H_s^{Pa}\|_1 + \|H_t^{Ins} - H_s^{Ins}\|_1\)$
2. Granularity HyperKD: Decoupling and Transferring Hierarchical Concept Representations
Objects in real-world videos exhibit complex hierarchical part-whole compositions, and single-level distillation frequently blurs fine details. Given \(M\) levels of segmentation masks \(m(i, j)\) predicted by the teacher, Granularity HyperKD initializes multi-granularity nodes \(V_t^{Gran}\) via masked feature pooling: $\(\mathcal{V}_t^{Gran} = \{\text{MaskedAveragePooling}(z_t^I(i, j), m(i, j)) \mid j = 1, \dots, M\}\)$ Hyperedges \(E_t^{Gran}\) connecting related nodes across disparate granularity levels (e.g., full object, semantic part, subpart) are generated via the same \(\epsilon\)-ball clustering. The student model constructs its corresponding incidence matrix \(H_s^{Gran}\) using teacher-predicted masks. The granularity distillation loss is formulated as: $\(\mathcal{L}_{GHKD} = \|H_t^{Gran} - H_s^{Gran}\|_1\)$ This formulation delivers orthogonal structural supervision that reinforces the student model's ability to segment arbitrary visual concepts at varying degrees of detail.
3. Progressive Model Contraction & Search: Pareto-Optimal Lightweight Hiera Backbones
Directly substituting off-the-shelf lightweight backbones (such as TinyViT) yields suboptimal synergy with SAM2's multi-scale decoding pipeline. MobileSAM2 parameterizes a dedicated Hiera contraction search space defined by four architectural factors: stage-wise embedding dimensions \(\Gamma_{Emb}\), block depths \(\Gamma_{Blk}\), attention window sizes \(\Gamma_{WinSiz}\), and attention head expansion ratios \(\Gamma_{HExp}\). Guided by the HyperKD objective, progressive contraction iteratively explores compact configurations that preserve maximal temporal and multi-granularity capacity, producing three model scales: MobileSAM2-5M (5.84M), MobileSAM2-10M (10.37M), and MobileSAM2-23M (23.74M).
Loss & Training¶
The overall objective function combines point-wise intermediate feature alignment with the two hypergraphical distillation losses: $\(\mathcal{L}_{total} = \|z_t^I - z_s^I\|_1 + \alpha \cdot \mathcal{L}_{THKD} + \beta \cdot \mathcal{L}_{GHKD}\)$ Empirically, balancing weights are set to \(\alpha = 1\) and \(\beta = 1\). The model is trained on 11K videos from the SA-V manual subset (~10% of total data) using the AdamW optimizer with an initial learning rate of \(5 \times 10^{-4}\). Training spans 5 epochs on a single NVIDIA A100 (80GB) GPU with batch size 5 (8-frame sequences per sample), completing in approximately 65 hours.
Key Experimental Results¶
Main Results¶
MobileSAM2 is benchmarked across two standard VOS datasets (MOSE val, DAVIS 2017 val), a long-term benchmark (LVOS val), and open-world video benchmarks (SA-V val, SA-V test).
| Model | Distillation Method | Params. (M) | MOSE val (J&F) | DAVIS 2017 val (J&F) | LVOS val (J&F) | SA-V val (J&F) | SA-V test (J&F) |
|---|---|---|---|---|---|---|---|
| SAM2 Base+ (100% Data / 256 GPUs) | - | 68.7 | 72.8 | 88.8 | 75.8 | 72.2 | 74.7 |
| SAM2 Large (100% Data / 256 GPUs) | - | 212.1 | 74.6 | 89.2 | 81.7 | 74.5 | 76.0 |
| SAM2-TinyViT-5M (~10% Data / 1 GPU) | No Distill | 5.4 | 26.0 | 30.4 | 30.8 | 28.4 | 27.7 |
| SAM2-TinyViT-5M (~10% Data / 1 GPU) | FitNet | 5.4 | 33.8 | 38.7 | 37.7 | 34.3 | 35.1 |
| MobileSAM2-5M (~10% Data / 1 GPU) | FitNet | 5.8 | 36.6 | 41.5 | 39.7 | 38.2 | 39.3 |
| MobileSAM2-5M (~10% Data / 1 GPU) | CIRKD | 5.8 | 36.8 | 43.3 | 41.2 | 38.3 | 38.6 |
| MobileSAM2-5M (~10% Data / 1 GPU) | FAKD | 5.8 | 37.7 | 42.5 | 41.0 | 39.6 | 39.9 |
| MobileSAM2-5M (Ours) | HyperKD | 5.8 | 44.6 | 51.0 | 48.8 | 49.2 | 49.4 |
| SAM2-TinyViT-11M (~10% Data / 1 GPU) | FitNet | 11.2 | 36.4 | 40.0 | 39.8 | 37.8 | 39.0 |
| MobileSAM2-10M (~10% Data / 1 GPU) | FAKD | 10.4 | 41.3 | 46.7 | 44.0 | 42.4 | 43.3 |
| MobileSAM2-10M (Ours) | HyperKD | 10.4 | 48.7 | 54.2 | 51.8 | 49.9 | 50.0 |
| SAM2-TinyViT-21M (~10% Data / 1 GPU) | FitNet | 21.2 | 44.9 | 49.0 | 47.0 | 45.8 | 45.6 |
| MobileSAM2-23M (~10% Data / 1 GPU) | FAKD | 23.7 | 60.6 | 65.0 | 63.4 | 62.0 | 62.7 |
| MobileSAM2-23M (Ours) | HyperKD | 23.7 | 65.8 | 72.1 | 69.0 | 66.4 | 67.8 |
Ablation Study¶
Ablation experiments on LVOS val isolate the contributions of individual HyperKD modules and evaluate hypergraph distance metrics against pairwise graph KD baselines.
| Model Variant | Patch-wise THKD | Instance-wise THKD | Granularity HyperKD | LVOS val (J&F) |
|---|---|---|---|---|
| MobileSAM2-23M (No Distill Baseline) | - | - | - | 44.8 |
| + Patch-wise THKD | โ | - | - | 62.9 |
| + Instance-wise THKD | - | โ | - | 64.4 |
| + Dual Temporal Hypergraphs | โ | โ | - | 64.4 |
| + Granularity HyperKD only | - | - | โ | 63.1 |
| MobileSAM2-23M (Full HyperKD) | โ | โ | โ | 69.0 |
Varying the hyperedge threshold \(\epsilon\) from 0.5 to 1.5 yields J&F scores of 65.2, 68.4, 69.0, 68.1, and 62.0, identifying \(\epsilon = 1.0\) as the optimal trade-off between connectivity and over-smoothing. Compared against pairwise graph distillation frameworks, Context Matters achieves 60.7 J&F and IntRA-KD achieves 60.9 J&F, both clearly trailing HyperKD's 69.0 J&F. On a mobile GPU (RTX 4060 8GB), SAM2-Large encounters OOM, while MobileSAM2-5M, 10M, and 23M sustain 13.3, 10.8, and 8.2 FPS respectively. HyperKD also transfers effectively to TrackAnything (TAM), improving DAVIS-2017 J&F from 60.8 to 66.5 on MobileTAM-23M.
Highlights & Insights¶
- Pioneering Hypergraphs in Video Foundation Model Distillation: Pairwise graphs cannot represent simultaneous multi-entity interactions across space and time. HyperKD reformulates inter-frame consistency and multi-level granularity as hyperedges, explicitly capturing high-order topological dependencies.
- Orthogonal Dual-Hypergraph Synergy: Temporal HyperKD ensures continuous spatial-temporal tracking, while Granularity HyperKD injects multi-scale conceptual fidelity. Their combination prevents semantic drift during complex occlusions and fast motion.
- Unified Architecture-Distillation Co-design: Rather than adapting generic backbones, MobileSAM2 relies on progressive contraction guided directly by HyperKD losses to derive lightweight Hiera architectures tailored for hierarchical visual attention.
- Practical Embodied AI Deployment: Integrated as the perception backbone in CaP-Agent0 on LIBERO-PRO, MobileSAM2-23M achieves a 21.8% average success rateโsubstantially outperforming SAM2-TinyViT-21M (19.3%) and closing in on full SAM3-based agents (23.2%) with minimal onboard compute overhead.
Limitations & Future Work¶
- Neighbor Search Overhead in Hyperedge Formation: Constructing hyperedges via pairwise distance matrix thresholding scales quadratically with node count, creating compute and memory bottlenecks on ultra-high-resolution inputs or prolonged video sequences.
- Static Spatial Prompt Sampling: The distillation protocol relies primarily on uniform 4ร4 spatial prompt grids from teacher predictions, leaving interactive multi-turn correction dynamics less thoroughly modeled during distillation.
- Future Directions: Exploring dynamic sparse hypergraphs to reduce topological computation and adapting HyperKD to open-vocabulary grounding and interactive multimodal reasoning agents.
Related Work & Insights¶
This work builds on SAM2's promptable video segmentation foundation and extends classical knowledge distillation (from FitNet's feature regression and relational KD's pairwise matrices) to high-order hypergraph structures inspired by Hypergraph Neural Networks (HGNN). A key takeaway for future research is that large foundation models encode complex relational invariants; distilling them through explicit high-order geometric representations enables lightweight models to approximate large-scale pre-trained behavior under modest compute and data regimes.
Rating¶
- Novelty: 4.5 / 5.0 (First work exploring hypergraph-based distillation for video foundation models)
- Experimental Thoroughness: 4.5 / 5.0 (Covers general VOS, long-term VOS, SA-V, embodied robotics, and cross-architecture evaluation)
- Writing Quality: 4.5 / 5.0 (Clear motivation, mathematically rigorous formulations, and thorough ablation analyses)
- Value: 4.5 / 5.0 (Crucial stepping stone for deploying interactive video segmentation on mobile and embodied edge devices)