Skip to content

Segmentation-Guided Homography Estimation for Long-Term Planar Tracking

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/serycjon/WOFTSAM
Area: Segmentation
Keywords: planar object tracking, homography estimation, segmentation-guided, SAM 2, Hough transform

TL;DR

This paper introduces SAM-H, a training-free geometric pipeline that extracts 8-DoF homographies from SAM 2 segmentation masks via Hough line intersections and DINOv2 symmetry disambiguation, and combines it with optical flow into WOFTSAM to achieve state-of-the-art performance on PlanarTrack and POT-210.

Background & Motivation

Planar object tracking seeks to estimate the full 8-degrees-of-freedom (8-DoF) homography matrix relating a planar target's coordinates across video frames. It serves as an indispensable geometric building block for augmented reality, film post-production visual effects, and robotic visual servoing. Conventional planar tracking methods primarily rely either on direct intensity-based template alignment or on sparse/dense correspondence matching (such as SIFT features or RAFT optical flow). However, in realistic unconstrained environments, tracked targets frequently encounter severe perspective distortion, wide-angle rotations, heavy motion blur, low-textured or reflective surfaces, transparent media, and dynamic appearance changes. Correspondence-based methods depend heavily on matchable surface texture; once tracking fails due to temporary full occlusion, rapid out-of-view motion, or severe blur, errors accumulate rapidly, leading to irrecoverable track loss.

In parallel, generic video object tracking has experienced a major paradigm shift driven by foundational segmentation models such as SAM 2. These models demonstrate exceptional open-world generalization and long-term tracking stability, maintaining high-fidelity masks across severe appearance variations and low-textured surfaces. Nevertheless, segmentation trackers output only region-level binary masks devoid of the rigid 8-DoF geometric structure required for planar homography estimation. Furthermore, directly estimating homographies from binary masks introduces severe geometric ambiguities: for quadrilateral planar targets (standard in benchmarks and industrial settings), any cyclic permutation of the four corner vertices yields the exact same bounding mask, making it impossible to uniquely determine a homography from the binary mask alone.

To bridge the gap between robust region segmentation and high-precision 8-DoF geometry, the paper investigates whether segmentation masks can serve as reliable re-detection anchors for long-term planar tracking. By decomposing mask boundary geometry via Hough line intersections and resolving cyclic ambiguities with motion priors and self-supervised feature matching, the geometric transformation can be recovered in a training-free manner. Core idea: develop SAM-H, a training-free segmentation-to-homography pipeline using Hough transform corner extraction and DINOv2 appearance symmetry disambiguation, and integrate it into WOFTSAM as a robust re-detection engine cascaded with dense optical flow to achieve both pixel-level precision and long-term tracking robustness.

Method

Overall Architecture

Given a video sequence and initial four control points forming a target quadrilateral, the goal is to estimate an accurate 8-DoF homography matrix \(H_t\) for each frame. The proposed system features two core contributions: first, the standalone SAM-H pipeline, which tracks the target mask using an adapted SAM 2 model, fits boundary lines using the Hough transform to identify corner intersections, and resolves 4-fold cyclic symmetry via motion continuity and DINOv2 feature matching; second, WOFTSAM, a three-stage cascade tracker that prioritizes high-precision pre-warped optical flow estimation, triggers SAM-H re-detection whenever flow inlier ratios drop below a failure threshold, and provides a direct fallback mechanism.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Current input frame and state"] --> B["Stage 1: Optical flow pre-warping tracking<br/>WFH dense optical flow"]
    B -->|Inlier ratio >= 20%| C["High-precision homography output"]
    B -->|Inlier ratio < 20% tracking loss detected| D["Stage 2: SAM-H mask homography re-detection<br/>SAM 2 mask + Hough intersections + DINOv2 disambiguation"]
    D --> E["Refined WFH pre-warping based on HSAM"]
    E -->|Refinement inlier ratio satisfied| C
    E -->|Refinement still fails| F["Stage 3: Fallback mechanism<br/>Direct HSAM homography fallback"]
    F --> C

Key Designs

1. Hough Transform Boundary Fitting and Corner Localization: Training-Free Quadrilateral Recovery from Masks Directly detecting geometric corners on discrete binary masks is vulnerable to boundary noise, jagged pixel edges, non-rigid mask expansions, and local contour concavities. To address this issue, SAM-H applies the classical Hough transform to the outer contour of the binary mask to detect the 4 most prominent boundary lines. The intersections of these four lines yield candidate corner coordinates. False intersections situated far from the mask boundary are automatically discarded, retaining only the four genuine boundary intersections. This approach requires no task-specific neural network training and exhibits inherent scale and perspective invariance, reliably reconstructing sharp quadrilateral edges even when the segmentation mask suffers from minor boundary leakage or contour noise.

2. Hybrid Symmetry Disambiguation: Combining Motion Priors with DINOv2 Self-Supervised Features Because a planar quadrilateral is geometrically invariant under cyclic permutations of its four vertices, the detected corner set cannot be mapped directly to the original control points. During continuous tracking at standard frame rates (\(\ge 25\) fps), inter-frame displacement is minimal; hence, the algorithm selects the cyclic permutation that minimizes Euclidean distance to the previous frame's corners. However, following tracking loss caused by occlusions or severe motion blur, historical motion priors become invalid. In these cases, the pipeline switches to global appearance-based disambiguation: the candidate patches warped under each cyclic permutation are extracted, and their self-supervised DINOv2 features are matched against the reference template via dot-product cosine similarity. Leveraging DINOv2's robustness to illumination, blur, and perspective distortion, the correct orientation is reliably identified. To prevent erratic switching during partial visibility or visual transitions, a temporal consistency filter requires the appearance-selected ordering to remain stable across \(\Theta_r = 5\) consecutive frames.

3. Partial Visibility Handling: Adaptive Similarity and Translation Transform Fallbacks In real-world sequences, foreground obstacles frequently cause partial occlusions, leaving only one or two boundary lines detectable and preventing closed-form 8-DoF homography computation. SAM-H handles this through an adaptive low-DoF pose composition: when all 4 corners are detected, the full 8-DoF homography \(H_t\) is computed directly; when 2 or 3 corners are visible, the current homography is composed from the previous frame's pose and a residual transformation \(\Delta H_{t-1 \to t}\), where \(\Delta H\) is constrained to a 4-DoF similarity transform (translation, rotation, uniform scale); when only a single corner is reliably tracked, \(\Delta H\) reduces to a 2-DoF spatial translation. This graded fallback prevents tracking divergence during extended partial occlusions.

4. WOFTSAM Three-Level Cascade: Fusing Correspondence Precision with Segmentation Robustness While SAM-H provides exceptional robustness against target loss, its boundary-derived homographies are limited by mask contour quantization, exhibiting lower precision at tight thresholds (e.g., 5 px) compared to dense correspondence matching. Conversely, correspondence-based trackers (such as WOFT) deliver high precision on textured surfaces but lack recovery mechanisms once tracking fails. WOFTSAM integrates both into a complementary cascade: the current frame is first pre-warped with the previous homography and refined via the Weighted Flow Homography (WFH) module. If the flow inlier ratio falls below \(\Theta_i = 20\%\), tracking failure is flagged, triggering re-detection: the frame is pre-warped using SAM-H's homography \(H_{\text{SAM}}\) and refined again with WFH. If the refined correspondence support remains insufficient, the tracker safely falls back to using \(H_{\text{SAM}}\) directly.

Loss & Training

SAM-H's geometric intersection extraction and symmetry disambiguation operate completely training-free, requiring zero task-specific parameter fine-tuning. The underlying vision models are deployed with off-the-shelf pre-trained weights: segmentation utilizes SAM 2.1 Hiera-Tiny incorporating DAM4SAM's enhancements (setting temporal memory stride to 5 and freezing memory updates when the target is absent to prevent feature contamination); appearance disambiguation uses pre-trained DINOv2 ViT-S/14 with registers producing 384-dimensional normalized embeddings; dense optical flow homography estimation reuses pre-trained RAFT weights from WOFT. On a single NVIDIA RTX A5000 GPU, WOFTSAM operates at approximately 2.4 FPS with 4.5 GB of GPU RAM usage.

Key Experimental Results

Main Results

The proposed trackers were comprehensively evaluated on the POT-210 benchmark and the challenging PlanarTrack test set (PTTST). Alignment error between ground-truth and estimated warped control points is evaluated under 5-pixel and 15-pixel precision thresholds (\(p@5\) and \(p@15\)). On PlanarTrack, the authors also introduced high-precision re-annotations of initial frame poses (PTTST re-annot).

Dataset Metric WOFTSAM (Ours) SAM-H (Ours) WOFT (Prev. SOTA) HVC-Net SIFT
POT-210 \(p@5\) (%) 91.9 64.4 90.0 91.4 65.8
POT-210 \(p@15\) (%) 97.5 89.0 95.4 95.8 69.6
PTTST (Original GT) \(p@5\) (%) 51.5 62.0 43.6 42.0 17.6
PTTST (Original GT) \(p@15\) (%) 77.2 80.0 64.8 63.1 26.7
PTTST (Re-annotated GT) \(p@5\) (%) 57.4 62.0 48.9 48.8 โ€“
PTTST (Re-annotated GT) \(p@15\) (%) 77.6 80.1 65.9 64.3 โ€“

Ablation Study

Ablations investigate segmentation backbone selection, the multi-frame temporal consistency parameter \(\Theta_r\) for symmetry disambiguation, and the failure-detection inlier threshold \(\Theta_i\) within WOFTSAM.

Ablation Dimension Config / Variant PTTST \(p@5\) (%) PTTST \(p@15\) (%) Note
Segmentation Module Plain SAM 2 59.6 77.2 Standard baseline video segmentation tracker
Segmentation Module Full DAM4SAM 58.6 77.7 Extra distractor modeling slightly degrades planar performance
Segmentation Module SAM2+ (Ours) 62.0 80.0 Retains stride 5 and freezes memory on target absence
Consistency Threshold \(\Theta_r = 0\) (Immediate switch) 58.9 76.6 Lacks temporal filtering; unstable under partial visibility
Consistency Threshold \(\Theta_r = 2\) 61.5 79.6 Smooths transitions; eliminates corner flip jitter
Consistency Threshold \(\Theta_r = 5\) (Default) 62.0 80.0 Optimal balance between stability and rapid re-detection
Consistency Threshold \(\Theta_r = 8\) 62.0 80.1 Performance saturates beyond 5 frames
Failure Inlier Threshold \(\Theta_i = 10\%\) 50.2 75.0 Tolerates too much drift before triggering re-detection
Failure Inlier Threshold \(\Theta_i = 20\%\) (Default) 51.5 77.2 Balances optical flow tracking and re-detection
Failure Inlier Threshold \(\Theta_i = 30\%\) 52.0 77.6 More aggressive re-detection slightly improves hard cases

Key Findings

  • On the unconstrained PlanarTrack benchmark, SAM-H outperforms the prior state-of-the-art WOFT by +18.4 percentage points on \(p@5\) and +15.2 pp on \(p@15\), demonstrating that long-term robust re-detection under severe appearance changes is the primary bottleneck in planar tracking.
  • On the high-precision POT-210 benchmark, WOFTSAM leverages correspondence precision together with SAM-H re-detection, nearly halving 15 px alignment errors compared to WOFT (reaching 97.5% \(p@15\)), with the largest gains observed under motion blur, occlusions, and unconstrained conditions.
  • Component selection breakdown within WOFTSAM indicates that normal optical flow tracking operates 85.3% of the time, re-detection is triggered 3.7% of the time, and direct fallback is utilized 11.0% of the time. The re-detection rescue rate reaches 69.8%, proving that the cascade effectively prevents permanent tracking failure.
  • A per-sequence oracle analysis reveals that approximately half of the sequences favor SAM-H while the other half favor WOFTSAM. An oracle selector reaches 86.9% \(p@15\) (+6.9 pp over SAM-H alone), confirming the profound complementarity between mask-based and correspondence-based homography estimation.
  • Initial frame ground truth analysis reveals that initial annotation errors (averaging only 1.9 px) are scaled up significantly as targets enlarge downstream (59% of frames feature larger targets than the initial view). Evaluating on re-annotated initial poses boosts WOFTSAM's \(p@5\) from 51.5% to 57.4%, closing more than half of the gap to SAM-H.

Highlights & Insights

  • Repurposing Segmentation Masks for 8-DoF Geometry: Rather than relying on black-box neural regressions from masks to poses, the method utilizes Hough transform line intersections on mask contours, turning foundational video segmentation models into zero-shot geometric trackers without requiring training data.
  • DINOv2 Semantic Descriptors for Cyclic Symmetry Disambiguation: Resolving the 4-fold ambiguity of quadrilateral corner sequences through self-supervised DINOv2 feature cosine similarity provides invariance to blur, reflections, and perspective distortions where low-level pixel matching fails.
  • Three-Stage Resilient Cascade: Combining "high-precision optical flow \(\to\) mask-guided re-detection with flow refinement \(\to\) direct geometric fallback" establishes an effective blueprint for fusing geometric precision with deep semantic tracking stability.

Limitations & Future Work

  • Quadrilateral Target Assumption: SAM-H relies on extracting four prominent straight boundary lines using the Hough transform; it cannot handle non-quadrilateral planar targets, circular/curved objects, or highly irregular planar regions.
  • Susceptibility to Segmentation Mask Spill: When the planar target is a sub-region of a larger 3D object that is only partially visible at initialization, SAM 2 can erroneously expand its segmentation to the entire 3D object as the view expands, distorting boundary line extraction.
  • Straight-Line Occluder Ambiguity: When an occluding object presents straight boundaries that maintain a pseudo-quadrilateral contour, SAM-H may incorrectly fit the homography to the visible partial region rather than maintaining the complete target extent.
  • vs WOFT (Serรฝch & Matas, WACV 2023): WOFT relies on RAFT dense optical flow pre-warping for weighted homography fitting, delivering high short-term precision but failing irrecoverably after tracking loss. WOFTSAM adopts WOFT's core flow tracker and introduces SAM-H as an automatic re-detection engine, overcoming the long-term tracking bottleneck.
  • vs HVC-Net (Zhang & Ling, CVPR 2022): HVC-Net regresses four control points end-to-end with deep networks, but its generalization is bound to synthetic training distributions. The proposed approach leverages foundational priors from SAM 2 and DINOv2, exhibiting superior zero-shot transfer across extreme real-world challenges.
  • vs PRTrack (Zhang et al., ICCV 2023): PRTrack estimates corner heatmaps and segmentation masks simultaneously for multi-object planar tracking. Evaluated on the MPOT-3K benchmark with strict 1-on-1 target evaluation, WOFTSAM achieves 95.1% \(p@50\), outperforming PRTrack's 92.85%.

Rating

  • Novelty: โญโญโญโญโ˜† (Ingenious, training-free formulation extracting 8-DoF homographies from segmentation masks via Hough transform and self-supervised symmetry disambiguation)
  • Experimental Thoroughness: โญโญโญโญโญ (Evaluated across POT-210, PlanarTrack, and MPOT-3K, accompanied by detailed error attribution and high-precision ground truth re-annotation)
  • Writing Quality: โญโญโญโญโญ (Exceptionally clear narrative, transparent discussion of failure modes and model complementarities, and reproducible system design)
  • Value: โญโญโญโญโญ (Establishes new state-of-the-art benchmarks on challenging planar tracking datasets and provides valuable code and refined annotations to the community)