Skip to content

InstrAct: Towards Action-Centric Understanding in Instructional Videos

Conference: ECCV 2026
Paper: ECCV Official Page
Code: https://zyyangzy.github.io/InstrAct/
Area: Self-Supervised Learning
Keywords: instructional video understanding, action-centric representations, static bias mitigation, dynamic time warping, masked action modeling

TL;DR

Addressing noisy web supervision and severe static shortcut bias where models rely on static objects rather than motion dynamics in instructional videos, InstrAct introduces a pretraining framework combining LLM-assisted action-centric hard negatives, an Action Perceiver with verb-guided distillation, Soft-DTW temporal alignment, and Masked Action Modeling (MAM), significantly enhancing fine-grained action discrimination and procedural temporal reasoning.

Background & Motivation

In recent years, video-language pretraining has expanded from trimmed clips containing single atomic actions to long, untrimmed instructional videos depicting complex multi-step procedures (such as HowTo100M and COIN). In realistic instructional activities, tasks are composed of sequentially ordered actions with causal and temporal dependencies; models must therefore not only recognize isolated objects and motions, but also model the temporal evolution of entire workflows. However, prevailing Video Foundation Models (VFMs) suffer from an acute "static bias" when handling instructional videos—models heavily exploit static visual cues such as background scenery and salient objects (nouns) as shortcuts, failing to comprehend true action semantics (verbs) and dynamic interactions.

This issue is particularly exacerbated in instructional video settings. Prior efforts mitigating static shortcuts predominantly targeted short, trimmed clips with clean captions or simple synthetic verb substitutions. In contrast, long untrimmed instructional videos rely on Automatic Speech Recognition (ASR) transcripts, which are inherently filled with conversational filler, non-instructional speech, and loose temporal alignment between audio narrations and visual actions. Such noisy and weak supervision makes it remarkably difficult to disentangle sequential actions from background context or to accurately ground temporal semantics. Directly transferring techniques designed for trimmed atomic actions to long-form instructional videos neither suppresses transcript noise nor establishes rigorous procedural reasoning.

The angle of attack in this paper is to unify data curation, motion feature disentanglement, and temporal sequence constraints into an action-centric pretraining paradigm. The authors develop an LLM-driven pipeline to filter conversational noise and extract structured verb phrases for synthesizing verb-altered and order-swapped hard negatives. On the visual side, an Action Perceiver distills dynamic tokens under verb-guided teacher supervision, while differentiable Soft-DTW temporal alignment and Masked Action Modeling jointly drive cross-modal grounding. Core idea: by combining LLM-curated verb-altered and order-swapped hard negatives with an Action Perceiver for motion token compression, Soft-DTW sequence alignment, and Masked Action Modeling (MAM), the framework systematically eliminates static object shortcuts and reinforces procedural temporal reasoning in instructional videos.

Method

Overall Architecture

InstrAct is designed to learn robust action-centric representations from weakly supervised instructional videos. The overall pipeline integrates structured action mining on the text side with a three-pronged objective on the multimodal representation side. Given an input video, a video backbone extracts dense spatio-temporal features, which are compressed into a compact set of latent Action Tokens via an Action Perceiver. Concurrently, an LLM filters non-instructional subtitle noise, standardizes narrations into discrete verb phrases, and generates verb-altered and order-swapped hard negatives. During pretraining, the framework optimizes global video-text contrastive alignment, applies Soft-DTW sequence alignment between action tokens and verb phrases, enforces Masked Action Modeling (MAM) with a causal multimodal text decoder, and distills verb semantics into the action token latent space.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Long Instructional Videos & Weakly Aligned ASR Subtitles"] --> B["LLM-Assisted Curation & Hard Negatives<br/>Filter conversational noise · Extract verb phrases · Construct HNs"]
    B --> C["Action Perceiver & Verb-Guided Distillation<br/>Temporal latent queries · Distillation from verb teacher"]
    C --> D["DTW-Align Temporal Sequence Alignment<br/>Soft-DTW alignment · Order-aware contrastive regularization"]
    C --> E["Masked Action Modeling MAM<br/>Autoregressive causal text decoder · Visual cross-attention"]
    D --> F["Action-Centric Video-Language Representation Output"]
    E --> F

Key Designs

1. LLM-Assisted Data Curation & Action Hard Negatives: Purifying Supervision and Penalizing Static Shortcuts

Raw ASR transcripts in instructional videos contain ubiquitous conversational fillers ("hi guys", "welcome back") that lack visual grounding. This design first leverages an LLM binary classifier to filter out non-instructional commentary, retaining only sentences describing concrete food preparation and physical manipulation. The LLM then standardizes the retained text into syntactic verb phrases following patterns like "V + N", "V-ing", or "V + Prep + N" (e.g., parsing "Crack the egg using the mug and separate the white" into ['crack egg', 'separate white']). Leveraging these structured units, two types of hard negatives (HNs) are constructed: verb-altered HNs replace the action verb while preserving the exact objects and scene context, penalizing reliance on static objects; order-swapped HNs permute the chronological order of actions while keeping identical vocabulary, breaking bag-of-words shortcuts and enforcing procedural order comprehension.

2. Action Perceiver with Verb-Guided Distillation: Distilling Motion Tokens from Redundant Video Features

Raw video representations are dominated by redundant background tokens that dilute transient action signals. To compress visual dynamics into an action-focused space, the authors introduce the Action Perceiver. Given visual embeddings \(V \in \mathbb{R}^{T \times N \times C}\), a bank of \(K\) learnable latent queries \(L \in \mathbb{R}^{K \times C}\) constrained by localized temporal window attention masks interacts with \(V\) via cross-attention to produce student Action Tokens \(S \in \mathbb{R}^{K \times C}\): $\(S = \text{Perceiver}(q = L, k = V, v = V)\)$ Because explicit verb labels are unavailable during downstream inference, a verb-guided teacher branch provides distillation supervision during pretraining: extracted verb phrase embeddings \(P \in \mathbb{R}^{M \times C}\) query visual features \(V\) to generate teacher tokens \(T\). A bidirectional soft contrastive distillation loss aligns student similarity predictions against teacher targets across both video-to-text (v2t) and text-to-video (t2v) directions: $\(\mathcal{L}_{\text{distill}} = \mathcal{L}_{\text{single}}(s_{v2t}, s'_{v2t}) + \mathcal{L}_{\text{single}}(s_{t2v}, s'_{t2v})\)$ This distillation injects motion semantics directly into the latent query space, enabling the model to extract action-sensitive representations at test time without requiring verb prompts.

3. Soft-DTW Temporal Alignment with Order-Aware Regularization: Flexible Alignment of Procedural Sequences

Individual actions in instructional videos vary widely in duration, and narration timestamps often drift from visual execution, making rigid monotonic frame-to-word matching brittle. The framework computes cosine distances between action tokens \(S = \{s_1, \dots, s_K\}\) and verb embeddings \(V = \{v_1, \dots, v_M\}\) to build a cost matrix \(C \in \mathbb{R}^{K \times M}\), optimizing a differentiable Soft-DTW objective to find the optimal monotonic alignment path: $\(\mathcal{L}_{\text{align}}(S, V) = \text{Soft-DTW}(C)\)$ To prevent degenerate trivial alignments, an order-aware contrastive regularization term is incorporated: a negative action sequence \(\tilde{S} = [s_K, \dots, s_1]\) is formed by reversing the token order, and a hinge loss enforces a margin \(\beta\) ensuring that the forward sequence alignment cost is strictly lower than the reversed sequence cost: $\(\mathcal{L}_{\text{order}} = \max\left(0, \mathcal{L}_{\text{align}}(S, V) - \mathcal{L}_{\text{align}}(\tilde{S}, V) + \beta\right)\)$

4. Masked Action Modeling: Deep Cross-Modal Reconstruction of Motion Dynamics

While contrastive objectives and DTW capture global and sequence-level correspondences, dual-encoder architectures offer limited token-level fine-grained multimodal interaction. InstrAct incorporates a multimodal causal text decoder that injects visual action tokens \(S\) via cross-attention following self-attention at each decoder layer. By randomly masking action verbs in the narration text, the decoder is trained via an autoregressive cross-entropy loss to reconstruct the masked tokens conditioned on visual action evidence: $\(\mathcal{L}_{\text{mam}} = -\sum_{t=1}^{N-1} \log P(y_{t+1} \mid h_{1:t}, S)\)$ This objective forces the multimodal layers to attend directly to visual motion features to resolve missing action verbs, eliminating language prior shortcuts where models guess actions purely from linguistic co-occurrence.

Key Experimental Results

Main Results

The evaluation is conducted across three diagnostic benchmarks: InstrAct-Semantic (a 10-choice MCQ isolating fine-grained verb identification by controlling identical object contexts), InstrAct-Logic (a 2–3 negative choice MCQ evaluating procedural temporal ordering), and InstrAct-Dynamics (an authentic cross-modal retrieval benchmark of 16,326 clips organized into 120 object-centric pools to suppress object-matching shortcuts).

Benchmark & Metric Metric Ours (InstrAct) Best Baseline (VideoPrism) Runner-up (PerceptionLM) Classic Baseline (MIL-NCE)
InstrAct-Semantic R@1 (%) 28.1 18.2 16.5 7.2
R@5 (%) 78.0 74.4 63.2 49.4
Mean Rank (↓) 3.0 4.0 4.0 6.0
InstrAct-Logic ACC (%) 40.0 31.3 31.4 5.1
InstrAct-Dynamics T2V R@1 / R@5 (%) 33.86 / 54.73 27.23 / 48.49 16.33 / 32.11 12.63 / 30.77
V2T R@1 / R@5 (%) 35.19 / 59.52 34.13 / 59.08 23.10 / 42.05 15.04 / 34.42
Mean R@1 / R@5 (%) 34.53 / 57.13 30.68 / 53.79 19.72 / 37.08 13.84 / 32.59

Ablation Study

The impact of hard negatives, the Action Perceiver module, and the pretraining objectives on the InstrAct Bench is summarized below:

Configuration / Variant Verb HNs Order HNs Action Perceiver MAM Loss DTW Align Semantic R@1 (%) Semantic R@5 (%) Logic ACC (%) Note
Full model (InstrAct) ✓ ✓ ✓ ✓ ✓ 45.1 90.8 44.7 Best overall performance across tasks
w/o DTW Align ✓ ✓ ✓ ✓ – 43.7 90.0 39.4 Logic ACC drops by 5.3% without DTW
w/o MAM Loss ✓ ✓ ✓ – ✓ 38.7 86.5 40.2 Semantic R@1 drops by 6.4% without MAM
w/o Action Perceiver (MAM) ✓ ✓ – ✓ – 41.0 89.4 37.6 Dense raw frames degrade cross-modal attention
w/o Action Perceiver (DTW) ✓ ✓ – – ✓ 37.3 87.2 37.4 Raw frame DTW easily collapses onto static objects
Verb-altered HNs only ✓ – ✓ ✓ ✓ 46.7 91.2 40.0 High verb specialization but limited temporal logic
Order-swapped HNs only – ✓ ✓ ✓ ✓ 28.1 78.0 44.2 Strong sequence sensitivity but lower verb discrimination

In the temporal alignment analysis (Normalized DTW Cost), raw frame embeddings yield an alignment cost of 0.97, whereas Action Perceiver latents drop the cost to 0.91. Qualitative heatmaps reveal that raw embeddings display diffuse vertical stripes over static objects (e.g., persistently attending to "jam" across all frames), whereas the Action Perceiver produces a clean, diagonal staircase pattern isolating sequential manipulation steps (e.g., "put on").

Key Findings

  • Complementary Objectives: MAM predominantly drives local fine-grained verb discrimination, yielding a 6.4% boost in Semantic R@1, whereas Soft-DTW sequence alignment specializes in macro temporal sequencing, providing a 5.3% gain in Logic ACC. Joint optimization achieves the superior trade-off.
  • Action Perceiver Disentangles Motion from Object Bias: Without the Action Perceiver, uncompressed visual tokens suffer from static visual shortcuts, leading to degenerate alignment paths. Localized window attention combined with teacher distillation allows the model to extract clean motion dynamics without requiring verb annotations at test time.
  • Broad Backbone Generalization: When integrated across four distinct backbones (MIL-NCE, CLIP-ViP, CLIP4Clip, and ViCLIP), InstrAct provides consistent improvements, boosting ViCLIP from 14.7% to 45.1% in Semantic R@1 and from 30.0% to 44.7% in Logic ACC.

Highlights & Insights

  • Object-Controlled Benchmarking against Shortcuts: By evaluating on object-controlled pools (InstrAct-Dynamics) and controlled MCQs, the paper effectively prevents models from using object names as retrieval shortcuts, providing a rigorous standard for action bias evaluation.
  • Inference-Free Teacher Distillation: Guiding visual latent queries with verb embeddings during training and removing the text branch during inference elegantly resolves the practical constraint that downstream tasks do not provide ground-truth action verbs.
  • Transferability to Procedural Embodied AI: The combination of Soft-DTW and order-swapped contrastive loss can be naturally transferred to robot learning from demonstration and first-person manipulation planning where temporal sequence fidelity is paramount.

Limitations & Future Work

  • Domain Scope: The empirical pretraining and verification are currently centered on the cooking subset of HowTo100M. Extending this paradigm to domains with longer temporal horizons and diverse manipulation tools (e.g., carpentry, industrial manufacturing, or multi-modal craftwork) remains an important frontier.
  • Reliance on External LLM Parsing: The verb phrase extraction currently relies on offline LLM prompting. Future work could investigate an end-to-end framework where visual action grounding iteratively bootstraps text parsing without requiring separate LLM inference.
  • vs MIL-NCE [21]: MIL-NCE relies on multiple-instance contrastive learning to handle uncurated instructional video pairs but lacks action-specific hard negatives. Consequently, it leans heavily on static noun shortcuts, achieving only 7.2% R@1 on InstrAct-Semantic; InstrAct bridges this gap via explicit action-centric curation and temporal alignment.
  • vs VideoPrism [41] & PerceptionLM [3]: While modern VFMs excel in general representation learning, their procedural reasoning on order-sensitive actions remains capped (~31% on InstrAct-Logic); InstrAct's order-aware DTW pretraining explicitly enforces procedural progression, outperforming them by a substantial margin (44.7%).

Rating

  • Novelty: ⭐⭐⭐⭐ [Well-motivated combination of verb distillation, Perceiver compression, and Soft-DTW sequence alignment for long instructional videos]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Introduces 3 diagnostic benchmarks, evaluates 4 foundation backbones, and provides extensive quantitative and qualitative ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear problem formulation, transparent methodology, and cohesive narrative structure]
  • Value: ⭐⭐⭐⭐⭐ [Provides valuable design insights and benchmarks for temporal action understanding, mitigating static shortcuts, and procedural video learning]