Towards Effective Long Video Understanding: Dynamic MAS Construction via Meta-Agent¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Video Understanding
Keywords: Long Video Understanding, Dynamic Multi-Agent System, Meta-Agent, MAS-as-Code, Reinforcement Learning with Verifiable Rewards (RLVR)
TL;DR¶
To eliminate the computational redundancy and structural rigidity of static multi-agent architectures in long video understanding, DyMAC introduces a Vision-Language Meta-Agent that generates instance-specific multi-agent systems via executable code compiled into directed computational graphs, optimized through SFT and Reinforcement Learning with Verifiable Rewards (RLVR) for superior accuracy and efficiency.
Background & Motivation¶
Long video understanding is fundamentally constrained by immense spatio-temporal data volume and strict context length limits, which demand fine-grained retrieval alongside high-level temporal reasoning. Agent-based methods powered by Large Language Models (LLMs) and Vision-Language Models (VLMs) have emerged as the prevailing paradigm to mitigate information overload and memory constraints. By orchestrating external tools, keyframe sampling, and multi-step reasoning, these methods successfully decompose long-horizon video perception into tractable local subtasks. However, nearly all existing multi-agent systems (MAS) rely on manually crafted, static "one-size-fits-all" architectures that force rigid workflows across every query.
The core tension lies in the extreme instance-level diversity of video understanding tasks. To accommodate the most demanding spatio-temporal reasoning cases, static architectures are intentionally designed with extensive pipelines, complex hierarchies, or multi-round debate loops. When applied to simple clip retrieval or local factual QA, these rigid workflows enforce redundant global perception and excessive inter-agent communication, incurring massive computational waste and severe inference latency. Conversely, when confronted with complex causal reasoning, fixed linear or hierarchical topologies lack the adaptive flexibility to dynamically spawn auxiliary verification branches based on emergent intermediate clues.
The angle of attack in this paper is to shift from static human-engineered topologies to automated, instance-adaptive system synthesis. Core idea: formulate multi-agent orchestration as a code generation problem driven by a central Vision-Language Meta-Agent that dynamically synthesizes instance-specific executable code (MAS-as-Code) compiled into a directed computational graph, optimized via a two-stage paradigm (SFT + RLVR) with joint accuracy and execution efficiency rewards.
Method¶
Overall Architecture¶
DyMAC (Dynamic MAS Construction) frames multi-agent system configuration as a single-pass code generation and deterministic computational graph execution pipeline. Given a joint multimodal input \(X = (V, Q)\) comprising a long video \(V\) and a user query \(Q\), a central Meta-Agent (parameterized by Qwen2.5-VL-7B) assesses the reasoning requirements and decodes an executable standard code block (MAS-as-Code) in a single forward pass. This code is deterministically compiled into a heterogeneous directed computational graph (MAS-as-Graph) \(G = (\mathcal{V}, \mathcal{E})\), where each node \(v_i \in \mathcal{V}\) designates an agent role configured with specialized meta-attributes, and directed edges \(\mathcal{E}\) define conditional data flow and execution transitions. The workflow initiates at a designated start node, progresses through parallel visual perception, hierarchical retrieval, or multi-agent debate, and concludes at a terminal node that aggregates intermediate findings to output the final answer.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
IN["Multimodal Input<br/>Long Video V + User Query Q"] --> META["Meta-Agent Generates MAS-as-Code<br/>Single-Pass Decoding via Qwen2.5-VL-7B"]
META --> GRAPH["Deterministic Compilation to MAS-as-Graph<br/>Heterogeneous Directed Graph G=(V,E)"]
GRAPH --> EXEC["Dynamic Workflow Execution<br/>Node Allocation of 3B/7B Models & Reasoning Patterns"]
EXEC --> RLVR["Two-Stage Training & Reward Alignment<br/>SFT Syntax Warmup + RLVR Policy Refinement"]
RLVR --> OUT["Final Answer Synthesis<br/>Optimal Trade-off Between Accuracy & Latency"]
Key Designs¶
1. MAS-as-Code Generation and Deterministic Graph Compilation: Transforming Multi-Agent Orchestration into Single-Pass Code Synthesis
Prior dynamic multi-agent architectures predominantly rely on multi-turn prompt-based search or runtime self-reflection, which induce prohibitive API overhead and severe latency in long video scenarios. DyMAC overcomes this by formalizing dynamic topology orchestration as a single-pass standard code generation problem \(C \sim \pi_\theta(\cdot \mid X)\). The synthesized code explicitly specifies LangGraph topology configurations, node definitions, state dictionary transitions, and conditional routing edges. Upon deterministic compilation, the resulting heterogeneous directed graph \(G = (\mathcal{V}, \mathcal{E})\) provides extreme structural flexibility, dynamically instantiating parallel fan-out topologies for multi-segment summarization, cyclical loopback structures for iterative causal deduction, or fan-in verification pipelines for cross-temporal confirmation.
2. Quadruple Meta-Attributes Specification for Agent Roles: Granular Allocation of Compute and Cognitive Patterns
To prevent the resource mismatches inherent in static multi-agent systemsโwhere every agent uniformly runs large models or identical prompting routinesโDyMAC defines each agent node \(v_i \in \mathcal{V}\) via a quadruple meta-attribute tuple: $\(v_i = \langle \mathcal{I}_i, \mathcal{M}_i, \mathcal{T}_i, \mathcal{P}_i \rangle\)$ where \(\mathcal{I}_i\) is the customized system prompt establishing specific subtask goals and constraints; \(\mathcal{M}_i\) is the dynamically allocated foundation model size, routing subtasks between lightweight Qwen2.5-VL-3B and full-scale 7B models based on computational difficulty; \(\mathcal{T}_i\) denotes the discrete set of authorized external tools or APIs (e.g., FrameQA, ClipQA, temporal localization modules); and \(\mathcal{P}_i\) designates the cognitive reasoning pattern selected from six paradigms: Chain-of-Thought (CoT), ReAct, Plan-and-Execute, Self-Refine, LLM-Debate, and Self-Consistency. This attribute routing guarantees that straightforward extraction subtasks consume minimal tokens, while visually ambiguous verifications trigger multi-agent debate.
3. Two-Stage SFT and RLVR Optimization: Joint Alignment of Accuracy and Execution Efficiency
Untrained Meta-Agents frequently generate syntactic errors or deadlocking topologies, while supervised fine-tuning alone cannot teach the policy to balance accuracy against execution cost. DyMAC establishes a progressive two-stage training scheme. First, it constructs an SFT demonstration corpus \(\mathcal{D}_{SFT}\) containing 3,519 samples derived from CG-Bench (>20 minutes video length), where candidates are filtered for ground-truth correctness and ranked by structural AST complexity to distill syntactic robustness into the Meta-Agent. Second, it optimizes the policy via Reinforcement Learning with Verifiable Rewards (RLVR) using Group Relative Policy Optimization (GRPO). The joint reward function is formulated as: $\(\mathcal{R}(C, X) = \lambda_1 \mathcal{R}_{acc}(C, X) + \lambda_2 \mathcal{R}_{eff}(C)\)$ where \(\mathcal{R}_{acc} \in \{0, 1\}\) serves as an objective verifiable metric indicating whether the terminal prediction matches ground truth. The efficiency reward \(\mathcal{R}_{eff}\) inverts the intra-group min-max normalized execution penalty \(P_{eff}(C) = \max_{p \in \mathcal{P}} \sum_{v_i \in p} w(\mathcal{M}_i) \cdot t_i\), where \(t_i\) measures total token consumption on node \(v_i\) and \(w(\mathcal{M}_i)\) weights the parameter scale of driving model \(\mathcal{M}_i\). This reward design actively incentivizes the Meta-Agent to prune redundant nodes and unnecessary tool calls while preserving reasoning precision.
A Worked Example¶
Consider a complex video query: "What is the color of the shoes worn by the man during his run?": - Meta-Agent Planning: The Meta-Agent recognizes that the video features extensive pre-run footage followed by rapid running sequences with heavy motion blur. It dynamically generates a 5-node "Fan-out & Fan-in" architecture. - Parallel Analysis (Fan-out): A Temporal Locator (Node I, powered by a 3B model) segments timestamps into preparation and running intervals. It routes them concurrently to Visual Analyzer-Preparation (Node IIa, 3B model) and Visual Analyzer-Running (Node IIb, 3B model) to gather local visual evidence. - Cross-Interval Debate and Synthesis (Fan-in): A Verification Reviewer (Node III, 7B model using the LLM-Debate pattern) cross-references the crisp close-up evidence of white sneakers from Node IIa with the continuous motion-blurred silhouette from Node IIb, resolving the visual ambiguity. A Final Synthesizer (Node IV) consolidates the consensus to output Option C (White).
Key Experimental Results¶
Main Results¶
On Video-MME, MLVU, LVBench, and LongVideoBench (LVB), DyMAC is benchmarked against end-to-end MLLMs, multi-turn tool invocation methods, and representative agent frameworks:
| Method | Model Size | Video-MME | MLVU | LVBench | LVB | Avg. Acc (%) | Avg. Time (s/min video) |
|---|---|---|---|---|---|---|---|
| Vanilla | 7B | 65.1 | 68.8 | 45.3 | 56.0 | 58.8 | 1.5 |
| Video-o3 | 7B | 66.5 | 72.1 | 47.6 | 60.5 | 61.7 | 0.9 |
| ReAct | 7B | 68.0 | 71.2 | 47.5 | 57.5 | 61.1 | 8.3 |
| VideoRAG | - | 62.1 | 72.4 | - | 58.7 | - | 20.8 |
| Vgent | 7B | 68.9 | 72.1 | - | 59.7 | - | 24.1 |
| VideoLucy | 7B+671B | 72.5 | 76.1 | 58.8 | - | - | 19.9 |
| DyMAC (Ours) | 7B (Meta) | 70.6 | 73.2 | 55.6 | 60.7 | 65.0 | 4.6 |
Ablation Study¶
Ablation of training stages and input modalities on MLVU, Video-MME, and LVB:
| Config | MLVU | Video-MME | LVB | Avg. Acc (%) | Note |
|---|---|---|---|---|---|
| Vanilla | 68.8 | 65.1 | 56.0 | 63.3 | Direct single-model QA |
| MAS-Text (Untrained) | 55.3 | 50.8 | 46.6 | 50.9 | Query-only input, massive code compilation errors |
| MAS-Text + SFT | 67.7 | 66.0 | 55.6 | 63.1 | Syntax stabilized, compilation failures resolved |
| MAS-Text + SFT + RL | 69.5 | 67.8 | 56.0 | 64.4 | Policy refined under text-only queries |
| MAS-Visual (Untrained) | 54.3 | 51.2 | 48.0 | 51.2 | Multimodal input without fine-tuning |
| MAS-Visual + SFT | 70.9 | 68.5 | 57.2 | 65.5 | Video visual priors significantly enhance topology |
| MAS-Visual + SFT + RL (Full) | 73.2 | 70.6 | 60.7 | 68.2 | Full two-stage training, achieving peak trade-off |
Furthermore, reward function ablation demonstrates: - Accuracy reward only: 72.7% on MLVU / 70.0% on Video-MME with 13.2s inference time. - Joint Accuracy + Efficiency rewards: 73.2% on MLVU / 70.6% on Video-MME with 4.6s inference time, accelerating execution nearly 3x while marginally improving accuracy.
Key Findings¶
- Crucial Role of Visual Multimodal Priors: Synthesizing agent graphs purely from textual queries (MAS-Text + RL, 64.4%) lags significantly behind multimodal generation (MAS-Visual + RL, 68.2%), demonstrating that the Meta-Agent must inspect video content directly to accurately gauge task difficulty.
- Regularization Effect of Efficiency Penalties: Penalizing cumulative graph execution cost not only slashes inference latency from 13.2s to 4.6s per minute of video, but also acts as a structural regularizer, pruning noisy intermediate agents and lifting MLVU accuracy from 72.7% to 73.2%.
- Dynamic Topology Scaling with Video Duration: On Video-MME subsets, DyMAC generates an average of 3.4 nodes for Short videos, 4.7 nodes for Medium videos, and scales up to 6.0 nodes for Long videos, verifying its adaptive resource allocation across temporal horizons.
Highlights & Insights¶
- Single-Pass Code Synthesis for Dynamic MAS: Formulating multi-agent orchestration as MAS-as-Code avoids expensive runtime multi-round negotiation loops, leveraging deterministic compilers to create scalable, executable agent systems instantly.
- Verifiable Reward Alignment with GRPO: Combining objective answer verification with normalized path-level token-parameter execution penalties establishes an effective reinforcement learning blueprint for structural agent pruning.
- Heterogeneous Model and Reasoning Allocation: Dynamically routing between 3B and 7B models and among diverse reasoning paradigms (e.g., CoT vs. LLM-Debate) offers a principled methodology for balancing capacity and inference cost in complex multimodal systems.
Limitations & Future Work¶
- Compiler Robustness and Edge Fallbacks: While SFT and RLVR drastically reduce syntax errors, edge-case code generation failures or execution deadlocks still necessitate robust fallback and recovery mechanisms.
- Sparse Sampling Dependency for Meta-Planning: The Meta-Agent currently relies on 32 sparsely sampled frames for initial graph orchestration, which could potentially misjudge task difficulty if crucial subtle evidence resides in extremely narrow temporal windows.
- Future Directions: Exploring lightweight runtime dynamic graph replanning (allowing nodes to dynamically prune or append local subgraphs) and incorporating specialized sensory modules (such as 3D spatial tracking and audio analysis) into the meta-attribute pool.
Related Work & Insights¶
- vs VideoLucy / VideoExplorer: VideoLucy relies on an enormous 671B model with deep memory backtracking resulting in 19.9s latency, whereas VideoExplorer optimizes policy within a static agent graph. DyMAC shifts optimization to the topological and attribute level, achieving superior average accuracy at 4.6s latency with a 7B controller.
- vs Vgent / LongVideoAgent: Vgent and LongVideoAgent enforce predetermined graph traversal or master-slave structures, causing computational redundancy on straightforward queries. DyMAC's Meta-Agent compiles tailored architectures on demand, offering unmatched efficiency and flexibility across varying video lengths.
Rating¶
- Novelty: โญโญโญโญโญ Pioneering framework combining instance-specific MAS-as-Code generation with RLVR in long video understanding.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluation across four mainstream benchmarks with exhaustive ablations on modalities, rewards, model allocation, and reasoning paradigms.
- Writing Quality: โญโญโญโญโญ Clear mathematical formulations, clean structural diagrams, and consistent terminology throughout.
- Value: โญโญโญโญโญ Provides a benchmark paradigm for building adaptive, efficient, and cost-aware multi-agent systems in complex vision-language tasks.