Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs¶
Conference: NeurIPS2026
arXiv: 2609.31193
Area: Multimodal VLM / Interpretability
Keywords: trimodal binding, symbolic IDs, causal mediation analysis, active speaker detection, audio-visual prompting
TL;DR¶
Through representational analysis and activation interventions on controlled successful trajectories, this paper explains how audio-visual LLMs bind “who says what” through temporal and position IDs, then uses an existing active speaker detector for visual prompting and optional lightweight fine-tuning to improve conversation understanding, with training-free gains depending on visual-marker grounding capabilities.
Background & Motivation¶
Audio-visual large language models can recognize visual content and process speech without reliably associating an utterance with the correct visible speaker. In single-shot videos containing several characters simultaneously, answering “Which word did the tiger say?” requires more than recognizing the tiger and transcribing the words: the model must connect a textual reference, a visual character, and an acoustic segment. Understanding the three modalities separately does not guarantee correct correspondence among them.
Prior binding research in language models and vision-language models suggests that models may rely on relatively content-independent structural variables rather than directly matching full semantic representations, such as using left/right positions to track visual objects. With audio, spatial position is no longer the only coordinate; speaking order may also index acoustic events. This paper therefore first asks whether models use modality-specific structural IDs as intermediates, and then asks which conversion fails without contextual demonstrations, rather than attributing every error to inadequate multimodal understanding.
The practical intervention uses external Active Speaker Detection (ASD): if a model can locate the sound or character referenced in text but cannot connect the two, the temporal correspondence can be drawn directly on the video. Core Idea: explain trimodal binding through “anchor ID retrieval—target ID selection—feature retrieval,” use controlled interventions to locate the cross-modal ID-conversion bottleneck, and supply the missing correspondence through bounding boxes marking the current speaker.
Method¶
Overall Architecture¶
The paper contains a mechanistic investigation and an application improvement; these should not be treated as one deployment network. The investigation defines bidirectional binding tasks on synthetic four-animal videos, examines internal representations during successful answers, and uses activation patching to test the causal roles of different information types. The application takes ordinary conversational video, uses an existing ASD model to locate the current speaker frame by frame, overlays red boxes, and explains their meaning before the AVLLM answers questions or generates dialogue descriptions.
The proposed three-stage account maps the textual anchor to its modality-specific structural ID, selects the target ID in the other modality, and retrieves the semantics indexed by that target ID. These IDs describe structural information inferred from hidden states; they are not explicit integer registers emitted by the model or newly introduced discrete encoding modules.
Optional ASD-FT applies supervised fine-tuning to boxed synthetic videos to strengthen marker interpretation and audio-visual correspondence. Boxes in real evaluation come from the detector, whereas synthetic training boxes are determined by generation. Activation caches obtained from correctly primed trajectories are diagnostic tools, not part of actual deployment.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Controlled videos and textual anchors"] --> B["Symbolic Binding Tasks"]
B --> C["Causal Bottleneck Diagnosis"]
C -.->|Motivates improvement; no correct activations transferred| E["Active Speaker Prompting"]
D["Conversation video and question"] --> E
E -->|Inference data flow: boxed video and explanation| F["AVLLM answer or description"]
E -.->|Optional training: synthetic videos and gold answers| G["Prompt-Aware Fine-Tuning"]
G -.->|Updates LoRA parameters| F
Key Designs¶
1. Symbolic Binding Tasks: separate semantics, space, and time to identify the index actually used
Each synthetic video contains four animals placed randomly in four quadrants, each uttering a different country name in random order. Animal species, quadrant, spoken word, and speaking order define a binding event. Position IDs 0–3 correspond to top-left, top-right, bottom-left, and bottom-right; temporal IDs 0–3 index the four utterances. Changing animals or words while controlling structure helps distinguish semantic-content tracking from structural indexing.
Acoustically-Anchored Visual Retrieval (AAVR) starts from a word in the text prompt and asks which animal said it. Visually-Anchored Audio Retrieval (VAAR) starts from an animal description and asks which word it said. The hypothesized AAVR path is “word→speaking order→screen position→animal”; VAAR reverses the relationship through “animal→screen position→speaking order→word.” The contribution is not another classification task, but an independently testable correspondence between two structural coordinate systems.
Because accuracy is low in the natural, unprimed setting, the main mechanistic analysis includes only correct predictions and supplies contextual demonstrations of the other three animal–word bindings. These do not directly provide the queried animal’s answer, but they give strong task structure and binding cues. Evidence from successful trajectories therefore should not be equated with a stable, correct full computation in every unprompted video.
2. Causal Bottleneck Diagnosis: measure representational structure, then alter intermediate information to test its effect
Representational Similarity Analysis (RSA) compares pairwise similarity patterns of hidden states with three hypothesis spaces—temporal IDs, position IDs, and target semantics—and computes their Pearson correlation. The authors extract per-layer attention outputs at the anchor token and the last prompt token across 800 samples. Mid-to-late anchor-token layers align more closely with the anchor-modality ID; late last-token layers align with the target-modality ID; the deepest layers align with answer semantics. This ordering supports the three-stage account, but correlation alone does not establish that the information determines the answer.
Causal Mediation Analysis (CMA) constructs original, modified, and patched forward passes. The modified video swaps utterance order, swaps animal positions, or changes target semantics, and its attention outputs are patched into the original pass. The first two manipulations preserve the actual association between the queried animal and its word, so the two complete passes can have the same correct answer; mixing their intermediate states instead predicts a shift toward the exchanged event. This tests whether structural IDs mediate the computation rather than merely reflecting a change in video content.
The CM Score in the original Equation (1) measures how patching changes the counterfactual answer’s logit advantage over the original answer:
Here \(c_1\) is the original context, \(c_1^*\) is the patched pass, \(y_1\) is the original answer, and \(y_1^*\) is the expected counterfactual answer; \(M(c)[y]\) is the logit of the corresponding token. A larger score indicates that the patched information more strongly steers the output toward the expected counterfactual. CMA aggregates 960 samples and patches windows spanning one-quarter of the total model depth. This is relatively coarse stage-level evidence, not a precise localization of a single-layer discrete-ID circuit.
The authors next compare primed and unprimed RSA, finding a small difference for anchor IDs and a larger difference for target IDs. To test the bottleneck, they select the 10 highest-scoring attention head–layer pairs per stage on a 100-sample subset, average their outputs over 800 primed instances into correct activation caches, and patch these into unprimed runs. Target-ID-selection interventions substantially recover accuracy. This diagnostic intervention depends on correct internal states and must not be marketed as an inference method requiring no correct information. Stages can also span both token positions and need not occupy strictly disjoint layer ranges across models.
3. Active Speaker Prompting: visualize cross-modal correspondence rather than merely increase salience
The application uses the existing LASER ASD model to draw a red bounding box around the detected current speaker in each frame. Because the box changes with utterance events, it supplies “the sound at this moment corresponds to this location,” not merely a static object box. The AVLLM still receives the original audio and video, but the video now carries a visible synchrony cue that reduces ambiguity when converting temporal IDs to position IDs or vice versa.
A textual explanation is also necessary. For question answering, an instruction defining the red box as the current speaker is prepended to the original question; for captioning, the explanation is appended. The model must interpret the box and connect it to the utterance rather than guess from a red region. Marking an inactive character, a random region, or omitting the explanation does not replace correct synchronized marking.
Training-free ASD does not update AVLLM parameters, but it still runs an external detector and therefore is not computation-free. Its effectiveness depends on existing localization and marker-grounding capabilities: Qwen2.5-Omni and MiniCPM-o-4.5 benefit, whereas video-SALMONN2+ declines on several training-free metrics. On the three general benchmarks, which often lack speech or visible speakers, the authors do not apply ASD boxes. The ASD rows are consequently identical to vanilla and cannot be read as evidence that prompting improves general tasks.
4. Prompt-Aware Fine-Tuning: teach marker interpretation through a small controlled supervision set
ASD-FT uses 400 synthetic human-character videos rather than animals as training subjects: 200 two-person single-shot videos, 100 four-person single-shot videos, and 100 two-person multi-shot videos. Characters, country words, and TTS voices are combined independently, and exact active-speaker boxes are generated during synthesis. This separates box errors from the model-learning problem in training, but also creates a distribution shift between accurate training boxes and real detector outputs.
The dataset contains 1,000 AAVR multiple-choice samples, 1,000 VAAR multiple-choice samples, and 40 open-ended captioning samples, totaling 2,040. The retrieval tasks require appearance–word correspondence, while captioning summarizes each character’s appearance and utterance. Supervision therefore targets not only the visual appearance of a red box but also retrieval of acoustic information from the correct visual region. After fine-tuning, the authors observe stronger target-ID correlations in the unprimed setting. This remains RSA evidence and does not establish a comprehensive causal attribution of the fine-tuning gains.
A Worked Example¶
Consider the main-text example: the panda at top-left says Mexico and speaks second; the tiger at top-right says Japan and speaks fourth; the monkey at bottom-left says Canada and speaks third; the rabbit at bottom-right says Egypt and speaks first. Spoken order is described here from 1, while IDs are indexed from 0.
For “Who said Japan?”, the AAVR account first maps Japan to temporal ID 3, then maps temporal ID 3 to position ID 1, and finally retrieves the tiger at top-right. For “What did the tiger say?”, VAAR retrieves position ID 1, selects temporal ID 3, and retrieves Japan. If the second stage incorrectly chooses bottom-left, the model can recognize every animal and transcribe Japan correctly while still attributing the word to the monkey.
In the mechanism experiment, swapping the tiger and monkey positions leaves their words unchanged, so the correct answer for the complete modified video remains the tiger. Patching only the modified target-position state into the original video can redirect retrieval to the monkey still occupying bottom-left in the original. This tests the causal role of an intermediate index; it does not claim that the tiger actually started saying Canada after the swap.
Prompted inference requires no such activation patching: when Japan is spoken, the red box highlights the top-right tiger and the text explains that the box marks the current speaker. Optional ASD-FT learns to use this cue. This synthetic animal example illustrates the process and introduces no additional evaluation result.
Loss & Training¶
Training uses LoRA supervised fine-tuning (SFT). The paper introduces no dedicated symbolic-ID loss or additional alignment loss, so no new loss formula is supplied here. LoRA rank is 16, batch size is 1, and gradient accumulation is 8 steps on a single NVIDIA RTX A6000.
video-SALMONN2+ receives 140 optimization steps and the other two models receive 250, all below 300 steps. Evaluation uses greedy decoding. Training-free ASD and ASD-FT are distinct settings; the latter is not training-free.
Key Experimental Results¶
Main Results¶
The table selects four conversation-centric benchmarks and one general benchmark from the original Table 1. AVSpeaker, DailyOmni, SocialOmni, and OmniBench report accuracy (%); SocialOmni uses the perceptual-task subset. DiaDemBench REF measures speaker-attribution consistency, while ASR measures transcription accuracy—not attack success rate, and not directly word error rate.
DiaDemBench is scored with gemini-2.5-Pro and gemini-2.5-Flash. Appendix B.2.2 explicitly states that ± reflects variation across repeated Gemini-judge evaluations of the same model outputs. The paper calls this variance but does not precisely define it as standard deviation or standard error; it must not be interpreted as training-seed uncertainty or a confidence interval.
| Model and setting | AVSpeaker | DailyOmni | SocialOmni | DiaDem REF | DiaDem ASR | OmniBench |
|---|---|---|---|---|---|---|
| Qwen2.5-Omni 7B | 44.86 | 64.33 | 38.95 | 19.1±0.4 | 28.1±0.1 | 50.79 |
| + ASD | 45.59 | 64.66 | 41.90 | 22.5±0.2 | 33.0±0.2 | 50.79 |
| + ASD-FT | 48.79 | 71.09 | 44.25 | 20.7±0.2 | 32.3±0.1 | 52.80 |
| MiniCPM-o-4.5 9B | 49.41 | 61.99 | 46.05 | 6.3±0.5 | 7.4±0.7 | 47.99 |
| + ASD | 50.34 | 62.49 | 50.00 | 9.7±0.4 | 10.8±0.1 | 47.99 |
| + ASD-FT | 51.77 | 63.41 | 49.70 | 24.8±0.2 | 30.9±0.0 | 51.84 |
| video-SALMONN2+ 7B | 44.74 | 65.13 | 43.49 | 14.0±0.3 | 19.3±0.1 | 37.04 |
| + ASD | 44.65 | 62.57 | 43.89 | 11.8±0.3 | 19.0±0.0 | 37.04 |
| + ASD-FT | 49.10 | 69.13 | 49.62 | 19.7±0.2 | 26.7±0.2 | 44.92 |
General tasks receive no ASD boxes. ASD-FT DAVE / WorldSense scores are: Qwen 33.91 / 53.73 versus vanilla 30.75 / 50.06; MiniCPM 57.97 / 57.37 versus 55.80 / 55.73; SALMONN 39.65 / 46.38 versus 35.28 / 45.04. These support general gains from fine-tuning over the original model, not a contribution from actual box prompting on these tasks.
Ablation Study¶
The following results come from the original Tables 2–4 and all use Qwen2.5-Omni 7B. AVSpeaker and DAVE report accuracy (%). The ordinary FT comparison on DAVE is retained to prevent interpreting ASD-FT as superior to ordinary fine-tuning on every task.
| Config or comparison | AVSpeaker | DAVE | Note |
|---|---|---|---|
| Vanilla | 44.86 | 30.75 | No additional prompting or fine-tuning |
| ASD: active character, red, medium, with explanation | 45.59 | 30.75 | No boxes on general tasks |
| Inactive character marked | 44.77 | Not applicable | Explanation, red, medium |
| Random region marked | 44.52 | Not applicable | Explanation, red, medium |
| Active character without textual explanation | 44.80 | Not applicable | Red, medium |
| Active character, green, medium, with explanation | 45.95 | Not applicable | Gains survive a color change |
| Active character, red, thick, with explanation | 45.70 | Not applicable | Thickness variant |
| Active character, red, thin, with explanation | 45.58 | Not applicable | Thickness variant |
| Ordinary FT, unboxed training videos | 46.89 | 35.15 | Training control without ASD prompts |
| ASD-FT | 48.79 | 33.91 | Higher AVSpeaker, but lower DAVE than ordinary FT |
Additional mechanistic evidence comes from Appendix Table 5 and A.1. Accuracy statistics use 800 samples per model; intervention values are reported in the appendix prose, and no missing MiniCPM value is inferred.
| Model | Unprimed | Primed | Primed text only, no audio/video | Unprimed after target-ID activation intervention |
|---|---|---|---|---|
| video-SALMONN2+ 7B | 27.0% | 99.2% | 1.00% | 51.0% |
| Qwen2.5-Omni 7B | 29.6% | 100.0% | 12.50% | 49.0% |
| Qwen2.5-Omni 3B | 28.7% | 99.94% | 12.50% | 45.4% |
| MiniCPM-o-4.5 9B | 0.0% | 99.31% | 0.06% | Not reported in this passage |
Key Findings¶
- Correct marking matters more than color alone: inactive, random, and unexplained boxes all fall below vanilla 44.86, whereas correctly placed green boxes reach 45.95. Gains are not simply a visual-salience effect.
- ASD-FT improvements over vanilla do not mean superiority over every setting: MiniCPM SocialOmni is 49.70, below training-free ASD at 50.00; Qwen DiaDem REF / ASR are 20.7 / 32.3, below ASD at 22.5 / 33.0.
- In the original Table 2, ASD at 45.59 on AVSpeaker exceeds AVCD 42.82, FMD 45.27, OutRo 45.24, NumPro 45.08, and VISER 44.55. This ordering applies only to that model and dataset.
- The extension to approximately 2,000 SocialOmni two-speaker clips is RSA of position and temporal IDs, without a full semantic-space comparison or equally sized real-data CMA. The appendix also checks RSA trends with five/six speakers and synthetic noise at 20 / 10 dB.
Highlights & Insights¶
- Separating object recognition and transcription from correspondence makes the failure diagnosis more specific. Accuracy recovery after target-ID intervention provides evidence beyond correlation curves.
- The external prompt carries a time-varying speaker location rather than a static detection result. Cross-modal correspondence may therefore require synchronized structural cues instead of additional appearance descriptions.
- Separating exact synthetic boxes from real detector boxes helps study marker grounding. Training success and detector reliability still require separate evaluation rather than being conflated into one end-to-end capability.
Limitations & Future Work¶
- The main mechanistic evidence comes from primed, correctly answered synthetic trajectories, leaving overlapping speech, moving characters, complex utterances, and occlusion underexplored. Real-world RSA supports the trend but does not replace causal validation in real scenes.
- IDs are abstract variables inferred through representations and interventions; quarter-depth patching windows do not establish a precise discrete circuit. Finer-grained, cross-context interventions could test the transferability of position and temporal variables.
- The method depends on the accuracy and fairness of existing ASD, as well as AVLLM marker interpretation. For authorized ordinary dialogue description, detector errors, occlusion, and performance across demographic groups should be documented alongside aggregate question-answering accuracy.
- The original Table 1’s all-benchmark gain claim should be read as ASD-FT versus vanilla, not versus ASD or ordinary FT; Table 4 explicitly reports superiority over FT on only 4/5 datasets. The main-text training-free comparison incorrectly cites Table 3, while its numerical results appear in Table 2.
- Mechanistic analysis excludes multi-shot scenarios, but training does include 100 two-person multi-shot videos; these are distinct scopes. The main text describes patching attention outputs, whereas Appendix C.2 refers to the residual stream. The cache is insufficient to resolve the exact hook location, so no implementation detail is inferred.
Related Work & Insights¶
- vs symbolic binding research in LLMs / VLMs: earlier studies examine entity–attribute binding or visual positions; this paper adds conversion between acoustic temporal order and visual position. It broadens the modalities of the binding hypothesis rather than proving one uniquely shared mechanism in all AVLLMs.
- vs NumPro / VISER visual prompting: generic numbering or visual structure assists localization, whereas this method’s boxes are determined by current speech events and directly supply audio-visual synchrony. The AVSpeaker comparison supports this choice but does not establish superiority on every visual reasoning task.
- vs AVCD / FMD / OutRo decoding improvements: these approaches modify inference-time output or modality utilization; ASD changes the input video and explains the marker. Complementarity should be tested rather than judging general superiority from one benchmark ranking.
- vs LASER active speaker detection: LASER is an existing component used by the paper, not a newly proposed detector. The contributions are mechanistic diagnosis, prompting, and lightweight training validation, not a redesigned object detector.
Rating¶
- Novelty: 4/5. Explains trimodal binding through modality-specific structural IDs and connects diagnosis to practical prompting.
- Experimental Thoroughness: 4/5. Includes multi-model interventions, seven application benchmarks, and prompt ablations, but limited causal coverage of complex real scenes.
- Writing Quality: 4/5. The three-stage account is clear; some table references and aggregate gain claims require checking the numerical results.
- Value: 4/5. Provides an actionable direction for improving audio-visual correspondence while retaining dependencies on external detection and marker grounding.