Binding Multiple Modalities via Multimodal Wasserstein Barycenter¶
Conference: NeurIPS 2026 Oral (according to the arXiv comment; the official acceptance page was not independently verified)
arXiv: 2609.33800
Area: Multimodal VLM
Keywords: multimodal alignment, Wasserstein barycenter, optimal transport, volumetric contrastive learning, missing modalities
TL;DR¶
BaryBind replaces a fixed modality anchor with a learnable Wasserstein barycenter map and applies volumetric contrastive learning to modality gap vectors around that anchor, improving zero-shot retrieval and classification on a VAST backbone, although its published dual derivation, potential cancellation, and some experimental numbers require clarification.
Background & Motivation¶
CLIP-style contrastive learning works well for two modalities, but extending it to text, video, audio, subtitles, and depth involves more than adding encoders. ImageBind uses vision as the binding hub, while LanguageBind uses language; other modalities primarily learn to approach an existing modality representation. Different sensors do not capture identical semantics: text can omit sounds, and audio need not identify visual entities. Making one modality the sole reference can favor information that it represents easily while compressing information from weaker modalities. VAST is the direct baseline here. The authors interpret asymmetric bidirectional retrieval recall as an empirical signal of this bias, but recall gaps also depend on candidate pools and annotation structure and are not direct measurements of bias.
GRAM and Triangle already replace pairwise cosine similarity with geometric quantities involving multiple vectors. The paper argues that changing similarity alone is insufficient: if the geometry still depends on an unoptimized modality-specific reference, other modalities must adapt to that reference. BaryBind therefore separates learning the reference from aligning representations around it. The former uses optimal transport to seek a distribution that is relatively central to several feature distributions; the latter uses the joint geometry of the barycenter and modality gap vectors to construct batch-level positive and negative comparisons. The shared semantic center is a representation-learning objective and interpretation, not a proven unique semantic object that eliminates all modality differences.
A barycenter should not be reduced to averaging the text, video, and audio vectors of one sample. The average is only one possible input to the map; the intended learned object is an output distribution and its generating map. The appendix compares direct averaging, mapping from an average feature, and mapping from text or video, aiming to distinguish reference optimization from initializer choice. Core Idea: first learn a reference constrained jointly by multimodal distributions, then organize multimodal representations through volumetric contrastive learning around that reference, rather than permanently centering the space on a fixed text or image representation.
Method¶
Overall Architecture¶
The input consists of multimodal observations of the same event; the outputs are aligned features and barycenter representations for retrieval, classification, and related tasks. BERT-B encodes text, BEATs encodes audio, and EVA-CLIP-ViT-G encodes visual inputs. The system inherits VAST's backbone but removes its modality fusion layers and adds lightweight MLPs. Subtitles can supply another modality; ChronoDepth generates depth, which enters through an additional lightweight head attached to the visual encoder.
During training, the mean of available modality features or a designated modality feature initializes the barycenter map. The MWB objective attempts to move the output distribution toward a multimodal Wasserstein barycenter. BVC then computes a volume from the reference vector and modality-to-reference gap vectors and contrasts matched samples with batch-level mismatches. DAM supplies a separate matching discriminator for instance-level correspondence; it is not the barycenter solver.
Inference does not require alternating maximization of the potentials. Learned encoders and the barycenter map produce representations for queries, candidates, or classification; the appendix illustrates query-to-barycenter-to-other-modality retrieval. When a modality is absent, only visible inputs are available. A complete-modality average used during training cannot be assumed available at test time.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Available modality inputs"] --> E["Modality encoders"]
E --> W["MWB"]
W -->|Training: barycenter and gap vectors| C["BVC"]
E -->|Training: modality features| C
E -->|Training: matched and mismatched inputs| M["DAM"]
W -.->|Barycenter objective| O["Joint encoder and map updates"]
C -.->|Contrastive supervision| O
M -.->|Matching supervision| O
W -->|Inference: learned barycenter representation| I["Retrieval or classification"]
E -->|Inference: learned modality representations| I
Key Designs¶
1. MWB: learn a distribution-level geometric reference rather than directly averaging samples
Let \(\mathbb{P}_k\) denote each encoded modality distribution. The intended barycenter distribution \(\mathbb{Q}\) minimizes the weighted sum of Wasserstein distances to those distributions. The paper uses a Euclidean-norm transport cost, not a default squared Euclidean cost. This distinction affects barycenter properties, so intuitions from Wasserstein-2 barycenters cannot all be transferred unchanged. The stated objective is:
This concerns transport between probability distributions, not the coordinate-wise arithmetic mean of each paired feature tuple. In practice, the model does not explicitly maintain a complete transport plan. An MLP map \(T_\theta\) transforms initializer \(\bm m_0\) into barycenter feature \(\bm b\), and its outputs collectively define an approximate barycenter distribution. The main method states that the default initializer is the mean of all available features; experimental descriptions and several appendix experiments instead use text. Both choices occur in the paper and should be interpreted according to the particular experimental protocol, not merged into one unambiguous configuration.
To approximate this objective, the authors introduce neural potentials for the modalities and enforce a hard zero-sum constraint by subtracting their weighted mean. Each raw potential is parameterized by a two-layer MLP. Potentials undergo gradient ascent and the map undergoes gradient descent, intending to replace a fixed original anchor and allow all modalities to influence reference updates.
However, the published formulation contains a consequential reproducibility issue. Equation (5) evaluates every potential at the same \(T_\theta(\bm m_0)\). If the terms use the same paired sample or the same barycenter distribution, their weighted potential contributions cancel exactly under the constraint above. A literal implementation would therefore remove any nontrivial role for potential maximization in MWB. It cannot be presented as a confirmed effective dual game without clarification. Modality-dependent transport maps or different sampling mechanisms could require a different analysis, but code or additional documentation is needed to establish whether either is implemented; this note does not invent them.
Appendix A.1 starts from a dual with modality-specific cost transforms, then moves the infimum outside the integral and exchanges summation with an infimum, citing a shared barycenter and linearity. Linearity alone generally does not justify these exchanges. Expectations over marginal distributions are also not automatically equivalent to optimization over valid transport couplings. The appendix temporarily writes the potential congruence condition as 1 and subsequently returns to 0. These are issues in the current derivation, not invitations to replace the authors' algorithm with a reader-reconstructed formula.
Theorem 3.2 relates output-distribution error to map optimization and dual errors, conditional on strong convexity of transport cost minus potential with respect to the barycenter variable. Ordinary MLPs, the Euclidean-norm cost, or weight decay do not automatically establish this assumption; the Euclidean norm itself lacks strict curvature in some directions. The appendix proof additionally relies on a shared optimal map and corresponding transport relationships. It therefore does not show that practical training must converge to a unique correct global semantic center. The result should be read as conditional analysis rather than a guarantee for every implementation.
2. BVC: organize positive and negative samples through the joint volume of barycenter and gap vectors
With a learned reference, the method goes beyond separately comparing text with video and text with audio. It computes each modality's gap vector \(\bm r_k\) relative to the barycenter and assembles those vectors together with the barycenter vector. The barycenter vector retains the reference direction, while gap vectors describe directions of deviation. Their Gram determinant supplies a geometric quantity that depends jointly on multiple vector relationships.
This compact notation follows main-text equation (9). Appendix A.3 actually discusses a cosine Gram matrix of normalized vectors: under that convention, each column must be normalized before constructing the Gram matrix. Raw inner products and cosine similarities should not be conflated. An ordinary Gram determinant gives squared parallelotope volume; the paper calls its quantity a barycenter simplex volume, although a literal simplex volume also involves a factorial factor. Appendix B.8 compares volume normalizations. Accordingly, \(V\) is treated here as the paper's geometric score, not an absolute physical volume directly comparable across different modality counts.
BVC treats a sample's barycenter and its complete group of gap vectors as a positive pair. Negatives pair the current barycenter with another sample's gap-vector group, and vice versa. Negative volume divided by a temperature acts as the similarity score, so a lower volume receives a higher score. The two batch-level contrastive terms respectively hold the barycenter and the gap-vector group fixed. This does not enumerate every possible multimodal mismatch or impose a separate cosine threshold on every modality pair.
Higher-order geometry can capture joint structure missed by pairwise scores, but a small volume is not equivalent to semantically correct alignment of every modality. Linearly dependent columns can make the determinant zero even when some angles are large; opposing directions can also produce degenerate geometry. Normalized gap vectors additionally require an explicit numerical rule for zero gaps. Contrastive negatives, MWB, and DAM are therefore training conditions that connect the geometric score to instance semantics, not incidental decorations.
The appendix reports a -0.9844 correlation between volume and MSR-VTT retrieval performance. This is an empirical association under a particular trajectory or sampling procedure, not a theorem guaranteeing semantic correctness on arbitrary data. Likewise, subtitles and estimated depth may add information derived from an existing video; these results do not establish equal effectiveness for arbitrary independent sensors.
3. DAM: supplement geometric alignment with instance matching supervision
A volume score summarizes feature geometry but still needs a connection to which observations describe the same event. DAM follows VAST's matching discrimination approach: encoded data from a designated anchor are combined with concatenated non-anchor features through cross-attention, and a two-layer MLP predicts a matching probability. Matched and mismatched inputs receive binary supervision, supplying a more direct correspondence signal than geometry alone.
DAM is an auxiliary training path. It neither requires every modality at inference nor reconstructs missing audio. Its contribution appears in the ablation: removing DAM while retaining MWB+BVC changes MSR-VTT T2V/V2T from 56.1/53.6 to 54.8/51.1. This supports the value of matching supervision but also prevents assigning the full system's entire gain to the barycenter solver.
Main-text equation (11) writes a binary log-likelihood without a leading minus sign, whereas equation (12) adds it to an objective that is minimized. Standard binary classification would require a negative log-likelihood or the opposite optimization direction. The current text does not resolve this sign conflict. This note describes the intended supervision and reported ablations without inserting a minus sign and presenting it as the authors' original equation.
A Worked Example¶
Consider a video, audio track, and text description of a cat meowing, with another batch sample describing sea waves. The encoders produce modality features. Their three-modality average can initialize the MLP that produces a barycenter reference for the cat sample; a specified protocol can instead initialize it from text. Averaging the input and learning the output are separate steps. Skipping the map and calling the mean a solved Wasserstein barycenter would erase the distinction the method intends to make.
BVC pairs the cat reference with the cat gap-vector group as a positive example, and with the sea-wave gap-vector group as a negative example, then constructs the reverse comparison. DAM checks instance correspondence in parallel. This example explains the training flow; it does not imply that zero volume identifies a cat and assigns no unreported numerical volume to either sample.
If audio is removed during subsequent continued training, only text and video can drive barycenter optimization. Audio is available again as an actual query or candidate input during AudioCaps retrieval testing. The experiment therefore tests cross-modal transfer after a particular training stage lacks a modality, not the spontaneous acquisition of audio semantics by a model that has never encountered sound.
Loss & Training¶
The overall objective combines MWB, BVC, and DAM. Default weights are 1 for BVC and 0.1 for DAM, the temperature is 0.07, and modality distribution weights are uniform. Algorithm 1 uses ascent for potentials and descent for encoders and the barycenter map; its potential step size is twice that of the other networks. This is the published update procedure, whose interpretation remains subject to the cancellation and sign issues above.
Main experiments continue training an existing VAST model on a 150,000-example subset of VAST-27M because the full dataset is unavailable. Continued training uses one frame and a 10-second audio clip per example, 10,000 steps, batch size 256, an initial learning rate of 1e-4 with linear decay, and two A100 GPUs. Separate training-from-scratch experiments use eight frames, batch size 64, and four epochs. They should not be conflated with the continued-pretraining zero-shot tables or treated as the same training budget.
Zero-shot here denotes downstream evaluation without conventional task-specific fine-tuning. It does not mean that the backbone lacks pretraining, modalities have never been paired, or the target test set supplied no information. Appendix B.2 explicitly states that MSR-VTT test performance is evaluated every 100 steps and used to select the best checkpoint. This introduces a test-set model-selection concern and should be distinguished from zero-shot evaluation with strict test isolation.
Key Experimental Results¶
Main Results¶
The following percentages come directly from main-text Tables 1 and 2. T2V/V2T entries are directional R@1 scores; T-VA denotes the text versus video-and-audio evaluation configuration. Results with additional subtitles or depth do not use the same input budget.
| Dataset and configuration | Metric | VAST | BaryBind | Gain (percentage points) |
|---|---|---|---|---|
| MSR-VTT, T-VA | T2V / V2T R@1 | 49.3 / 43.7 | 56.1 / 53.6 | +6.8 / +9.9 |
| DiDeMo, T-VA | T2V / V2T R@1 | 49.5 / 48.2 | 56.3 / 54.0 | +6.8 / +5.8 |
| ActivityNet, T-VA | T2V / V2T R@1 | 51.4 / 46.8 | 60.6 / 57.2 | +9.2 / +10.4 |
| VATEX, T-VA | T2V / V2T R@1 | 80.0 / 77.3 | 84.7 / 81.8 | +4.7 / +4.5 |
| VGGSound5K, A | Acc@1 / Acc@5 | 40.3 / 71.7 | 45.7 / 75.2 | +5.4 / +3.5 |
| VGGSound5K, V | Acc@1 / Acc@5 | 46.3 / 72.7 | 48.3 / 76.4 | +2.0 / +3.7 |
| VGGSound5K, A+V | Acc@1 / Acc@5 | 48.1 / 79.6 | 55.6 / 83.4 | +7.5 / +3.8 |
The MSR-VTT directional gap falls from 5.6 to 2.5 percentage points, a reduction of 3.1, calculated directly from Table 2's 56.1 and 53.6. The prose instead reports a BaryBind gap of 2.8 and gives DiDeMo T2V as 56.1, whereas Table 2 reports 56.3. This note follows the table while explicitly retaining the discrepancy rather than silently harmonizing the source.
Ablation Study¶
The following subset of main-text Table 5 uses Acc@1 for classification and R@1 for retrieval. TV+TA CL denotes pairwise text-video and text-audio contrastive losses. Each row is a separately trained configuration, not an inference-time modality removal.
| Config | VGGSound A / V / A+V | MSR-VTT T2V / V2T |
|---|---|---|
| TV+TA CL | 38.1 / 42.8 / 44.5 | 46.8 / 40.1 |
| TV+TA CL + DAM | 40.3 / 43.6 / 46.3 | 49.3 / 43.7 |
| TV+TA CL + MWB + DAM | 44.3 / 46.8 / 49.8 | 50.6 / 48.8 |
| BVC + DAM | 42.9 / 47.1 / 50.1 | 51.5 / 46.8 |
| MWB + BVC | 45.2 / 47.8 / 54.7 | 54.8 / 51.1 |
| MWB + BVC + DAM | 45.7 / 48.3 / 55.6 | 56.1 / 53.6 |
Table 5's TV+TA CL+DAM classification values differ from Table 1's VAST values: V/A+V are 43.6/46.3 in the former and 46.3/48.1 in the latter. Ablation gains should therefore be calculated within Table 5, not by substituting classification baselines from Table 1. The ablation prose calls 55.6 audio classification, although Table 5 assigns it to A+V; this note retains the correct column assignment.
Missing-modality results come from main-text Tables 3 and 4 and must distinguish training-time from inference-time absence.
| Task and protocol | Metric | VAST | BaryBind |
|---|---|---|---|
| AudioCaps, trained with T-VA | T2A / A2T R@1 | 32.1 / 26.1 | 35.7 / 32.5 |
| AudioCaps, trained with T-V only | T2A / A2T R@1 | 10.4 / 6.7 | 21.2 / 14.5 |
| VGGSound5K, trained with A+V, inference with A+V | Acc@1 / Acc@5 | 48.1 / 79.6 | 55.6 / 83.4 |
| VGGSound5K, trained with A+V, inference with A only | Acc@1 / Acc@5 | 40.8 / 71.6 | 49.4 / 78.3 |
Key Findings¶
- Within Table 5's protocol, adding MWB to TV+TA CL+DAM raises V2T from 43.7 to 48.8. Relative to BVC+DAM, the full configuration raises V2T from 46.8 to 53.6, supporting complementarity between reference learning and geometric alignment.
- Without audio in training, BaryBind obtains T2A/A2T of 21.2/14.5, below 35.7/32.5 with complete training inputs. Its advantage is reduced degradation relative to baselines, not cost-free modality absence. The protocol does not fully explain how pre-existing VAST audio knowledge is retained.
- With T-VASD, main-text Table 2 reports MSR-VTT 57.8/54.6, but appendix Tables 11 and 22 give different values for related configurations. These appendix results should not all be treated as exact reproductions of one evaluation.
- Appendix Table 24 reports TA2I FID of 38.6 versus VAST's 50.6. Generation uses a separate CLIP ViT-H/14 encoder, ImageBind audio encoder, and Stable UnCLIP decoder; it does not establish generation capability in the main retrieval backbone itself.
Highlights & Insights¶
- Separating reference learning from alignment around that reference provides a clearer intervention than changing similarity alone. The appendix's mean-feature initializer comparison is especially useful because symmetric fusion inputs should not be mistaken for completed distribution-level barycenter optimization.
- Reporting bidirectional retrieval alongside unimodal and joint-modality classification helps reveal gains limited to one query direction. These remain proxies for modality balance and require interpretation in light of candidate distributions, annotation granularity, and task difficulty.
- The missing-modality experiments impose two distinct stresses: no audio during continued training and no video during inference. They should not be collapsed into an unrestricted claim of robustness to arbitrary modality absence.
Limitations & Future Work¶
- The interface between theory and implementation is unclear: hard zero-sum potentials cancel at a shared barycenter, and infimum exchanges in the dual derivation need additional conditions. Executable code and sampling details should precede conclusions about whether the solver implements the stated distribution-level transport objective.
- Degenerate BVC volumes do not guarantee semantically correct alignment of every modality pair. Further evidence should include worst-pair performance, mismatched or contradictory inputs, and explicit handling of normalization, zero vectors, and determinant stability.
- Several prose-table and main-text-appendix numbers disagree, and error bars are absent. The appendix acknowledges fixed random seeds; a fixed seed does not substitute for statistical significance analysis across repeated runs.
- Test-set checkpoint selection weakens strict zero-shot generalization claims. An independent validation split is needed, together with clarification of which pretrained audio modules are frozen, updated, or retained when audio is absent from training.
- Appendix Table 19 reports per-step time increasing from about 8.8 seconds for VAST to about 9.7 seconds for BaryBind, with parameters rising from 1.28B to 1.34B. The overhead is not zero. Transfer to large multimodal language models, reconstruction objectives, and genuinely independent sensors remains future work rather than demonstrated capability.
Related Work & Insights¶
- vs VAST: VAST supplies the pretrained backbone and principal comparison. BaryBind removes fusion layers and adds a barycenter map and joint geometric objectives. The evidence is primarily continued-training improvement under that backbone and data budget.
- vs ImageBind / LanguageBind: These bind other modalities through a designated modality, whereas BaryBind seeks to optimize the reference itself. A text initializer is compatible with a final reference that is no longer fixed to the original text feature, but does not guarantee freedom from bias.
- vs GRAM / Triangle: All use geometric relationships among multiple vectors. BaryBind adds barycenter optimization and scores a structure built from the barycenter and gap vectors. The substantive question is whether reference optimization supplies additional stable gains, not merely a different geometric name.
Rating¶
- Novelty: 4/5 โ The combination of learned geometric reference and volumetric alignment is distinctive, but its dual solver implementation needs clarification.
- Experimental Thoroughness: 3/5 โ Broad tasks and ablations are weakened by test-set selection, missing error bars, and numerical discrepancies.
- Writing Quality: 2/5 โ The main narrative is clear, but formulas, column assignments, and cross-table consistency contain substantive issues.
- Value: 4/5 โ A useful perspective on multimodal anchor bias, with implementation verification required before practical reuse.