MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Full Text: ECCV Open Access
Area: Audio & Speech
Keywords: self-supervised learning, audio-visual representation learning, joint-embedding predictive architecture, cross-modal alignment, frozen evaluation
TL;DR¶
MJEPA introduces an elegant and scalable audio-visual self-supervised learning framework that trains a single, unified ViT encoder using only joint-embedding predictive objectives in latent space; by combining intra-modal multi-level feature prediction with cross-modal bidirectional alignment, it sets new state-of-the-art records across seven frozen evaluation benchmarks without requiring contrastive objectives, pixel/spectrogram reconstruction, or heavy data augmentations.
Background & Motivation¶
Learning visual representations from large-scale unlabeled video data has become a foundational pillar of modern computer vision. Because natural video streams are inherently multimodalโwhere visual frames and complementary acoustic signals physically co-occurโextending self-supervised learning to jointly capture both modalities is a compelling and natural frontier. Nevertheless, existing audio-visual self-supervised learning (AV-SSL) approaches (such as CAV-MAE, EquiAV, and MAViL) remain encumbered by complex, multi-component systems: they predominantly deploy separate, modality-specific encoders and stitch together intricate loss mixtures spanning contrastive objectives, masked pixel/spectrogram reconstruction, cross-modal distillation, and hand-crafted data augmentations. This structural segregation and objective complexity severely constrain cross-modal synergy while hindering seamless scaling to billion-parameter architectures and heterogeneous unpaired data.
A profound empirical tension lies in model parameter sharing across modalities. Prior investigations (including VATT, XKD, and AVSiam) observed that simply replacing modality-specific encoders with a shared backbone causes individual unimodal representations to degrade sharply below their respective unimodal baselines. Without proper constraints, shared encoder weights are pulled in opposing directions by disparate acoustic spectral patterns and visual scene textures. While Joint-Embedding Predictive Architectures (JEPAs) have demonstrated remarkable efficacy in unimodal domains (I-JEPA in images, V-JEPA in video, and A-JEPA in audio) by predicting abstract representations of masked tokens directly in latent space without low-level pixel reconstruction, their formulation has until now remained strictly confined to unimodal paradigms.
This paper approaches the problem through the lens of the Platonic Representation Hypothesis: high-level representations across distinct modalities naturally converge toward a unified statistical model of the physical world. Therefore, a shared encoder is not only structurally viable but fundamentally advantageous, provided an explicit, lightweight cross-modal predictive mechanism aligns high-level semantic manifolds and resolves representational competition. The core idea is to train a single, unified audio-visual encoder exclusively with JEPA predictive objectivesโcoupling multi-level intra-modal masked feature prediction with global mean-pooled bidirectional cross-modal alignmentโthereby resolving parameter interference, driving bidirectional positive transfer, and enabling seamless scaling to a 1-billion-parameter ViT-g on heterogeneous video datasets.
Method¶
Overall Architecture¶
The architecture of MJEPA is characterized by radical simplicity and structural unity. Continuous audio inputs (standardized to 10-second segments and converted into \(1024 \times 128\) log-mel spectrograms) and video sequences are first projected into a shared hidden dimension via minimal modality-specific tokenizers (a 2D convolution for audio spectrogram patches and a 3D convolution for video spatio-temporal tubelets). Modality-specific learnable tokens and absolute sincos positional embeddings are added to inject spatial, temporal, and frequency topologies. The model natively accepts three input streams: audio-only (\(a\)), video-only (\(v\)), and concatenated audio-video sequences (\(av\)). A single shared context encoder \(E_\theta\) processes visible tokens, while an exponential moving average (EMA) target encoder \(E_{\bar{\theta}}\) processes full, unmasked sequences to provide stable target representations under a stop-gradient constraint.
The training framework is driven by two complementary predictive mechanisms: (1) an intra-modal joint-embedding prediction pipeline where a shared narrow Transformer predictor \(P_\phi\) fuses multi-level intermediate encoder features to recover target representations at masked locations; and (2) a cross-modal bidirectional prediction pipeline where six lightweight 3-layer MLP predictors \(C^{m_1 \to m_2}_\psi\) take mean-pooled last-layer source features and regress to unmasked mean-pooled last-layer target features. The entire system is trained end-to-end using an unweighted sum of nine smooth \(L_1\) distances, avoiding negative pairs, codebooks, or reconstruction decoders.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Continuous Multimodal Inputs<br/>Audio Spectrograms / Video Clips / Paired Streams"] --> B["Modality-Specific Tokenization & Shared Encoder<br/>2D/3D Conv Projections + Spatiotemporal Embeddings + Unified ViT"]
B --> C["Multi-Level Intra-Modal Latent Prediction<br/>Concatenate Multi-Layer Features + MLP + Masked Token Recovery"]
B --> D["Global Pooled Cross-Modal Predictive Alignment<br/>Last-Layer Mean-Pooling + 3-Layer MLP Bidirectional Regressions"]
C --> E["Unified Optimization & Heterogeneous Scaling<br/>9 L1 Predictive Objectives + Distributed GPU Workload Partitioning"]
D --> E
E --> F["Downstream Frozen Feature Probing<br/>Single Frozen ViT Backbone + 4-Layer Attentive Probe"]
Key Designs¶
1. Modality-Specific Tokenization & Shared Encoder: A Single Physical Carrier for All Modalities
To overcome the architectural redundancy of dual-tower networks that fail to share fundamental spatio-temporal abstractions, MJEPA deploys a single Vision Transformer backbone for both audio and visual processing. Raw audio signals are mapped into \(1024 \times 128\) log-mel spectrograms and tokenized using 2D convolutions, accompanied by 2D sincos absolute positional embeddings (time-frequency). Video clips are tokenized via 3D convolutions into tubelets, accompanied by 3D sincos positional embeddings (time-height-width). Modality-specific learnable embeddings are added to retain modality identities. Whether fed unimodal audio tokens, unimodal video tokens, or concatenated multimodal sequences, all tokens pass through the exact same Transformer blocks. This design establishes a unified representation space while reducing the model footprint to a single cohesive backbone.
2. Multi-Level Intra-Modal Latent Prediction: Capturing Structural and Abstract Hierarchies
Restricting latent feature prediction solely to the final encoder layer often causes intermediate representations to miss fine-grained spatio-temporal dynamics and acoustic textures. MJEPA incorporates multi-level feature extraction within the intra-modal objective across all three input modes (\(a \to a\), \(v \to v\), \(av \to av\)). For the ViT-L variant, intermediate features are extracted from layers \(\{5, 11, 17, 23\}\); for ViT-g, layers \(\{9, 19, 29, 39\}\) are sampled. These intermediate representations are concatenated along the channel dimension, projected back to the model dimension via a 2-layer MLP, and supplied alongside learnable mask tokens to a narrow ViT predictor \(P_\phi\) (embedding dimension 384). The target representations are extracted from the unmasked sequence via the stop-gradient EMA target encoder \(E_{\bar{\theta}}\):
where \(M\) denotes the set of masked token indices, \(\Delta_m\) represents positional mask tokens, and \(\text{sg}(\cdot)\) designates the stop-gradient operator. Predicting across multiple latent tiers enables the encoder to retain rich structural detail without regressing onto high-frequency pixel noise.
3. Global Pooled Cross-Modal Predictive Alignment: Resolving Interference and Bridging Modalities
A naive shared encoder trained exclusively with intra-modal objectives exhibits catastrophic representational interference, as individual modality gradients destabilize one another. Furthermore, audio and video streams lack explicit token-to-token spatial correspondences, rendering dense local cross-modal matching suboptimal and prone to spurious correlations. MJEPA introduces a global predictive alignment objective operating exclusively on top-layer semantic features. Six bidirectional cross-modal predictive objectives (\(a \leftrightarrow v\), \(a \leftrightarrow av\), \(v \leftrightarrow av\)) are trained using minimal 3-layer MLPs:
where \(\text{pool}(\cdot)\) denotes global mean pooling across the token sequence, and \(C_\psi^{m_1 \to m_2}\) represents the cross-modal predictor. By forcing the pooled output of a masked source modality to predict the unmasked pooled representation of a target modality, top-layer representations are aligned onto a shared semantic manifold. This turns negative parameter conflict into bidirectional positive transfer and spontaneously enables high-precision cross-modal retrieval without negative contrastive mining.
4. Heterogeneous Data Partitioning & Model Scaling: Positive Transfer from Unpaired Video
High-quality paired audio-visual video clips are scarce, whereas silent or uncurated video footage is abundant. Because MJEPA's shared encoder natively accepts pure video inputs without architectural modifications, the framework naturally scales across heterogeneous datasets. Beyond the paired AudioSet-2M dataset (\(\sim\)1.8M 10-second clips, \(\sim\)5k hours), MJEPA directly incorporates the unpaired VideoMix2M dataset (\(\sim\)2M videos, \(\sim\)136k hours from HowTo100M, Kinetics-710, and SSv2). In distributed training, compute is partitioned: one GPU subset processes paired AV data across all nine loss terms, while the remaining GPUs process video-only samples optimizing exclusively \(\mathcal{L}_{v \to v}\), weighted by a factor of 5.0 to balance gradient magnitudes. Crucially, the semantic alignment bridge ensures that adding unpaired video data significantly improves not only video features but also audio-only representation quality, while scaling seamlessly to a 1-billion-parameter ViT-g.
Loss & Training¶
The total pre-training objective of MJEPA is the unweighted sum of three intra-modal multi-level prediction losses and six cross-modal global prediction losses:
All models are optimized with AdamW (\(\beta_1=0.9, \beta_2=0.995\)) in bfloat16 precision for 150k steps. Audio inputs undergo 75%โ80% patch masking. For base ViT-L, video uses 8 small temporal blocks (15% frame area); scaled ViT-L and ViT-g adopt a multi-block strategy featuring 8 small blocks alongside 2 large blocks (70% frame area). Scaled models subsequently undergo a 12k-step high-resolution cooldown phase (64 frames with linear learning rate decay). Downstream evaluation strictly adheres to the frozen probing protocol, employing an attentive probe featuring 1 learnable query token and 4 Transformer blocks.
Key Experimental Results¶
Main Results¶
MJEPA was thoroughly benchmarked on AudioSet-20K (audio-visual, audio-only, video-only) and dedicated environmental sound classification benchmarks (ESC-50, FSD50K). All evaluations employ the attentive probe on frozen representations. As reported below, frozen MJEPA substantially outperforms previous frozen baselines and rivals or surpasses full-model finetuning benchmarks.
| Method | Params | Pre-train Data | Evaluation Protocol | AS20K (A) mAP | AS20K (V) mAP | AS20K (A-V) mAP | ESC-50 Acc (%) | FSD50K mAP |
|---|---|---|---|---|---|---|---|---|
| CAV-MAE | 170M | IN+AS | Frozen Feature | 19.38 | 18.14 | 34.59 | 77.5 | 46.1 |
| MAViL | 170M | IN+AS | Frozen Feature | 30.00 | โ | โ | 90.8 | โ |
| EquiAV | 170M | IN+AS / AS | Frozen Feature | 34.25 | 18.60 | 38.60 | 93.2 | 57.9 |
| CAV-MAE Sync | 170M | AS | Frozen Feature | 21.66 | 16.20 | 28.50 | 89.2 | 55.5 |
| MJEPA ViT-L (Ours) | 300M | AS | Frozen Feature | 38.89 | 25.38 | 42.90 | 95.2 | 63.9 |
| MJEPA ViT-L + Data Scaling (Ours) | 300M | AS+VM2M | Frozen Feature | 40.00 | 29.63 | 45.31 | 96.8 | 65.5 |
| MJEPA ViT-g + Data Scaling (Ours) | 1B | AS+VM2M | Frozen Feature | 40.97 | 29.82 | 45.44 | 96.9 | 65.8 |
| CAV-MAE (Reported Finetuned) | 170M | IN+AS | Full Finetuning | 37.70 | 19.80 | 42.00 | โ | โ |
| MAViL (Reported Finetuned) | 170M | IN+AS | Full Finetuning | 41.80 | 24.80 | 44.90 | 94.4 | โ |
| EquiAV (Reported Finetuned) | 170M | AS | Full Finetuning | 42.40 | 25.70 | 46.60 | 96.0 | 62.6 |
Note: Finetuned baseline metrics are cited directly from original publications; AS = AudioSet-2M, IN = ImageNet-1K, VM2M = VideoMix2M.
Ablation Study¶
The author conducted an incremental ablation study on AudioSet-20K to dissect the impact of each architectural stage. Across all dimensions, cross-modal alignment and joint encoding proved indispensable.
| Staged Configuration | Description & Objectives | AS20K (A) | AS20K (V) | AS20K (AV) | Key Takeaway & Analysis |
|---|---|---|---|---|---|
| 1. Unimodal Baselines | Separate A-only (\(L_{a \to a}\)) & V-only (\(L_{v \to v}\)) encoders | 30.87 | 19.84 | โ | Standard single-modality baselines; no interaction |
| 2. Naive Shared Encoder | Single shared ViT; only \(L_{a \to a} + L_{v \to v}\) (no alignment) | 28.70 | 17.90 | 32.38 | Severe degradation in both modalities (A: -2.17, V: -1.94) due to parameter conflict |
| 3. + Cross-Modal Alignment | Shared ViT + top-layer bidirectional \(L_{a \leftrightarrow v}\) | 33.52 | 19.58 | 36.34 | Eliminates degradation immediately; audio exceeds unimodal baseline (+2.65) |
| 4. + Joint AV Encoding (Full MJEPA) | Supports concatenated inputs; adds \(L_{av \to av}\) and \(L_{a/v \leftrightarrow av}\) | 38.89 | 25.38 | 42.90 | Massive multimodal synergy (+8.02 A, +5.54 V over unimodal baselines) |
| 5. + Data Scaling | ViT-L trained on AS2M + unpaired VM2M | 40.00 | 29.63 | 45.31 | Unpaired video data directly enhances audio features (A: 38.89 \(\to\) 40.00) |
| 6. + Model Scaling | Scaled to ViT-g (1B) with high-res Cooldown | 40.97 | 29.82 | 45.44 | Sets state-of-the-art frozen benchmark; rivals top finetuned models |
On downstream video benchmarks, MJEPA ViT-L trained on AS+VM2M achieves 84.7% top-1 accuracy on Kinetics-400 and 73.3% on SSv2, closely matching dedicated video models like V-JEPA 2 ViT-L (85.1% / 73.7%) despite the latter requiring \(10\times\) more video training volume (VM22M). In cross-modal retrieval on AudioSet (without contrastive learning), MJEPA attains 29.3% R@1 on Audio\(\to\)Video and 30.6% R@1 on Video\(\to\)Audio, rivaling EquiAV's 29.6% and 30.1% achieved via explicit contrastive losses.
Key Findings¶
- Cross-modal predictive alignment is essential for shared encoders: Sharing parameters across modalities without cross-modal objectives triggers catastrophic parameter interference. Introducing global predictive alignment not only eliminates this degradation but fosters strong positive transfer.
- Unpaired video transfers positively to audio representations: Incorporating video-only data into training boosts audio downstream classification monotonically across AS20K, ESC-50, and FSD50K, verifying that cross-modal predictive alignment creates genuine semantic synergy.
- Latent predictive architectures suffice for cross-modal retrieval: Contrary to the common belief that cross-modal retrieval requires explicit negative pair repulsion (InfoNCE), pure \(L_1\) prediction between mean-pooled global embeddings induces a structured, highly discriminative cross-modal metric space.
Highlights & Insights¶
- Radical architectural simplicity: Completely avoids negative queues, memory banks, pixel/spectrogram reconstruction decoders, and handcrafted geometric augmentations. The entire training loop relies solely on smooth \(L_1\) distance in embedding space.
- Hierarchical task separation (multi-level local vs. pooled global): Confines intra-modal learning to intermediate spatial/spectral masked token recovery, while restricting cross-modal alignment to global pooled representations, aligning with high-level conceptual convergence.
- Native flexibility for heterogeneous data ingestion: The unified tokenization and shared backbone accept arbitrary unimodal and multimodal streams without architectural redesign, opening an efficient pathway for incorporating massive web-scale non-paired data.
Limitations & Future Work¶
- Coarse cross-modal alignment granularity: Relying entirely on global mean pooling omits fine-grained spatio-temporal audio-visual correspondences (e.g., localized sound sources and motion rhythms). Incorporating lightweight cross-attention or deformable routing could enhance fine-grained perception.
- Absence of language and textual semantics: MJEPA focuses exclusively on acoustic and visual physical signals. Expanding this framework into a tri-modal Audio-Visual-Language JEPA represents a natural and promising next frontier.
- Training compute overhead: While avoiding pixel-space decoders, optimizing nine predictive losses and scaling a 1B ViT-g model with high-resolution cooldown requires substantial distributed GPU infrastructure.
Related Work & Insights¶
- vs. CAV-MAE / MAViL: CAV-MAE and MAViL rely on dual-stream backbones combining contrastive learning with masked autoencoding and require ImageNet pre-training. MJEPA trains a single shared encoder from scratch using pure latent prediction, achieving superior frozen representations.
- vs. EquiAV: EquiAV demands aggressive equivariance-based data augmentations and contrastive pairs. MJEPA relies strictly on uncorrupted physical correspondences and JEPA prediction without augmentations.
- vs. V-JEPA / A-JEPA: While V-JEPA and A-JEPA established predictive SSL in unimodal settings, MJEPA is the first to prove that a single shared encoder can bridge both modalities without negative transfer, highlighting the indispensable role of cross-modal prediction.
Rating¶
- Novelty: โญโญโญโญโญ Pioneering unification of audio-visual self-supervised learning under a single JEPA encoder with pure predictive objectives.
- Experimental Thoroughness: โญโญโญโญโญ Evaluated across seven benchmarks spanning audio, video, multimodal classification, and cross-modal retrieval with progressive ablations up to 1B parameters.
- Writing Quality: โญโญโญโญโญ Outstanding clarity, rigorous logical progression, and transparent experimental reporting.
- Value: โญโญโญโญโญ Establishes a highly scalable, minimal-complexity paradigm for multimodal foundation models.