Skip to content

Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

Conference: ECCV 2026
Paper: ECCV Official
Area: Video Understanding
Keywords: text-to-video retrieval, probabilistic embedding, distribution bridge, uncertainty modeling, contrastive learning

TL;DR

Addressing the failure of deterministic point matching to account for cross-modal ambiguity and one-to-many semantic associations, this paper proposes the Distribution-Alignment Bridge (DAB) framework, which models texts and videos as Gaussian distributions and aligns them via a sampling-free deterministic diffusion bridge and directional KL contrastive learning, achieving marked recall improvements and uncertainty-calibrated retrieval.

Background & Motivation

Text-to-video retrieval (TVR) fundamentally contends with the modality gapโ€”the profound representational disparity between compact, symbolic textual descriptions and dense, spatio-temporally rich video sequences. Although pretrained vision-language models such as CLIP combined with cross-modal attention mechanisms have significantly advanced the field, prevailing methods uniformly project queries and video items into single deterministic embedding vectors. Such point-wise matching implicitly presumes a deterministic one-to-one mapping, overlooking the intrinsic ambiguity and one-to-many correspondences of multimodal data: a single concise query can truthfully depict diverse video scenes, while a complex video can be described by multiple semantically distinct captions.

Recent efforts like DITS borrow concepts from diffusion models to recast cross-modal alignment into iterative refinement rather than one-shot projection. While DITS successfully avoids the training instability and poor ranking behavior caused by generative diffusion models with pure Gaussian noise and \(L_2\) objectives, its optimization remains confined to deterministic feature vectors, ignoring semantic uncertainty and contextual variance. Conversely, earlier probabilistic embedding frameworks like UATVR introduce Gaussian distributions to capture semantic variance, but they rely on stochastic prototype sampling during similarity computation, collapsing distributions back into noisy point comparisons and impairing training stability.

Addressing this tension between point-level limitations and stochastic sampling instability, this work constructs a deterministic trajectory directly within the probabilistic distribution parameter space. Text and video modalities are explicitly represented as Gaussian distributions characterized by semantic means and uncertainty variances, guided by a lightweight truncated drift network that transports the text distribution toward the video distribution. Core Idea: Reframe text-to-video retrieval from deterministic point matching into sampling-free distribution alignment, leveraging a deterministic distribution bridge to jointly refine semantic centers and uncertainty scales, optimized via directional KL divergence contrastive loss to achieve calibrated, globally coherent retrieval.

Method

Overall Architecture

The DAB framework comprises three primary components: a foundational feature extraction backbone with cross-modal attention pooling, a Gaussian probabilistic encoder that disentangles semantic centers from variance, and a sampling-free Distribution Bridge. First, a pretrained CLIP extracts frame-level and text-level embeddings, followed by a transformer-based cross-attention module that pools video frames conditioned on the text query. Next, a shared probabilistic encoder maps features into Gaussian parameters \((\mu, \log \sigma^2)\), with a dedicated token-based uncertainty head capturing temporal variance across video frames. Finally, conditioned on discrete timesteps, the Distribution Bridge applies a lightweight MLP to predict residual drift terms that iteratively transport the text distribution toward the video distribution, optimized end-to-end via a directional KL divergence-based Probabilistic Alignment Contrastive (PAC) loss.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multimodal Inputs<br/>Text queries and uniformly sampled video frames"] --> B["CLIP Foundational Features & Cross-Attention Pooling"]
    B --> C["Probabilistic Embedding Module: Mean & Variance Disentanglement"]
    C --> D["Distribution Bridge: Time-Conditioned Deterministic Drift"]
    D --> E["Probabilistic Alignment Contrastive Loss: Directional KL Optimization"]
    E --> F["Retrieval Ranking & Calibrated Uncertainty Margin Output"]

Key Designs

1. Probabilistic Embedding Module: Disentangling Semantic Centers from Modality Uncertainty To resolve the incapacity of deterministic point vectors to represent ambiguity and one-to-many associations, this module maps each modality into a diagonal Gaussian distribution \(\mathcal{N}(\mu, \operatorname{diag}(\sigma^2))\). Input embeddings are \(L_2\)-normalized and fed into two independent linear heads (preceded by LayerNorm) to predict the semantic mean \(\mu \in \mathbb{R}^{B \times D}\) and log-variance \(\log \sigma^2 \in \mathbb{R}^{B \times D}\). The log-variance bias is initialized to -2.0 to foster low initial uncertainty and clipped to \([-6, 2]\) for numerical stability. For video features, a token-based uncertainty head captures intra-video temporal dynamics: frame-level tokens pass through single-head self-attention, are mean-pooled, and residually combined with global video representations to yield the final video variance \(\log \sigma_v^2 = \frac{1}{2}(\log \sigma_{v,\text{base}}^2 + \log \sigma_{v,\text{tok}}^2)\). This decouples core semantics (mean) from temporal variability (variance), allowing the video modality to naturally maintain higher contextual entropy than text.

2. Distribution Bridge: Time-Conditioned Deterministic Iterative Drift To overcome the excessive variance and ranking instability caused by random noise injection in generative diffusion models, DAB establishes a sampling-free, deterministic distribution bridge. Given a text distribution \(P_t = \mathcal{N}(\mu_t, \sigma_t^2)\) and a target video distribution \(Q_v = \mathcal{N}(\mu_v, \sigma_v^2)\), the bridge executes truncated refinement across discrete timesteps \(t \in [1, T']\). At each step, a sinusoidal time embedding and the current distribution parameters are fed into DriftDistMLP to predict residual parameter drifts \([\Delta \mu_t, \Delta(\log \sigma_t^2)]\). The updates are scaled by a linear step-size schedule \(\chi_t \in [0.2, 0.6]\): $$ \mu_{t-1} = \mu_t - \chi_t \cdot \Delta \mu_t, \quad \log \sigma_{t-1}^2 = \log \sigma_t^2 - \chi_t \cdot \Delta(\log \sigma_t^2) $$ The drift output layers are zero-initialized to prevent abrupt early disturbances, and log-variances are consistently clipped within \([-6, 2]\). Guided by contractive updates (contraction rate at most 0.8), the text distribution deterministically and smoothly converges toward the target video distribution \(\hat{P}_t\) without any stochastic sampling.

3. Probabilistic Alignment Contrastive Loss: Directional KL Divergence for Ranking Optimization Cross-modal distribution discrepancy is highly sensitive to the chosen divergence metric. While symmetric metrics like Wasserstein-2 (\(W_2\)) provide uniform geometric distances, they lack sensitivity to asymmetric cross-modal uncertainty and fail to aggressively penalize tail variance mismatches. DAB adopts the closed-form directional Kullback-Leibler (KL) divergence between diagonal Gaussians as the pairwise cost \(C_{ij} = D_{\mathrm{KL}}(Q_{v_j} \parallel \hat{P}_{t_i})\): $$ D_{\mathrm{KL}}(Q \parallel P) = \frac{1}{2} \sum_{d=1}^D \left[ \frac{(\mu_{Q,d} - \mu_{P,d})^2}{\sigma_{P,d}^2} + \frac{\sigma_{Q,d}^2}{\sigma_{P,d}^2} - \log \frac{\sigma_{Q,d}^2}{\sigma_{P,d}^2} - 1 \right] $$ The model then minimizes a bidirectional Probabilistic Alignment Contrastive (PAC) loss: $$ \mathcal{L}{\mathrm{PAC}} = - \frac{1}{B} \sum \right] $$ This loss penalizes directional mismatches relative to in-batch negative pairs. The variance-sensitive gradients of directional KL divergence actively pull the bridged text variance to subsume video variance without enforcing variance homogenization, thereby preserving authentic modality-specific uncertainty.}^B \left[ \log \frac{\exp(-C_{ii}/\tau)}{\sum_{j=1}^B \exp(-C_{ij}/\tau)} + \log \frac{\exp(-C_{ii}/\tau)}{\sum_{j=1}^B \exp(-C_{ji}/\tau)

Loss & Training

The entire networkโ€”including CLIP projection layers, cross-attention pooling, probabilistic heads, and the distribution bridgeโ€”is trained end-to-end. We employ AdamW with weight decay 0.02, setting learning rates to \(1\times 10^{-6}\) for CLIP encoders and \(1\times 10^{-4}\) for probabilistic and bridge modules, with a 10% linear warmup ratio. Truncated steps \(T'\) are set to 32 out of \(N=1000\) diffusion steps. Batch size is fixed to 32, and models are trained for 10 epochs across MSR-VTT, MSVD, and VATEX on a single NVIDIA A6000 GPU.

Key Experimental Results

Main Results

DAB is evaluated against representative deterministic baselines (CLIP4Clip, X-Pool), stochastic probabilistic frameworks (UATVR, T-MASS), and diffusion-based retrieval methods (DiffusionRet, DITS) on MSR-VTT, MSVD, and VATEX.

Dataset Backbone Method R@1 (%) R@5 (%) R@10 (%) MdR MnR
MSR-VTT (1K-A) ViT-B/32 CLIP4Clip (Neurocomputing'22) 44.5 71.4 81.6 2 15.3
MSR-VTT (1K-A) ViT-B/32 UATVR (ICCV'23) 47.5 73.9 83.5 2 12.9
MSR-VTT (1K-A) ViT-B/32 DITS (NeurIPS'24) 51.9 75.7 84.6 1 11.6
MSR-VTT (1K-A) ViT-B/32 DAB (Ours) 56.2 85.4 92.4 1 4.1
MSR-VTT (1K-A) ViT-B/16 DITS (NeurIPS'24) 55.0 79.8 87.1 1 10.0
MSR-VTT (1K-A) ViT-B/16 DAB (Ours) 57.6 86.2 92.4 1 4.2
MSVD ViT-B/32 CLIP4Clip (Neurocomputing'22) 45.2 75.5 84.3 2 10.0
MSVD ViT-B/32 Cap4Video (CVPR'23) 51.8 80.8 88.3 1 -
MSVD ViT-B/32 DAB (Ours) 54.5 87.8 93.9 1 3.4
VATEX ViT-B/32 DITS (NeurIPS'24) 64.1 92.7 97.0 1 2.9
VATEX ViT-B/32 DAB (Ours) 65.4 93.1 97.2 1 2.4
VATEX ViT-B/16 NarVid (CVPR'25) 68.4 94.0 97.1 1 -
VATEX ViT-B/16 DAB (Ours) 70.7 95.7 98.3 1 1.9

Ablation Study

On MSR-VTT (CLIP ViT-B/32, 5 training epochs), component contributions, distance metrics, and truncated step counts \(T'\) are thoroughly ablated.

Dimension / Setting R@1 (%) R@5 (%) R@10 (%) MnR Latency (s) FLOPs (T)
Component: Baseline (X-Pool) 46.9 72.8 82.2 14.3 - -
Component: + Probabilistic Embedding (w/ Prob.) 44.3 76.6 87.2 7.4 - -
Component: + Deterministic Bridge (w/ Bridge) 43.2 76.6 86.3 8.5 - -
Component: Full DAB (Prob + Bridge) 52.6 82.7 90.8 5.3 51.8 278.93
Metric: Wasserstein-2 Distance (\(W_2\)) 46.8 79.7 88.4 6.6 - -
Metric: Directional KL Divergence (\(D_{\mathrm{KL}}\)) 52.6 82.7 90.8 5.3 - -
Step Count: \(T'=8\) 49.0 82.0 89.0 5.8 39.5 69.73
Step Count: \(T'=16\) 50.6 80.7 89.4 5.6 43.9 139.47
Step Count: \(T'=32\) (Default) 52.6 82.7 90.8 5.3 51.8 278.93
Step Count: \(T'=64\) 51.9 84.2 91.1 5.6 68.9 557.85

Key Findings

  • Dramatic Drop in Mean Rank (MnR): On MSR-VTT, DAB reduces MnR from 11.6 (DITS) to 4.14โ€”a 64% relative reduction. This indicates that DAB does not merely flip top-1 rankings on easy samples, but systematically reorganizes the global cross-modal space, eliminating distant negative outliers.
  • Symbiotic Role of Probabilistic Modeling and Bridge Refinement: Applying probabilistic modeling alone or bridge refinement alone degrades R@1 from 46.9% to 44.3% and 43.2%, respectively. Only their combination unlocks distribution-level iterative drift, surging R@1 to 52.6%.
  • Preserved Modality-Specific Asymmetry: The converged video variance (\(\sigma_v^2 = 2.979\)) remains 5.87 times larger than the text variance (\(\sigma_t^2 = 0.507\)), while the bridge intermediate variance (0.935) settles between them, proving that the bridge respects modality entropy differences without collapsing variance scales.
  • KL Margin Calibrates Retrieval Confidence: Sorting queries by their top-1 vs. top-2 KL divergence margin yields 85.6% R@1 at 25% coverage and 94.0% R@1 at 10% coverage. The Area Under the Risk-Coverage curve (AURC) drops from 0.429 to 0.256, verifying that bridge-induced divergence margins serve as a reliable selective retrieval signal.

Highlights & Insights

  • Sampling-Free Deterministic Diffusion: Discards high-variance Monte Carlo sampling from diffusion architectures, operating directly on compact Gaussian distribution parameters \((\mu, \log \sigma^2)\) via contractive residual drift.
  • Directional KL Divergence as an Adaptive Scaling Metric: Directional KL naturally respects the asymmetric information capacity of text and video, where ratio-based variance gradients act as implicit dynamic hard-negative weighting.
  • Local Semantic-Neighborhood Clustering: For queries with broad semantic scope, DAB clusters related candidate videos in the immediate top-ranked neighborhood rather than treating them as disconnected negatives, demonstrating robust semantic tolerance.

Limitations & Future Work

  • Iterative Refinement Latency Overhead: Although much faster than 1000-step generative diffusion, 32 refinement steps introduce ~30% computational latency overhead compared to single-step projection. One-step bridge distillation represents a promising future avenue.
  • Independent Diagonal Covariance Simplification: The diagonal covariance assumption overlooks inter-channel visual-semantic correlations; low-rank or full covariance formulations could potentially model richer cross-modal manifolds.
  • vs DITS (Wang et al., NeurIPS 2024): DITS introduced truncated deterministic refinement in point embedding space; DAB elevates this concept into distribution parameter space, jointly updating semantic centers and uncertainty scales.
  • vs UATVR (Fang et al., ICCV 2023): UATVR relies on stochastic sampling of prototype vectors during similarity evaluation; DAB replaces sampling with closed-form directional KL divergence, stabilizing optimization.
  • vs DiffusionRet (Jin et al., ICCV 2023): DiffusionRet formulates retrieval via generative denoising of latent representations; DAB demonstrates that discriminative retrieval only requires deterministic drift without stochastic generation.

Rating

  • Novelty: โญโญโญโญโ˜† Pioneers deterministic, sampling-free diffusion bridges over Gaussian distribution parameters for cross-modal retrieval.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluations across three benchmarks, rigorous ablations, variance asymmetry audits, and selective prediction calibration.
  • Writing Quality: โญโญโญโญโญ Well-grounded motivation, mathematically transparent formulation, and seamless alignment between text, math, and diagrams.
  • Value: โญโญโญโญโ˜† Sets a compelling, reproducible paradigm for uncertainty-aware multimodal retrieval and distribution-level alignment.