Bridging Vision and Language Concepts through Optimal Transport Semantic Flow¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/ChenyangZhang00/OTF-CBM
Area: Multimodal VLM
Keywords: Concept Bottleneck Model, optimal transport, inverse optimal transport, flow matching, cross-modal alignment
TL;DR¶
Addressing fine-grained localization loss and cross-modal geometric distortion caused by fixed embeddings and global cosine similarity in traditional vision-language CBMs, OTF-CBM adapts cross-modal costs via inverse optimal transport, captures many-to-one matching and background suppression via unbalanced optimal transport, and evaluates instantaneous flow matching velocities at the midpoint for accurate concept activation without numerical ODE integration.
Background & Motivation¶
Concept Bottleneck Models (CBMs) promise interpretable and transparent decision-making by routing neural predictions through an intermediate bottleneck layer of human-understandable concepts, making complex models amenable to post-hoc diagnosis, error inspection, and targeted human intervention. With the emergence of foundation vision-language models such as CLIP alongside modern LLMs, recent CBM variants have automated concept discovery and supervision by projecting image features onto concept banks via cosine similarity in shared multimodal embedding spaces. However, existing vision-language CBMs suffer from three persistent representational bottlenecks: first, reusing a shared CLIP latent space tethers visual representations to linguistic geometry that fails to capture fine-grained perceptual granularity; second, relying on global pooled image embeddings blurs localized visual evidence and severely degrades spatial grounding; third, standard cosine similarity optimized for instance-level category discrimination provides an inadequate proxy for measuring transport costs between heterogeneous local patches and textual concepts.
The core tension behind these limitations is that establishing faithful concept bottlenecks requires region-to-concept correspondences that are simultaneously semantically accurate and geometrically consistent across modalities. Yet in real-world scenarios, component-level annotations are rarely available, and visual-textual alignments are fundamentally unbalanced: multiple visual patches frequently correspond to a single semantic concept (many-to-one), extensive background regions possess no semantic counterparts, and specific textual concepts are entirely absent from a given image. Directly enforcing balanced optimal transport with mass conservation or relying on uncalibrated Euclidean/cosine metrics inevitably forces irrelevant background patches to absorb textual concept mass, generating distorted couplings and hallucinated associations.
This paper's angle of attack is to reformulate concept alignment as a dynamic, unbalanced transport process rather than a static geometric projection: first learning a data-driven metric landscape between heterogeneous feature spaces via inverse optimal transport, and subsequently parameterizing the continuous semantic flow from visual prototypes to concept embeddings via flow matching. Core idea: learn data-driven cross-modal costs through inverse optimal transport, establish flexible many-to-one couplings with background suppression via unbalanced optimal transport, and infer concept activations directly from instantaneous midpoint velocity agreement under a trained flow field without ODE integration.
Method¶
Overall Architecture¶
The forward pipeline of OTF-CBM consists of four coordinated stages: region prototype aggregation, adaptive cross-modal cost learning via inverse optimal transport, geometric coupling via unbalanced optimal transport, and continuous concept activation via flow matching velocity alignment. Given an input image, patch embeddings extracted from a visual encoder are clustered into region prototypes using K-means. Next, an adaptive cost matrix between visual prototypes and fixed textual concept embeddings is constructed using a multi-basis metric learned through inverse optimal transport, combined with CLS-attention background penalties. Solving the unbalanced optimal transport problem produces a flexible coupling plan that naturally accommodates many-to-one alignments and background mass shrinkage. Finally, this coupling supervises a continuous conditional velocity field under flow matching. At inference time, the model bypasses ODE numerical integration entirely, computing concept activation scores directly from instantaneous velocity agreement at the trajectory midpoint \(t=0.5\) before feeding the aggregated activations into a linear classifier.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image Patch Prototypes<br/>and Textual Concept Set"] --> B["Inverse Optimal Transport Cross-Modal Cost<br/>Adaptive Multi-Basis Fitting"]
B --> C["Unbalanced Optimal Transport Geometric Coupling<br/>KL-Relaxed Many-to-One and Background Suppression"]
C --> D["Flow Matching and Midpoint Velocity Activation<br/>Continuous Velocity Field and t=0.5 Instantaneous Alignment"]
D --> E["Concept Activation Vector Aggregation<br/>Linear Classifier Class Prediction"]
Key Designs¶
1. Inverse Optimal Transport Cross-Modal Cost: Adaptive Metric Calibration
Visual encoders (e.g., DINOv2) and textual encoders (e.g., CLIP text transformer) generate representations residing in heterogeneous geometric spaces, rendering predefined ground costs like squared Euclidean distance or cosine similarity inherently misaligned and prone to severe geometric distortion. To infer an authentic distance landscape that mirrors cross-modal semantic affinities, OTF-CBM adopts Inverse Optimal Transport (IoT) to learn the underlying cost function from empirical component-level correspondences. A lightweight visual adapter \(f_\psi\) projects visual prototypes into a shared \(d_p\)-dimensional space, and the parameterized cost function is defined as a linear combination of \(K=18\) candidate kernel bases \(\Phi = \{\phi_k\}_{k=1}^K\):
The basis set encompasses squared angular distance, dot product similarity, Dot-RBF hybrid kernels, inverse quadratic metrics, and root exponential functions. This formulation captures both local correlation and global semantic separation while preserving linear interpretability. To avoid optimization instabilities associated with sparse Fenchel-Young objectives, the model optimizes an absolute-weighted discrepancy between the induced transport plan \(\pi_\theta\) and empirical supervisory coupling \(\hat{\pi}\):
Detaching gradients through \(\pi_\theta\) during early training ensures smooth convergence, yielding a calibrated and frozen metric landscape \(c_{\theta^*}\) that serves as a dependable geometric substrate for subsequent transport and flow modeling.
2. Unbalanced Optimal Transport Geometric Coupling: Flexible Mass Conservation Relaxation
Patch-level transport across dense image grids is computationally prohibitive and sensitive to spatial redundancy. The framework therefore clusters local patch embeddings into \(K\) region prototypes \(\tilde{x}_{1:K}\) via K-means. In cross-modal matching, standard balanced optimal transport enforces strict marginal constraints (\(\pi \mathbf{1} = \mu\) and \(\mathbf{1}^\top \pi = \nu\)), compelling irrelevant background regions and absent concepts to exchange artificial mass. OTF-CBM addresses this by formulating a Vision-Language Unbalanced Optimal Transport (VLOT) problem with generalized Kullback-Leibler (KL) divergence penalties:
The relaxed column marginal enables multiple visual prototypes to associate with a single concept (many-to-one alignment), while the relaxed row marginal permits surplus visual mass to shrink at finite cost. Furthermore, to eliminate residual background interference, the model extracts the visual backbone's CLS self-attention map as a foreground prior. Prototypes exhibiting low attention are indexed into a background set \(B\), and their transport costs are explicitly penalized:
Because background mass transfer becomes systematically more expensive, the unbalanced Sinkhorn solver opts to diminish background mass rather than forcing spurious matches to semantic concepts, achieving clean spatial disentanglement.
3. Flow Matching and Midpoint Velocity Activation: ODE-Free Dynamic Inference
While the discrete transport plan \(\pi_{\mathrm{VLOT}}\) determines static assignments, it fails to capture continuous semantic transitions and cannot evaluate out-of-coupling feature instances. OTF-CBM lifts discrete transport displacements into a continuous vector field using conditional Flow Matching (FM). Given endpoint pairs \((x_0, x_1) \sim \pi_{\mathrm{VLOT}}\) where \(x_0 = f_{\psi^*}(\tilde{x}_k)\) and \(x_1 = c_j\), straight-line trajectories are defined as \(x_t = (1-t)x_0 + t x_1\) with target velocity \(u_t = x_1 - x_0\). The conditional velocity field \(v_\phi(x, t, \mathrm{cond})\) is trained via regression:
In generative modeling, sampling requires multi-step numerical ODE integration. For Concept Bottleneck Models, however, the goal is evaluating concept compatibility rather than synthesizing text tokens. By leveraging the theoretical properties of Brownian bridges and Schrödinger bridges, the interpolating process \(X_t = (1-t)x_0 + tx_1 + B_t\) exhibits conditional variance \(\sigma_t^2 = t(1-t)\), which is uniquely maximized at the midpoint \(t = \frac{1}{2}\). The midpoint represents the region of maximal semantic uncertainty and highest discriminative information along the transport path. Consequently, OTF-CBM evaluates instantaneous velocity consistency directly at \(t = 0.5\) without numerical integration:
where \(x_{1/2}^{(k,j)} = \frac{1}{2}(x_0 + x_{1,j})\) and \(u^{(k,j)} = x_{1,j} - x_0\). Concept activation scores are computed by aggregating top prototype responses \(a_j = \frac{1}{|TopK|} \sum_{k \in \mathrm{TopK}(S_{\cdot, j})} S_{k,j}\), normalized via LayerNorm, and passed to a linear classifier to produce class predictions \(\hat{y} = f_{\mathrm{cls}}(\mathrm{LayerNorm}(a))\).
Key Experimental Results¶
Main Results¶
The authors evaluated OTF-CBM on five diverse classification benchmarks: ImageNet-1K, CUB-200-2011, CIFAR-100, Animals with Attributes 2 (AwA2), and Places365. All baseline models were re-implemented using identical pre-trained ViT backbones and standardized Label-Free CBM concept banks to eliminate architectural confounders.
| Dataset | Metric | Ours (OTF-CBM) | Prev. SOTA (DOT-CBM) | Gain |
|---|---|---|---|---|
| ImageNet-1K | Top-1 Accuracy (%) | 85.62 | 83.84 | +1.78% |
| CUB-200-2011 | Top-1 Accuracy (%) | 89.92 | 85.39 | +4.53% |
| CIFAR-100 | Top-1 Accuracy (%) | 90.21 | 85.83 | +4.38% |
| AwA2 | Top-1 Accuracy (%) | 98.88 | 96.83 | +2.05% |
| Places365 | Top-1 Accuracy (%) | 55.13 | 50.65 | +4.48% |
Across all benchmarks, OTF-CBM consistently outperforms previous state-of-the-art CBM architectures, demonstrating substantial improvements on fine-grained datasets (+4.53% on CUB) and complex scene datasets (+4.48% on Places365).
To quantitatively benchmark the fidelity of the continuous cross-modal semantic flows, the paper evaluates four dynamic metrics: Transport Reconstruction Error (TRE), Velocity Mean Squared Error (VMSE), Mean Cosine Ratio (MCR), and Negative Pair Error (NPE).
| Semantic Flow Configuration | Modeling Paradigm | VMSE (↓) | MCR (↑) | NPE (↓) |
|---|---|---|---|---|
| Static OT Alignment | Static matching baseline (no flow) | 2.467 | 0.739 | 0.434 |
| OT Flow Matching | Linear flow via balanced transport | 2.001 | 0.803 | 0.299 |
| IoT + OT Flow Matching | Learned cost + balanced flow | 1.782 | 0.884 | 0.197 |
| IoT + UOT Flow Matching (Ours) | Learned cost + unbalanced many-to-one flow | 1.000 | 0.999 | 0.021 |
Ablation Study¶
The ablation study systematically analyzes the cumulative impact of each core component on classification accuracy across the five benchmark datasets.
| Config | ImageNet | CUB | CIFAR-100 | AwA2 | Places365 | Note |
|---|---|---|---|---|---|---|
| Vanilla-CBM | 79.17 | 78.32 | 80.04 | 93.15 | 44.80 | Baseline cosine bottleneck |
| + Classical OT | 80.42 | 80.31 | 85.23 | 94.38 | 47.70 | Explicit discrete coupling |
| + Vision-Language UOT (VLOT) | 82.68 | 84.31 | 87.44 | 95.61 | 50.59 | Marginal relaxation & background suppression |
| + Inverse OT Cost (IoT) | 84.87 | 87.44 | 89.09 | 96.56 | 52.99 | Data-driven metric landscape |
| + Semantic Flow (Full Model) | 85.62 | 89.92 | 90.21 | 98.88 | 55.13 | Midpoint velocity flow activation |
Regarding metric fitting fidelity, standard squared 2-Wasserstein distance (\(W_2^2\)) yields a TRE of 2.371, and cosine distance achieves 1.826. In contrast, the multi-basis IoT cost reduces TRE to 0.824, confirming the effectiveness of adaptive metric learning.
Under out-of-distribution (OOD) testing—where foreground objects are segmented via SAM and backgrounds are perturbed with randomized recolored noise—OTF-CBM attains an OOD accuracy of 82.0% on CUB (+14.5% over DOT-CBM's 67.5%) and 50.5% on Places365 (+8.4% over DOT-CBM's 42.1%), demonstrating robust resilience against contextual shortcuts.
Key Findings¶
- Data-driven metric learning via IoT is fundamental to cross-modal transport fidelity: combining angular distance, dot products, and multi-scale Dot-RBF kernels enables the model to reduce Transport Reconstruction Error (TRE) by 54.8% compared to cosine distance.
- Unbalanced OT with CLS-guided geometric suppression eliminates background hallucination: dropping the Negative Pair Error (NPE) from 0.434 to 0.021 and driving the Mean Cosine Ratio (MCR) to 0.999.
- Evaluating instantaneous flow velocities at the trajectory midpoint provides an optimal balance between expressiveness and efficiency: grounded in the maximum variance property of Schrödinger bridges, this single-step readout avoids costly ODE solver iterations while preserving continuous dynamical geometry.
Highlights & Insights¶
- From Static Projections to Dynamical Velocity Fields: Rather than treating concept prediction as static inner products in a frozen latent space, the paper re-envisions concept inference as a continuous flow process, measuring directional alignment in vector fields.
- ODE-Free Inference via Schrödinger Bridge Theory: By exploiting the maximal uncertainty property of Brownian bridge variance at \(t=0.5\), the model turns an otherwise slow generative flow process into a lightning-fast single-forward activation metric.
- A Decoupled Perspective on Multimodal Representation: The authors highlight that joint end-to-end representation learning often forces incompatible cross-modal geometries prematurely; learning structured unimodal embeddings first and bridging them with adaptive transport geometry offers a compelling alternative paradigm.
Limitations & Future Work¶
- Reliance on Component-Level Annotation Priors: The inverse optimal transport cost functional requires initial fitting on datasets equipped with component-level associations, which may restrict deployment in zero-shot or extreme long-tail regimes lacking prior pairings.
- Fixed Cluster Count in Prototype Extraction: The K-means clustering budget \(K\) remains a fixed hyperparameter, which may under-represent densely packed visual scenes or over-partition sparse images.
- Future Directions: Developing self-supervised inverse optimal transport algorithms without component labels and extending the midpoint velocity flow mechanism to dense spatial reasoning tasks (e.g., open-vocabulary segmentation and embodied navigation).
Related Work & Insights¶
- vs Vanilla-CBM / Post-hoc CBM (P-CBM): Prior CBMs map global pooled feature vectors directly to concept scores via cosine similarity, erasing local spatial evidence; OTF-CBM establishes spatially grounded, region-to-concept unbalanced transport plans.
- vs DOT-CBM: While DOT-CBM explores discrete optimal transport for concept discovery, it assumes fixed Euclidean/cosine metrics and static bipartite assignment; OTF-CBM learns data-driven non-Euclidean cost metrics via IoT and lifts discrete matching into a continuous flow matching dynamical system.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering integration of inverse optimal transport and flow matching velocity fields for concept bottleneck models, accompanied by an elegant ODE-free midpoint inference mechanism.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation across five major benchmarks, dynamic flow metric evaluations, extensive ablations, and SAM-based OOD robustness tests.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulations, compelling motivation, and clear structural narrative.
- Value: ⭐⭐⭐⭐⭐ Provides profound theoretical and practical insights for multimodal interpretability, geometric deep learning, and continuous concept reasoning.