Spherical Interpolation for Backward-Compatible Multimodal Representations¶
Conference: NeurIPS2026 (task-list grouping; reading version: arXiv v1, 2026-09-30)
arXiv: 2609.39836
Code: https://github.com/miccunifi/SLERP_backward_compatibility
Area: Multimodal VLM
Keywords: backward compatibility, cross-modal retrieval, orthogonal Procrustes, spherical interpolation, model upgrades
TL;DR¶
Without rebuilding the old gallery index, Procrustes first aligns new-model queries to the old space, then spherical interpolation combines them with old-model queries for the same input; endpoint complementarity accounts for most average retrieval gains, while support-set weight selection further improves compatibility success rates.
Background & Motivation¶
Vision-language models (VLMs) such as CLIP and SigLIP encode images and text as normalized vectors and retrieve items using cosine similarity. Upgrading the encoder is not as simple as replacing a classification head: the gallery may already contain millions or billions of old vectors, while an independently trained new model uses different representation coordinates. Searching the old index with new queries directly may disrupt existing matches rather than exploit stronger semantics. Re-encoding may even be impossible when raw gallery data have been deleted or are subject to privacy restrictions.
Backward-compatible training incorporates old-space constraints into new-model training, requiring a dedicated pipeline and potentially limiting the new model. Post-hoc orthogonal alignment offers a lightweight alternative: shared inputs estimate a new-to-old coordinate transformation without changing either encoder. However, a global transformation cannot remove every fine-grained difference between models. The paper observes residual angular discrepancies on both CC3M and zero-shot transfer to Flickr30k, showing that sharing coordinates does not make the two embeddings of an input identical.
The old query retains its matching behavior against the deployed gallery, while the aligned new query may supply discriminative information missing from the old representation. The residual need not be entirely noise, but moving all the way to the new endpoint is not necessarily beneficial. Core idea: treat the old query and aligned new query as endpoints on a shared sphere, select a transferable interpolation position between them, and exploit query-side complementarity for compatible upgrades without modifying the old gallery.
Method¶
Overall Architecture¶
Inputs are frozen old and new VLMs, a support set, and the deployed old-model gallery; the output is a single unit query vector that still searches the old index. โProcrustes Alignmentโ and โSupport-Set Weight Selectionโ run offline; โSpherical Query Fusionโ runs online and requires both models' embeddings of the same query. Offline fitting is not encoder training, and dashed edges in the diagram carry fitted results rather than back-propagation signals.
The gallery retains its old embeddings throughout: I2T searches an old text gallery with an image query, and T2I searches an old image gallery with a text query. Weights are selected separately for the two directions. Given a model pair, support modality, and retrieval direction, the weight is fixed across all target datasets and queries, rather than chosen by a per-query oracle.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Offline support set<br/>Frozen old and new encoders"] --> B["Procrustes Alignment"]
B --> C["Support-Set Weight Selection"]
Q["Online query"] --> O["Old-query encoding"]
Q --> N["New-query encoding<br/>Align and normalize"]
B -.->|Fixed map| N
O --> D["Spherical Query Fusion"]
N --> D
C -.->|Fixed weight| D
D --> S["Cosine retrieval"]
G["Old gallery and index<br/>No re-encoding"] --> S
Key Designs¶
1. Procrustes Alignment: establish shared coordinates before addressing residual differences
Each support input is encoded by both models, with old and new vectors stacked row-wise into \(U\) and \(\bar V\). In the equal-dimensional case, the objective is \(R^\star=\arg\min_{R^\top R=I}\|\bar V R-U\|_F^2\); if \(\bar V^\top U=P\Sigma Q^\top\), its closed-form solution is \(R^\star=PQ^\top\). The experiments call this alignment baseline SVD because it requires a singular value decomposition, not gradient-based optimization.
An equal-dimensional orthogonal map preserves the new model's internal inner products and angles, avoiding arbitrary rescaling or shearing merely to approach old support embeddings. Support can consist of text only, images only, or stacked embeddings from both modalities. Shared cross-modal structure allows a map estimated from one modality to transfer to the other, but independently trained models need not satisfy this correspondence exactly. The method exploits the remaining discrepancy rather than assuming that SVD resolves compatibility completely.
Actual model dimensions differ: CLIP ViT-B/32, L/14, and H/14 use 512, 768, and 1024 dimensions, respectively, while SigLIP1/2 use 1152. Appendix B distinguishes the directions of a rectangular map: a lower-dimensional new representation can be embedded isometrically into the old space, whereas a higher-dimensional new representation retains only the subspace aligned with the old space and cannot preserve all new-space inner products. The projected new query is therefore normalized as \(v=\phi_{\mathrm{new}}(x)R^\star/\|\phi_{\mathrm{new}}(x)R^\star\|_2\), requiring a nonzero denominator. Not every cross-dimensional upgrade is a full-space isometry.
2. Support-Set Weight Selection: calibrate deployment positions without test labels
The authors use 12,637 available imageโtext pairs from CC3M validation both to estimate the alignment map and to select interpolation weights. For each model pair and support modality, I2T and T2I separately maximize support-set Recall@1 over \(\alpha\in\{0,0.1,\ldots,1.0\}\). The resulting \(\hat\alpha\) is frozen for direct evaluation on Flickr30k, COCO2014, and NoCaps. Support is separate from target queries and galleries, but map fitting and weight selection share the same support set rather than using an additional independent calibration split.
The fixed weight addresses the unobservable best position at deployment. It does not access the current query's relevance labels or search the target test set dynamically. Dataset-level oracle \(\alpha^\star\) and per-query oracle results are reported separately as upper-bound analyses. Appendix K selects weights from a small labeled sample of the target distribution in a different setting; that analysis must not be conflated with the main protocol's absence of target-test tuning.
Including both endpoints in the support grid permits fallback to the old query when the new endpoint is unreliable. This provides a support-set fallback, not a guarantee of outperforming the old model on a target distribution. If \(\hat\alpha=0\), the system equals the old system and does not satisfy the paper's strict-improvement compatibility criterion.
3. Spherical Query Fusion: retain old-query matching while incorporating aligned new information
Online inference computes the old query \(u=\phi_{\mathrm{old}}(x)\) and normalized aligned new query \(v\), both on the old-space unit sphere. With \(\theta=\arccos\langle u,v\rangle\), the fused vector is \(q_\alpha=\frac{\sin((1-\alpha)\theta)}{\sin\theta}u+\frac{\sin(\alpha\theta)}{\sin\theta}v\). This is spherical linear interpolation (SLERP): \(\alpha=0\) gives the old query, \(\alpha=1\) gives the SVD query, and \(\angle(u,q_\alpha)=\alpha\theta\) gives the weight a consistent interpretation as a fraction of the total endpoint angle.
The formula and theory require \(0<\theta<\pi\). Coincident endpoints offer no useful path; antipodal endpoints have no unique minor geodesic and cannot use the zero-denominator formula directly. Near-coincident and near-antipodal cases also require numerical safeguards, but the nondegenerate theorem does not automatically cover every implementation fallback. The output remains one old-dimensional unit vector and requires one old-index search. Avoiding gallery re-encoding does not mean that old-query encoding can be retired.
Why can an interior point outperform both endpoints? Let \(q^\ast\) be an idealized retrieval direction used only for analysis. Its projection onto the plane spanned by \(u,v\) has magnitude \(\rho\) and angular coordinate \(\psi\) relative to the old endpoint. Theorem 1 gives the angular gap \(G(\alpha)=\arccos(\rho\cos(\alpha\theta-\psi))\): an interior direction strictly better than both endpoints exists exactly when the nonzero normalized projection lies in the relative interior of the minor arc. A zero projection gives angle \(\pi/2\) along the entire arc; if the projected direction lies outside the arc, an endpoint is optimal. Large endpoint separation alone does not imply useful interpolation, and interpolation cannot recover all information outside the plane.
Relating angles to Recall requires a retrieval margin. Appendix F defines \(m_{K,x}(q)=\max_{j\in P(x)}\langle q,g_j\rangle-\tau^-_{K,x}(q)\), where \(P(x)\) is the relevant-item set and \(\tau^-_{K,x}\) is the \(K\)-th largest negative score. A positive margin guarantees that at least one relevant item enters the top-\(K\); a zero margin depends on tie-breaking. Let \(q_K^\star(x)\) maximize this margin. Only when its optimal margin \(m^\star_{K,x}>0\) and \(G_{K,x}(\alpha)<2\arcsin(m^\star_{K,x}/4)\) does the local sufficient condition certify single-query success. It follows from the perturbation bound \(|m_{K,x}(q)-m_{K,x}(q')|\leq2\|q-q'\|_2=4\sin(\angle(q,q')/2)\). This is neither compatibility for all queries or datasets nor a deployment algorithm for finding the ideal direction.
SLERP's advantage should be interpreted narrowly. Normalized linear interpolation (NLERP) traverses the same minor arc with a different angular speed; continuous search for an individual query can reach the same optimal directions. With a shared coefficient and finite grid, parameterization changes the positions sampled for different queries, but Appendix N finds almost identical average results. At a fixed coefficient, NLERP and linear interpolation of old/new scores also produce identical rankings because normalization multiplies all scores for that query by the same positive constant. The evidence supports endpoint complementarity, not a performance gain unique to the spherical operator.
A Worked Example¶
Consider upgrading CLIP ViT-B/32 with ViT-L/14 using text-support alignment for I2T. The gallery retains text vectors encoded by ViT-B/32. Offline fitting estimates a 768-to-512-dimensional map on CC3M and selects image-query weight 0.7. Online, an incoming image passes through both image encoders; the new vector is projected and normalized, then fusion moves 70% of the angular distance along the minor arc. The resulting 512-dimensional vector searches the original text gallery.
Under the full Flickr30k protocol, the fixed procedure reaches 48.00% Recall@1, compared with 40.62% for old queries and 42.89% for aligned new queries alone. Weight 0.7 is shared by all images, not chosen after observing the correct caption. The dataset oracle's weight 0.5 and 48.61% are separate upper-bound results. This example distinguishes dual query encoding from zero gallery re-encoding.
Loss & Training¶
The method introduces no encoder loss, fine-tunes neither CLIP nor SigLIP, and learns no gradient-trained compatibility layer. Fitted quantities are only the closed-form Procrustes map and the support-grid weight. โTraining-freeโ means no model retraining or gradient optimization, not an absence of support data or offline computation.
The main retrieval benchmarks use the union of all Flickr30k Karpathy splits and the COCO2014 and NoCaps validation splits. They are not directly comparable with common Flickr30k 1K or COCO 5K protocols. I2T/T2I Recall@1/5/10 measure the fraction of queries retrieving at least one relevant item.
Key Experimental Results¶
Main Results¶
The table selects text-support T results from Tables 1โ2. All scores are Recall@1 percentages, and weights are ordered I2T/T2I. Upgrade arrows run from the new model to the old gallery model. SLERP uses fixed CC3M weights without test-set tuning.
| Model upgrade and dataset | Old I2T/T2I | SVD I2T/T2I | SLERP I2T/T2I | Fixed weight |
|---|---|---|---|---|
| CLIP L/14 โ B/32, Flickr30k | 40.62 / 21.73 | 42.89 / 21.49 | 48.00 / 23.33 | 0.7 / 0.5 |
| CLIP L/14 โ B/32, COCO | 28.76 / 14.47 | 30.27 / 14.03 | 33.40 / 15.26 | 0.7 / 0.5 |
| SigLIP2 โ B/32, Flickr30k | 40.62 / 21.73 | 42.80 / 23.56 | 51.01 / 25.79 | 0.6 / 0.5 |
| SigLIP2 โ B/32, NoCaps | 71.29 / 45.24 | 72.16 / 46.41 | 78.56 / 50.59 | 0.6 / 0.5 |
| SigLIP2 โ SigLIP1, COCO | 46.99 / 30.88 | 36.30 / 30.53 | 46.76 / 32.50 | 0.3 / 0.4 |
| SigLIP1 โ H/14, Flickr30k | 59.38 / 43.07 | 60.50 / 29.91 | 64.81 / 43.38 | 0.5 / 0.1 |
For same-family L/14 โ B/32, Flickr30k I2T rises from 42.89% with SVD to 48.00%, a gain of 5.11 percentage points. However, the corresponding COCO T2I score of 15.26% remains below XBT's 15.55%, so the method does not uniformly surpass training-based baselines. SigLIP2 โ SigLIP1 COCO I2T remains 0.23 percentage points below the old model, distinguishing substantial recovery from strict compatibility.
Empirical compatibility requires \(M_{\mathrm{new}\to\mathrm{old}}>M_{\mathrm{old}\to\mathrm{old}}\), comparing average metrics on the same old gallery rather than requiring every query to avoid regression. The positive-flip rate PFR is the fraction of all queries that fail with the old system and succeed with the updated system; NFR is the fraction that succeed before and fail afterward. Thus \(\Delta\mathrm{R@}K=\mathrm{PFR}-\mathrm{NFR}\), and higher average Recall can coexist with negative flips.
Ablation Study¶
Table 10 aggregates five model pairs, three support modalities, and three datasets: 45 configurations and 90 direction-level compatibility evaluations. The selected comparisons below retain the source table's mean gains over SVD and compatibility counts.
| Config | Mean I2T R@1 | Mean T2I R@1 | Mean gain over SVD (percentage points) | Compatible evaluations |
|---|---|---|---|---|
| SVD / Procrustes | 51.72 | 33.77 | 0.00 | 45/90 |
| Affine alignment | 31.69 | 23.10 | -15.35 | 11/90 |
| Ridge alignment | 33.96 | 28.89 | -11.32 | 14/90 |
| Whitened Procrustes | 49.06 | 32.78 | -1.83 | 37/90 |
| SVD + fixed normalized midpoint | 57.76 | 37.23 | +4.75 | 72/90 |
| SVD + support-selected NLERP | 57.73 | 37.38 | +4.81 | 84/90 |
| SVD + support-selected SLERP | 57.72 | 37.38 | +4.80 | 85/90 |
| SVD + adaptive NLERP | 57.86 | 37.42 | +4.89 | 88/90 |
The fixed midpoint already provides +4.75 percentage points, while SLERP's +4.80 adds only 0.05. Weight selection primarily improves compatibility robustness rather than producing a large additional mean gain. Adaptive NLERP selects positions from endpoint cosine agreement, with parameters calibrated only on CC3M. Its additional 0.09 percentage points over SLERP belong to an appendix extension, not the main method's per-query oracle.
The source contains a count conflict: Section 4.3 and Table 10 report 85/90 for SLERP, whereas Appendix N.3 states 84/90 and discusses a single-query floating-point tie at an old-endpoint configuration in a footnote. This note retains the table's 85/90 without silently reconciling the two counts. Exact compatibility totals require verification against the authors' code with consistent tie-breaking.
Key Findings¶
Appendix M's per-query oracle uses Flickr30k I2T, L/14 โ B/32, text support, and 11 weights, accessing individual test-query labels and therefore remaining non-deployable. The results distinguish useful interior directions from weights that can be selected online.
| Query-weight rule | Recall@1 (%) | Deployability and interpretation |
|---|---|---|
| CC3M fixed weight 0.7 | 48.00 | Main method; no target-test labels |
| Flickr30k dataset oracle 0.5 | 48.61 | Test-dataset-level upper bound |
| Per-query choice of two endpoints | 54.61 | Non-deployable oracle; 16,936 successful queries |
| Per-query search over the discrete arc | 59.15 | Non-deployable oracle; 18,344 successful queries |
- There are 1,408 interior-only successful queries, contributing 4.54 percentage points unavailable to the endpoint oracle. This is more direct evidence of interior-direction utility than the assignment of interior weights to 97.65% of achievable queries: the latter depends on relevant-item similarity and tie-breaking and is not the fraction requiring interpolation for success.
- For 99.87% of queries, similarity to at least one relevant caption is strictly higher at an interior grid point than at either endpoint, but only some of these changes improve rankings. Negative scores also change; relevant-item similarity alone does not establish Recall success.
- Auxiliary zero-shot classification keeps old text prototypes fixed and separately fits and selects weights on ImageNet validation. In Table 3, L/14 โ B/32 mean accuracy across ten datasets rises from SVD's 51.89% to 63.57%, exceeding the old model's 61.99%. SigLIP2 โ SigLIP1 reaches 79.95%, below the old model's 80.71%, so compatibility is not always recovered.
- In Appendix P's A100 IVF-PQ test with a \(10^7\)-item synthetic gallery, one-GPU SLERP I2T latency for L/14 โ B/32 is 31.04 ms versus 25.59 ms after full re-indexing; parallel execution on two GPUs reduces it to 25.87 ms. Gallery re-encoding is avoided, but dual-query encoding is not. Latency sums separately measured component p50 values, rather than measuring full-service end-to-end p50.
Highlights & Insights¶
- Alignment error is treated as a potentially complementary direction rather than only a bias to eliminate. Shared coordinates retain old-model matching stability while allowing the new model to contribute semantics without rebuilding the gallery.
- The normalized midpoint isolates fusion itself, while support-selected weights isolate position calibration. This ablation attributes gains to endpoint complementarity rather than presenting the nearly identical benefits of NLERP and score fusion as unique to SLERP.
- Retrieval margins remain the bridge between geometric theory and discrete rankings. Moving closer to a reference direction is not automatically successful; sufficient margin is needed for local certification, explaining why mean gains and negative flips can coexist.
Limitations & Future Work¶
- The old encoder must remain available online unless usable old-query embeddings are cached. When the gallery cannot be rebuilt, this can become a long-term cost. Query throughput, dual-model memory, and cache hit rates matter alongside avoiding index reconstruction.
- Results depend on support coverage and endpoint complementarity. Interpolation cannot create missing domain information when both endpoints are weak, and support-set fallback to the old endpoint does not ensure non-degradation in the target domain.
- The theory uses an unobserved ideal direction, and certification is a single-query local sufficient condition with a positive-margin assumption. A fixed weight does not guarantee regression-free queries or compatibility across an entire distribution.
- The v1 compatibility count conflicts between 85/90 and 84/90 as noted above. Appendix J also explains the jointly transformed query/gallery SVD endpoint as equivalent to the original new model, but preservation of all inner products directly holds only for isometric cases. It cannot be applied unconditionally to Appendix B's high-to-low-dimensional projection; the cross-dimensional re-indexing protocol warrants separate verification.
- Main results are point estimates without confidence intervals for retrieval gains; target-subset sensitivity uses only three random seeds. The checklist also acknowledges that the current code is a minimal implementation, with complete run commands and baseline reproduction scripts not yet supplied.
- Position prediction from endpoint agreement, query difficulty, and local negative-item margins is promising but requires independent calibration data. Label-based oracles cannot become deployment strategies. Serving old and new galleries during partial backfilling is another migration question.
Related Work & Insights¶
- vs BCT / XBT: Compatibility training constrains the difference between new and old representations; XBT additionally uses projection and LoRA-based training. This method freezes models and trades dual-query encoding for the absence of compatibility retraining, fitting upgrades with independent pretrained models and temporarily immutable galleries.
- vs Canonicalizing Multimodal Contrastive Representation Learning: A shared orthogonal transformation explains possible cross-modal alignment transfer. This paper investigates angular residuals after that global transformation: alignment establishes coordinates, while fusion addresses remaining discrepancies.
- vs model soups / parameter-space SLERP: Parameter fusion usually requires compatible architectures and weights. This method fuses query representations of the same input rather than model parameters, allowing cross-family combinations but retaining both encoders.
- vs NLERP / score fusion: For an individual query, these methods and SLERP share the reachable minor arc; NLERP and score fusion also have identical rankings at the same coefficient. The transferable lesson is to test whether combining endpoints helps before introducing more complicated weighting rules.
Rating¶
- Novelty: 4/5. Connects post-hoc alignment residuals, spherical paths, and retrieval margins, though the fusion operator itself is established.
- Experimental Thoroughness: 4/5. Covers model families, support modalities, retrieval directions, alternative operators, and costs; count discrepancies and reproduction details remain unresolved.
- Writing Quality: 4/5. Clearly separates the main procedure from oracles, but cross-dimensional re-indexing geometry and compatibility counts are inconsistent.
- Value: 4/5. Practical for VLM retrieval systems that cannot immediately rebuild a gallery, provided gains are weighed against continuing dual-query encoding costs.