Skip to content

XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication

Conference: NeurIPS2026
arXiv: 2608.11676
Code: https://github.com/WooseongYang/XBridge
Area: Multi-Agent / LLM Efficiency
Keywords: heterogeneous model communication, entity grounding, lexical anchor mapping, cross-attention, latent representation transfer

TL;DR

XBridge transfers the full context token sequence through deterministic lexical anchor mapping and reads sender hidden states through a pairwise-trained cross-attention bridge, outperforming 128-token text-summary communication on all seven tasks for three heterogeneous model pairs and reducing per-sample latency from 1.70 to 0.15 seconds in the specified H200 test.

Background & Motivation

Different model families have different training data, tokenizers, and representation spaces, potentially offering complementary capabilities to multi-agent systems, but arbitrary hidden states or KV caches cannot simply be shared between them. Text summaries provide a universal interface, yet require the sender to generate a message token by token and compress a full document into a bounded string. The paper's NLComm baseline lets the sender see the document but not the question and greedily generate a 128-token summary; this setting can omit the exact names, numbers, or relations that the receiver subsequently needs.

Projecting hidden states to the receiver's dimensionality does not automatically solve the problem. Equal dimensionality does not imply matching representation distributions, and continuous signals may retain relational semantics without reliably distinguishing the rare names involved. The authors call this rare-token compression collapse: “compression” here refers to an information bottleneck in continuous bridging, not a demonstrated reduction in network payload bytes. A latent-only bridge reaches 30.3% HotpotQA F1, well below the dual-channel model's 78.8%, illustrating that exact lexical identity and contextual reasoning are different requirements.

The paper therefore stops asking one continuous channel to preserve identity and transfer knowledge simultaneously. It maps the original context tokens into the receiver's own vocabulary, supplying recognizable lexical anchors before letting the receiver retrieve the sender's deeper representations according to its current processing state. Core idea: fix entity identity through a discrete full-context channel and supplement contextual information through a receiver-driven latent bridge, rather than making continuous representations alone recover exact entities across vocabularies.

Method

Overall Architecture

The task is a single asymmetric exchange: the sender receives a document, the receiver initially receives only a question, and the receiver ultimately generates the answer. One sender prefill forward pass produces the original document tokens and last-layer hidden states; XBridge transfers both, rather than first generating a summary.

Lexical Anchor Mapping (LAM) converts document tokens into receiver-vocabulary tokens and places their embeddings, obtained from the receiver's frozen embedding matrix, before the question. The Latent Enrichment Bridge (LEB) performs cross-attention at several receiver layers, using receiver states as queries and sender hidden states as keys and values. The channels meet within receiver computation: LAM shapes the entity information in the queries, while LEB supplies continuous contextual signals about those entities, without a separate explicit fusion network.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    C["Sender document"] --> S["Single sender prefill"]
    S -->|Original context tokens| L["Lexical Anchor Mapping LAM"]
    S -->|Last-layer hidden states| B["Latent Enrichment Bridge LEB"]
    L -->|Receiver-native embeddings| R["Receiver autoregressive generation"]
    Q["Receiver question"] --> R
    R -->|Receiver layer states as queries| B
    B -->|Gated residual| R
    R --> A["Answer"]
    T["Training: gold answer tokens"] -.->|Answer prediction supervision; update LEB only| B

Both branches originate from the same sender prefill, but LEB does not independently produce a fixed message in advance: it receives queries during receiver prefill and answer generation. The dashed edge denotes training supervision, not access to gold answers at inference time.

Here, decode-free means only that the sender does not autoregressively generate a communication message. LAM's fallback includes string decoding and re-tokenization, and the receiver still generates answers autoregressively; the protocol therefore does not eliminate all decoding or remove the receiver's need to process a long context.

Key Designs

1. Lexical Anchor Mapping LAM: place entity identity in receiver-native inputs

LAM processes all original document tokens, rather than running named entity recognition and selecting only names, or training a summary encoder. Each tokenizer pair has a precomputed deterministic mapping: a token whose surface string occurs in both vocabularies is converted by ID lookup; otherwise, that token is decoded into a string and split into one or more tokens by the receiver tokenizer, expanded in place. In an appendix example, Llama's 199 maps to Qwen's 1, 9, and 9, illustrating one-to-many remapping rather than regenerated digits.

The mapped IDs index the receiver's own frozen embedding matrix, and the full mapped context precedes the question. Receiver self-attention can thus reuse its pretrained handling of native vocabulary instead of treating heterogeneous hidden vectors as token embeddings. The channel also retains extensive non-entity text because relations and entity positions affect receiver queries; “entity anchors” describe its role, not an entity-only sparse message format.

On 100 HotpotQA validation samples, the reported token-level direct mapping rates are 97.1% for Llama→Qwen and 87.8% for Mistral→Qwen, leaving 2.9% and 12.2% for string fallback. These rates are consistent with vocabulary overlap ratios of 85.4% and 32.2%: overlap counts vocabulary entries, whereas mapping rates count tokens actually encountered in the samples.

The paper calls the mapping “lossless”; its concrete description supports preservation of surface strings in the presented examples, with possible sequence expansion. Tokenwise fallback is not a single global re-tokenization of the complete document, so identical token boundaries cannot be guaranteed from that description. The cache also provides no proof covering every Unicode case, byte fragment, special token, or normalization combination. The claim is best understood as the authors' lexical-content preservation claim, not an unconditional losslessness theorem for arbitrary tokenizer pairs.

2. Latent Enrichment Bridge LEB: let the receiver select the sender context it needs

Even after receiving all lexical content, the receiver lacks the sender model's deep processing of the document. LEB reads the sender's complete last-layer hidden-state sequence, forms queries from separately normalized receiver states, and projects sender states into keys and values. The query projection operates within the receiver dimensionality, while key and value projections accommodate the sender–receiver dimension mismatch. Each insertion layer has independent parameters but retrieves from the same sender states.

Cross-attention does not declare heterogeneous vectors compatible; it lets the receiver's current context determine which sender positions to retrieve. Through self-attention, LAM has already incorporated document entities and their context into receiver states, so LEB queries no longer contain only a question detached from the document. With LEB alone, the model may recover a broad relation without resolving lexical identity; with LAM alone, it has text but no additional sender representations. This also explains why adding LEB to LAM yields 22.3 percentage points, whereas LEB alone gains only 5.7 points over NoComm.

In the default 28-layer Qwen receiver configuration, four modules are inserted after layers 7, 14, 21, and 28. Each has approximately 66M parameters, totaling roughly 264M, or 3.8% of that receiver's size. These layer indices and the parameter ratio belong to that configuration and should not be directly assigned to the Llama receiver in reverse communication. Dimensions are adapted and a separate bridge is trained for each model pair; the paper does not establish that one bridge directly serves all models.

At each layer, the attention output enters the receiver state through a gated residual. The paper's core update is:

\[ h_R^{\prime(\ell)}=h_R^{(\ell)}+\tanh(\alpha^{(\ell)})\cdot A^{(\ell)}. \]

Here, \(A^{(\ell)}\) is the cross-attention output retrieved from the sender at that layer. The gate parameter starts at \(\alpha^{(\ell)}=1.0\), giving an initial coefficient of \(\tanh(1.0)\approx0.76\), not zero-initialized identity preservation. The residual retains the original computation path, but adding a nonzero signal still changes hidden representations; querying with receiver-native states does not mean leaving receiver representations undisturbed.

A Worked Example

An entity substitution example in the appendix asks which port city lies approximately 25 km north of the Lingnan Fine Arts Museum. The original document supports Keelung, and the intervention replaces the answer entity with Majuro. The example illustrates information flow rather than claiming that every sample is answered correctly.

The sender first processes the document, encoding the city name and surrounding relations in the original tokens and last-layer hidden states. LAM maps the entire document vocabulary, including the city name, to the receiver. Processing those tokens with the question gives the receiver queries that reflect the available entities and their context. LEB retrieves continuous sender states according to those queries, and the receiver then generates the city name.

In this example, original inputs to both channels produce Keelung; replacing only LEB's document states still produces Keelung; replacing only LAM's input changes the answer to Majuro. This demonstrates strong control of output entity identity by the explicit lexical channel in this intervention, but does not prove that LEB never affects entity selection or that relations and answers are always correct. In the aggregate experiment on 100 samples, 55–59 responses fall into the “neither original nor substituted entity” category, preventing generalization from the three successful examples to every sample.

Loss & Training

Both sender and receiver backbones are frozen; only bridge modules, including projections, normalizations, and gate parameters, are updated. Training uses next-token prediction supervision on gold answer tokens. LAM remains deterministic rather than learning to generate summaries. The stated objective is:

\[ \mathcal{L}(\theta)=-\sum_{t=1}^{|A|}\log P_R(a_t\mid a_{<t},e_{\text{ctx}},Q;\theta_{\text{bridge}}). \]

Here, \(e_{\text{ctx}}\) denotes the context embeddings supplied by LAM. Sender hidden states participate in conditioning through the bridge, although the printed conditional expression does not list them separately. Gold answers supervise training; they are not supplied at inference time.

The default balanced training set contains 587 examples from the seven task domains. The authors state that training and evaluation splits do not overlap, but applying the trained bridge across seven tasks does not mean all seven are unseen tasks: these domains participate in supervised training. Each sender→receiver pair is trained separately once, after which its bridge remains fixed across these tasks without task-specific adaptation.

Because the sender is frozen, document tokens and hidden states can be cached in advance, reducing subsequent bridge training to under 10 minutes on one GPU. That time assumes an existing cache and is not an end-to-end cost including cache generation, model loading, downloads, or training every model pair.

Key Experimental Results

Main Results

The table below selects the Llama-3.1-8B-Instruct→Qwen2.5-7B-Instruct pair from the paper's Table 1. Scores are answer F1 (%), and gains are percentage points. NLComm uses a greedily generated 128-token summary. FullComm is a single-model reference in which the receiver directly reads the full document, not a competing communication method, and is excluded from the paper's best-method ranking.

Task NoComm FullComm NLComm XBridge Gain over NLComm
Countries 0.0 50.8 22.6 72.5 +49.9
Tipsheets 66.5 98.5 96.3 99.8 +3.5
HotpotQA 24.6 78.4 68.5 78.8 +10.3
QASPER 7.1 31.3 33.7 47.5 +13.8
MuSiQue 10.6 31.5 32.5 48.2 +15.7
MFldQA 17.3 48.2 24.8 44.2 +19.4
2Wiki 19.3 47.6 14.2 51.1 +36.9
Seven-task average 20.8 55.2 41.8 63.2 +21.4

The other two heterogeneous pairs also outperform their corresponding NLComm on all seven tasks: Qwen→Llama averages 61.9 F1 versus 47.6, and Mistral→Qwen averages 65.0 versus 44.1. The three pairs cover three model families, not every directed pairing among them.

The advantage over FullComm is not “outperforming document reading without reading the document”: LAM already places the full mapped context in the receiver's input, with additional sender states. A higher average does not imply higher scores on every task; MFldQA above is 44.2 versus FullComm's 48.2.

Ablation Study

All results below use HotpotQA with Llama→Qwen. Component removals come from Table 3, module counts and gate initialization from Appendix Table 9, and grounding format from Appendix Table 8. Their control conditions differ, so they should not be merged into a ranking of one experimental factor.

Analysis dimension Config F1 (%) Note
Communication reference NoComm 24.6 Receiver sees only the question
Remove LAM LEB only 30.3 Continuous bridge lacks native lexical anchors
Remove LEB LAM only 56.5 Full mapped lexical input without latent enrichment
Full model LAM + four-module LEB 78.8 Complementary channels
Module count 1 module / 66M 66.3 Smaller bridge
Module count 2 modules / 132M 74.6 Greater capacity
Module count 7 modules / 462M 75.8 More modules do not improve further
Gate initialization Zero initialization 72.7 Below default warm initialization
Grounding format NLComm text + LEB 81.0 Still incurs sender summary generation cost

A separate entity intervention uses 100 HotpotQA samples and counts whether the answer contains the original entity, the substituted entity, or neither. These are sample counts, not F1 scores.

LAM input LEB input Original entity Substituted entity Neither
Original document Original document 45 0 55
Substituted document Substituted document 4 37 59
Substituted document Original document 5 37 58
Original document Substituted document 45 0 55

Key Findings

  • LAM and LEB are complementary: the full model improves over LAM alone by 22.3 points, while LEB alone improves over NoComm by just 5.7. Entity interventions support lexical identity primarily following LAM, but the many “neither” responses limit the evidence's coverage.
  • Balanced training averages 63.2 across seven tasks, compared with 49.5 for 20K HotpotQA-only training and 59.3 for 42K unbalanced training. Task coverage matters more than indiscriminately adding data; this does not establish that 587 examples cover arbitrary new tasks or models.
  • On H200, HotpotQA per-sample latency is 0.15 seconds for XBridge and 1.70 seconds for NLComm, roughly an 11-fold speedup; NoComm takes 0.09 seconds, and FullComm and KVComm each take 0.13 seconds. The main saving is sender-side autoregressive generation of a 128-token summary, not a demonstrated 11-fold speedup in bandwidth-limited, cross-machine deployments.
  • Homogeneous Qwen→Qwen uses a separately trained bridge and averages 63.1, compared with 51.1 for KVComm and 55.2 for FullComm. XBridge beats KVComm on six tasks, but its MFldQA score of 45.6 is below KVComm's 48.3.
  • Two independently trained bridges require no joint retraining in the dual-sender test and achieve 70.4 HotpotQA F1, compared with that test's single-sender full-context baseline of 67.0 and dual NLComm's 56.8. Zero-shot composition means combining existing bridges, not communicating across untrained model pairs; 67.0 is also not Table 1's FullComm score of 78.4.

Source consistency note: the main ablation discussion cites entity grounding ablations as Table 5, but they are in Table 3; Table 5 reports latency. The sender-scaling discussion similarly cites Table 5 instead of Table 4. Appendix B.2 says 587 samples are approximately 30 times fewer than 42K, whereas the listed sizes imply approximately 71.6 times fewer, and claims balanced training is best on every metric even though Tipsheets is 99.8 versus 100.0 for unbalanced training. Appendix C.2 labels 44.2−37.3 on MFldQA as +7.0, although the displayed difference is 6.9, possibly reflecting undisplayed precision. That discussion also mixes heterogeneous XBridge's 44.2 with homogeneous KVComm's 48.3, and its stated HotpotQA gain of +13.5 does not match Table 11's 87.1−65.3. This note preserves table values rather than converting those statements into a nonexistent unified comparison.

Highlights & Insights

  • Exact lexical identity and continuous knowledge need not be alternatives. Retaining a receiver-native lexical interface lets the latent bridge focus on contextual enrichment rather than recovering every rare token.
  • Receiver-driven retrieval is more targeted than fixed injection of heterogeneous states. Different layers can retrieve different information from the same sender states, while residual gates control the enrichment amplitude without providing a formal representation-compatibility guarantee.
  • The anchor format is not uniquely determined. Text anchors plus LEB reach 81.0, exceeding default LAM plus LEB at 78.8; the default favors removing sender generation latency rather than maximizing accuracy over every format.

Limitations & Future Work

  • Evaluation primarily covers one-way, single-turn document QA, not long-running multi-turn latent dialogue. Composition of two senders also does not establish stable scaling to arbitrary numbers of agents.
  • LEB is a supervised adaptation module with roughly 264M parameters, not a training-free universal translator. Freezing the backbones does not remove training, caching, and storage costs for each directed pair.
  • The payload contains full document tokens and the complete last-layer hidden-state sequence, and the receiver still processes a long input. The paper does not demonstrate reduced communication bytes or bandwidth savings on a real network; future evaluations should separately report payload bytes, network time, receiver prefill, and answer generation costs.
  • NLComm has a fixed 128-token summary budget and a question-unaware sender. Longer summaries, question-aware summaries, and other communication budgets require additional experiments; the present comparison does not establish superiority over all text protocols.
  • Representation analysis uses 196 HotpotQA samples, and lexical ranking uses the first answer token as a diagnostic proxy for multi-token answers. Improved cosine similarity and first-token rank do not guarantee faithfulness of all entities, complete answers, or relations.
  • Further tests should cover numbers, dates, multilingual text, special tokens, and context boundaries to delimit tokenwise mapping's content-preservation scope. Analyzing entity-substitution failures and long-document retrieval would also be more informative than presenting only successful examples.
  • vs NLComm / multi-agent debate: text messages are interpretable and architecture-agnostic; XBridge avoids sender summary generation and additionally exposes hidden states. The paper tests asymmetric, single-turn summary communication, not direct replacement of a full debate system.
  • vs KVComm / activation sharing: homogeneous KV or activation communication relies on architectural compatibility. XBridge uses lexical mapping and trainable cross-attention to accommodate heterogeneous interfaces, at the cost of full-context inputs and pairwise bridge training; KVComm can still be stronger on homogeneous retrieval tasks.
  • vs C2C / MoT: according to the paper's related work, the former projects caches across models and the latter lets a primary expert read peer representations, but their receivers already hold relevant inputs. XBridge emphasizes entity recovery when the receiver initially lacks the document, making the task setting as important as the module structure.
  • vs Flamingo: XBridge borrows gated cross-attention adaptation on frozen backbones, but retrieves another LLM's document states rather than visual representations and uses nonzero warm gate initialization. The transferable principle is to retain a native lexical interface before exploiting external continuous information through supervised retrieval.

Rating

  • Novelty: 4/5 — Separates entity identity from contextual enrichment in heterogeneous communication, targeting a specific failure mode.
  • Experimental Thoroughness: 4/5 — Includes three heterogeneous pairs, seven tasks, component ablations, and interventions, but lacks broad unseen-domain, network-cost, and multi-turn evaluation.
  • Writing Quality: 3/5 — The central mechanism is clear, but table references, data-size ratios, and appendix comparison scopes are inconsistent.
  • Value: 4/5 — Provides a practical interface for low-generation-latency heterogeneous collaboration, subject to pairwise-training and full-payload deployment constraints.