Q-BridgeNet: A Quantization Network for Cross-Lingual Sign Language Translation¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/FengLiQ/Q-BridgeNet
Area: Human Understanding
Keywords: Sign Language Translation, Residual Vector Quantization, Adaptive Temporal Segmentation, Multilingual Modeling, Large Language Model
TL;DR¶
Q-BridgeNet tackles negative cross-lingual transfer and rigid temporal boundary cutting in multilingual sign language translation by introducing adaptive temporal segmentation and shared–private residual vector quantization to learn discrete Q-units, fine-tuning a multilingual LLM for many-to-many translation.
Background & Motivation¶
Over 360 million people worldwide rely on sign languages for everyday communication. As visual languages possessing independent phonological, syntactic, and discourse systems, sign languages differ fundamentally from spoken languages. The overwhelming majority of existing sign language translation (SLT) research focuses on isolated, native sign–spoken language pairs (such as American Sign Language ASL–English, German Sign Language DGS–German, or Chinese Sign Language CSL–Chinese), mapping continuous video or skeleton pose sequences into target spoken text. Although these single-language models have achieved notable progress in in-pair settings, practical communication demands seamless interaction across diverse sign languages and multiple spoken language communities, giving rise to multilingual SLT.
However, existing multilingual sign language translation models face severe negative transfer and representational interference when unifying multiple sign languages into a single representation space. Forcing heterogeneous visual signing realizations into a shared continuous embedding manifold causes language-specific motion patterns, execution speeds, and grammatical structures to compete for representational capacity, degrading the shared manifold. Concurrently, signing motions exhibit irregular temporal dynamics; conventional tokenization schemes impose uniform downsampling or fixed-window slicing, which inevitably slices across semantic or lexical boundaries, producing fragmented gesture tokens that align poorly with spoken language texts.
To resolve the core tension between heterogeneous signing variations and asynchronous temporal boundaries, Q-BridgeNet decouples signing motions into language-agnostic semantic primitives and language-specific refinement details, discovering semantic units through data-driven adaptive segmentation. Core idea: decompose continuous sign poses into discrete Q-unit sequences via alternating block-coordinate descent between adaptive temporal segmentation and shared–private residual vector quantization, fine-tuning a multilingual LLM on these discrete units to eliminate cross-lingual representational conflicts.
Method¶
Overall Architecture¶
Q-BridgeNet establishes an end-to-end multilingual many-to-many sign language translation pipeline. The framework operates in two distinct stages: the first stage performs iterative Q-unit discovery, where continuous sign pose sequences are partitioned into semantically coherent variable-length segments and quantized into discrete Q-units through a shared–private residual vector quantizer (RQ-VAE); the second stage freezes the quantization network and fine-tunes a multilingual LLM (Qwen3-1.7B) via LoRA to autoregressively decode target spoken translations from the discrete Q-unit tokens under multilingual supervision.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input sign pose sequence<br/>79 keypoint temporal trajectories"] --> B["Adaptive temporal segmentation<br/>dynamic programming for optimal boundaries"]
B --> C["Shared-private residual vector quantization<br/>shared base codebook and private residual codebooks"]
C --> D["Block-coordinate descent alternating optimization<br/>iterative codebook and boundary convergence"]
D --> E["Multilingual LLM re-alignment<br/>discrete Q-units autoregressive decoding"]
E --> F["Target spoken text generation<br/>German / English / Chinese many-to-many translation"]
Key Designs¶
1. Adaptive temporal segmentation: dynamic programming eliminating rigid boundary cuts
Sign language gestures exhibit highly non-uniform temporal pacing, making fixed-length time windows prone to bisecting gestures across lexical boundaries. This design identifies semantically coherent variable-length segments by minimizing cumulative reconstruction error. For an input sequence of \(T\) skeleton frames \(\mathbf{P}=\{\mathbf{p}_1, \dots, \mathbf{p}_T\}\), the objective is to partition the time axis into \(N\) non-overlapping half-open intervals \(\mathcal{S}(\mathbf{P})=\{\mathbf{p}_{s_0:s_1}, \dots, \mathbf{p}_{s_{N-1}:s_N}\}\) (with \(s_0=1, s_N=T+1\), and maximum segment duration constrained to 32 frames). To prevent combinatorial search explosion, the model exploits the optimal substructure property to formulate a dynamic programming recurrence:
with base case \(\mathcal{L}_{\text{global}}(\mathbf{p}_{0:1}\mid\mathbb{V})=\mathcal{L}_{\text{elem}}(\mathbf{p}_{0:1}\mid\mathbb{V})\), where the element-wise cost \(\mathcal{L}_{\text{elem}}\) measures the reconstruction error of an individual segment under the corresponding RQ-VAE module. Sequential evaluation and backtracking yield optimal breakpoints corresponding to structurally complete signing units.
2. Shared-private residual vector quantization: decoupling shared primitives from language-specific nuances
Compressing multiple sign languages into a single codebook induces severe cross-lingual interference, whereas completely isolated codebooks eliminate cross-lingual knowledge sharing. The framework constructs \(L\) distinct RQ-VAE modules for \(L\) sign languages (ASL, CSL, DGS). Given the latent feature \(\mathbf{z}=\mathcal{E}(\mathbf{p}_{s_n:s_{n+1}})\), quantization decomposes \(\mathbf{z}\) across \(R+1\) ordered residual layers:
Crucially, the base quantizer \(\mathcal{Q}_0\) and its codebook \(\mathcal{Z}_0\) (containing 1,024 96-dimensional embeddings) are universally shared across all sign languages to capture language-agnostic semantic primitives. Subsequent residual quantizers \(\{\mathcal{Q}_i^r\}_{r=1}^R\) employ private, language-specific codebooks (512 embeddings per language) to absorb distinct articulatory amplitudes and signing styles. Concatenating all residual vectors produces the unified discrete token \(\hat{\mathbf{z}} = \mathbf{z}^0 \oplus \dots \oplus \mathbf{z}^R\), which reconstructs the pose segment via decoder \(\mathcal{D}_i(\hat{\mathbf{z}})\).
3. Block-coordinate descent alternating optimization: resolving boundary-codebook interdependency
Precise boundary discovery requires a well-trained codebook to provide reliable reconstruction metrics, while high-quality codebook learning relies on pure, low-variance segments. Because boundaries and codebooks are mutually dependent, the joint objective is formulated as:
where \(\lambda N\) regularizes segment frequency. Optimization proceeds via block-coordinate descent: fixing segmentation \(\mathcal{S}^t\) updates the codebook parameters \(\mathbb{V}^{t+1}\), after which fixing the updated codebook guides dynamic programming to update temporal boundaries \(\mathcal{S}^{t+1}\). Alternating over 5 rounds ensures monotonic loss reduction and converges to a stable local optimum, successfully harmonizing motion boundaries with discrete tokenization.
4. Multilingual LLM space re-alignment: bridging discrete sign units to target spoken languages
Once discrete Q-unit modeling is complete, all RQ-VAE weights are frozen. To project the discrete symbolic tokens into multilingual natural language, a pretrained multilingual LLM (Qwen3-1.7B) is adapted via LoRA. The training corpus is augmented by using GPT-5 to translate native spoken captions across German, English, and Chinese, creating trilingual parallel targets paired with explicit prompts (e.g., Translate <DGS> to German). Feeding the discrete token sequence \(\hat{\mathbf{Z}}=(\hat{\mathbf{z}}_1, \dots, \hat{\mathbf{z}}_N)\) into the LLM, training minimizes autoregressive language modeling cross-entropy loss:
By capitalizing on the broad cross-lingual priors inherent in the LLM, the model handles arbitrary sign–spoken language pairs through a single unified architecture.
Loss & Training¶
RQ-VAE training spans 1,000 epochs with loss weights \(\beta=1.0\) and \(\gamma=0.25\), using a 96-dimensional latent space. Skeleton trajectories are represented by 79 keypoints extracted via OpenPose, augmented with Gaussian noise (\(\sigma=0.002\)) to withstand detection jitter. Alternating block-coordinate descent runs for 5 iterations from a random partition. LLM fine-tuning is executed over 100 epochs across the augmented trilingual dataset on four NVIDIA RTX 4090 GPUs using LoRA.
Key Experimental Results¶
Main Results¶
Evaluation is conducted across three benchmarks: PHOENIX14T (DGS), How2Sign (ASL), and CSL-Daily (CSL). Translation quality is evaluated with BLEU-4 (B-4), BLEU-1 (B-1), and ROUGE-L (RG).
| Dataset / Pair | Method | B-4 | B-1 | RG | Note |
|---|---|---|---|---|---|
| How2Sign (ASL-EN) | SL-Trans. (CVPR'20) | 9.93 | 28.37 | 28.94 | Early single-language baseline |
| GloFE (ACL'23) | 2.24 | 14.94 | 12.61 | End-to-end gloss-free model | |
| YT-ASL (NeurIPS'23) | 12.39 | 37.82 | – | Large-scale pretraining baseline | |
| Uni-Sign (ICLR'25) | 14.50 | 40.40 | 34.30 | Unified multilingual SLT model | |
| ShuBert (ACL'25) | 16.20 | – | – | Multi-stream cluster learning prior SOTA | |
| Ours (ASL-EN Native) | 16.55 | 41.81 | 34.97 | Outperforms all prior methods | |
| Ours (Multilingual Avg.) | 16.44 | 41.73 | 35.20 | Robust across non-native targets | |
| Ours (ASL-ZH Non-native) | 16.83 | 42.04 | 35.52 | Outperforms native English target | |
| Ours (ASL-DE Non-native) | 15.96 | 41.34 | 35.11 | Strong cross-lingual generalization | |
| CSL-Daily (CSL-ZH) | SL-Trans. (CVPR'20) | 14.15 | 39.72 | 38.44 | Standard Transformer baseline |
| MMSLT (ICCV'25) | 21.11 | 49.87 | 48.92 | MLLM-based SLT framework | |
| ICTCA (COLING'25) | 22.47 | 52.37 | 51.87 | Text-CTC alignment approach | |
| MSKA (PR'25) | 25.52 | 56.37 | 54.04 | Multi-stream keypoint attention | |
| Uni-Sign (ICLR'25) | 25.61 | 53.86 | 54.92 | Multilingual unified prior SOTA | |
| Ours (CSL-ZH Native) | 26.91 | 57.39 | 54.85 | Establishes new SOTA benchmark | |
| Ours (Multilingual Avg.) | 26.33 | 57.09 | 55.04 | Superior multi-target translation | |
| Ours (CSL-EN Non-native) | 26.40 | 56.98 | 54.30 | High-quality non-native output | |
| Ours (CSL-DE Non-native) | 25.68 | 56.89 | 55.97 | Robust transfer to distant language | |
| PHOENIX14T (DGS-DE) | SL-Trans. (CVPR'20) | 19.45 | 43.72 | 45.44 | Classical two-stage model |
| TwoStream (NeurIPS'22) | 28.42 | 54.22 | 53.19 | Dual-stream visual modeling | |
| CV-SLT (AAAI'24) | 29.27 | 54.88 | 54.33 | Conditional VAE formulation | |
| SCOPE (AAAI'25) | 32.84 | 61.74 | 60.06 | LLM contextual embedding prior SOTA | |
| Ours (DGS-DE Native) | 33.62 | 62.80 | 61.76 | Reaches highest recorded accuracy | |
| Ours (Multilingual Avg.) | 33.40 | 62.56 | 61.39 | Consistently strong performance | |
| Ours (DGS-ZH Non-native) | 33.36 | 62.67 | 61.67 | Direct translation to Chinese | |
| Ours (DGS-EN Non-native) | 33.20 | 62.21 | 60.73 | Exceeds 33 BLEU-4 on English |
Ablation Study¶
The ablation study validates the effectiveness of each architectural component, alternating optimization, supervision quality, and LLM backbone capacity.
| Config | How2Sign (B-4 / B-1 / RG) | CSL-Daily (B-4 / B-1 / RG) | PHOENIX14T (B-4 / B-1 / RG) | Note |
|---|---|---|---|---|
| Full model | 16.44 / 41.73 / 35.20 | 26.33 / 57.09 / 55.04 | 33.40 / 62.56 / 61.39 | Shared base + private residual + alternating optimization |
| Uni-codebook VQ | 7.23 / 21.61 / 18.09 | 16.18 / 33.97 / 35.12 | 21.38 / 42.34 / 39.22 | Severe cross-lingual conflict with single codebook |
| Non-shared VQ | 9.29 / 25.60 / 22.52 | 18.90 / 37.82 / 39.01 | 22.62 / 42.03 / 43.14 | Fragmented latent space lacking cross-lingual pivot |
| Non-residual VQ | 12.04 / 31.48 / 27.75 | 21.49 / 41.44 / 44.32 | 27.32 / 50.98 / 52.45 | Flat codebooks fail to refine fine-grained residual errors |
| Single-pass Seg. & RQ | 13.39 / 33.24 / 31.56 | 24.11 / 46.47 / 47.36 | 28.61 / 52.03 / 49.46 | Lacks bidirectional feedback between boundaries and codes |
| GPT-5 Translation Control | 14.20 / 36.61 / 31.82 | 23.18 / 51.47 / 50.02 | 28.36 / 54.13 / 52.26 | Monolingual SLT + GPT translation confirms gains stem from sign space |
| GPT-4.1 Created Caption | 15.58 / 41.36 / 34.07 | 26.30 / 55.92 / 54.48 | 33.07 / 62.34 / 60.79 | Marginal drop, demonstrating robustness to translation quality |
| Qwen3-0.6B Fine-Tuning | 15.64 / 41.01 / 34.20 | 25.84 / 55.41 / 53.54 | 32.25 / 60.76 / 59.01 | Maintains competitive accuracy with reduced parameter footprint |
Key Findings¶
- Shared-private residual decoupling is indispensable: Eliminating the shared base codebook (Non-shared VQ) or collapsing into a vanilla single codebook (Uni-codebook VQ) leads to catastrophic performance drops, plunging BLEU-4 by 10.78 and 12.02 points respectively on PHOENIX14T. This confirms the necessity of isolating language-agnostic semantic anchors from language-specific nuances.
- Residual stacking and alternating updates deliver consistent gains: Residual quantization outperforms parallel flat codebooks (Non-residual VQ) by 4.40 BLEU-4 on How2Sign, while 5-round block-coordinate descent achieves a 3.05–4.79 BLEU-4 improvement over single-pass segmentation, underscoring the mutual dependency between segmentation boundaries and code assignments.
- Multilingual supervision benefits native translation: The multilingual unified model outperforms monolingual models trained strictly on native pairs (e.g., ASL-EN increases from 14.35 to 16.55, and DGS-DE improves from 30.16 to 33.62), demonstrating that multilingual objectives regularize the discrete signing space and expand semantic coverage.
- Emergence of gloss-like distributions in unsupervised Q-units: Token frequency analysis reveals that unsupervised private Q-units follow a long-tailed decay distribution matching ground-truth manual gloss annotations on CSL and DGS, proving that hierarchical quantization captures authentic linguistic regularities rather than arbitrary clustering artifacts.
Highlights & Insights¶
- Shared-private residual codebook design: The architecture elegantly pairs a shared base codebook to ground cross-lingual semantic invariants with private residual codebooks to absorb language-specific articulatory variations, preventing manifold collapse and negative transfer.
- Bi-directional synergy between segmentation and quantization: Moving beyond arbitrary fixed-window frame slicing, dynamic programming segmentation and residual quantization are jointly optimized via block-coordinate descent, naturally aligning motion boundaries with semantic units.
- Decoupled discrete representation for seamless LLM transfer: Discretizing continuous visual gestures into symbolic Q-units shields the multilingual LLM from high-frequency visual noise, enabling seamless adaptation across different model scales (1.7B and 0.6B) and zero-shot cross-lingual generalization.
Limitations & Future Work¶
- Reliance on pre-extracted skeleton quality: The framework relies on pose estimates derived from OpenPose; severe hand occlusions, rapid motions, or lighting shifts that corrupt joint coordinates directly impact downstream segmentation and quantization fidelity.
- Dependence on synthetic multilingual supervision: Cross-lingual supervision relies on LLM-translated parallel captions; subtle translation artifacts or cultural idiom mismatches could potentially inject label noise into non-native tuning.
- Extension to sign language production (SLP): While currently designed for sign-to-text translation (SLT), the structured, discrete nature of Q-units offers a promising foundation for reverse text-to-sign generation (SLP), paving the way toward bidirectional, symmetric multilingual communication.
Related Work & Insights¶
- vs Uni-Sign (ICLR'25): While Uni-Sign achieves unified multilingual SLT using continuous visual features, Q-BridgeNet resolves cross-lingual feature entanglement by establishing a discrete, hierarchically quantized Q-unit space that decouples shared primitives from language-specific residuals.
- vs MMSLT (ICCV'25) / SCOPE (AAAI'25): These works utilize multimodal LLMs or contextual LLM embeddings for monolingual sign language translation; Q-BridgeNet breaks through language-specific boundaries and provides a unified discrete symbolic interface across multiple sign and spoken languages.
- vs Traditional Gloss-based SLT: Classical pipelines rely heavily on expensive, labor-intensive gloss annotations as an intermediate bottleneck; Q-BridgeNet learns discrete, gloss-like symbolic units completely without gloss supervision, matching the structural properties of human annotations.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneers a shared-private residual quantization and variable-length segmentation framework for multilingual sign language tokenization]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive comparisons across three major sign language benchmarks, comprehensive ablations, and structural distribution analyses]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical formulation, cohesive narrative, and well-structured empirical validation]
- Value: ⭐⭐⭐⭐⭐ [Provides a practical and principled foundation for multi-pair sign-to-spoken translation, accelerating global accessibility technology]