Beyond Script Family Boundaries: Towards Unified Open-Set Scene Text Recognition¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Code: https://github.com/lancercat/uni-zero
Area: Multimodal VLM
Keywords: Scene Text Recognition, Open-Set Recognition, Zero-Shot Learning, Multi-Script Representation, Glyph Prototyping
TL;DR¶
This paper introduces the first unified cross-script-family open-set scene text recognition framework that combines CAM-CTC alignment, dynamic prototype matching, and heterogeneous character-word co-training to enable zero-shot recognition and rejection across CJK, Indic, and low-resource Yi scripts within a single model.
Background & Motivation¶
Modern vision-language models (VLMs), such as the Gemini and ChatGPT families, have achieved remarkable capabilities in transcribing document and scene text images in high-resource languages like English and Chinese. However, when faced with low-resource scripts featuring vast, long-tailed character sets and scarce digital corpora, commercial VLMs exhibit severe performance breakdowns. Even though their subword tokenizers technically encompass Unicode code points for minority writing systems like the Yi script, they struggle fundamentally with fine-grained visual comparison of distinct glyph shapes and contour details, leading to complete hallucination. This degradation indicates that general-purpose foundation models fail to reliably establish the visual metric boundaries of "to be" (identifying identical characters) and "not to be" (differentiating distinct characters).
Open-set and zero-shot text recognition paradigms offer a promising route to transcribe unseen characters without parameter retraining. Nonetheless, prior research has remained almost entirely confined within the Chinese-Japanese-Korean (CJK) family. Existing zero-shot Chinese character recognition methods rely heavily on language-specific priors, such as radical decompositions, stroke structures, or ideographic description trees. Such priors cannot generalize to non-ideographic scripts and lack the essential capability to reject "unknown unknown" characters absent from the registered vocabulary. While recent open-set text recognition (OSTR) methods incorporate rejection mechanisms, their quantitative benchmarks are restricted to closely related scripts like Japanese and Korean, failing to offer a script-agnostic architecture capable of handling vastly different writing systems like Brahmic/Indic scripts or syllabic Yi. Furthermore, naively co-training across multiple script families causes acute representational interference and gradient conflicts that degrade performance.
To transcend script family boundaries and answer whether open-set text recognition can genuinely scale beyond CJK, this paper attacks the problem at the level of representation decoupling and flexible spatial-temporal alignment: since human perception identifies unfamiliar glyphs via universal contour comparison and visual primitives, a deep model should likewise build a unified metric space through shared visual backbones, isolated domain statistics, and flexible feature sampling. Core idea: develop a unified, script-family-agnostic multi-task co-training open-set text recognition framework that leverages CAM-CTC alignment to resolve complex layout segmentation, paired with heterogeneous character-and-word synthetic co-training to achieve zero-shot recognition and open-set rejection across CJK, Indic, and Yi scripts simultaneously.
Method¶
Overall Architecture¶
The proposed unified open-set text recognition framework employs a pool of partially shared modules to accommodate heterogeneous typographies and diverse character layouts across multiple script families. The system comprises two main task categories: a synthetic character recognition task and word-level scene text recognition tasks. Each task is structured as a dedicated script group branch equipped with its own Region of Interest (RoI) Feature Extractor, dynamic prototype generation module (\(mkproto\)), and Open-Set Classifier (\(OSC\)).
Input images are first fed into a partially shared convolutional backbone. The convolutional kernels are shared across all tasks to exploit transferable low-level stroke and edge features, while task-specific Domain-specific Batch Normalization (BN) layers are dynamically swapped in according to the active script domain to prevent statistical interference across languages. Next, the extracted feature map tower is passed to task-dependent samplers: the character task uses a lightweight convolutional foreground sampler for single glyphs, whereas the word recognition task employs a Convolutional Attention Module (CAM) to produce an attention mask sequence over the spatial domain. Departing from conventional position-wise decoders that predict fixed word lengths, the sampled RoI feature sequence is supervised with a Connectionist Temporal Classification (CTC) loss, bypassing rigid character-level segmentation constraints. During inference, the prototyping module transforms standard font glyph images into normalized class centers, and the open-set classifier computes cosine similarities alongside a learnable hyperbolic tangent unknown threshold to output either recognized character labels or the [unk] rejection token.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Scene Word Images and Glyph Templates"] --> B["Script-Specific Normalization & Task Grouping<br/>Shared Convolutional Backbone + Task-Specific BN Layers"]
B --> C["CAM-CTC Alignment Pipeline<br/>Spatial Attention Mask Sampling + CTC Sequence Loss"]
C --> D["Dynamic Prototyping Module & Open-Set Classifier<br/>Normalized Glyph Centers + Learnable tanh Unknown Threshold"]
E["Heterogeneous Synthetic Co-training<br/>Full Unicode Glyph Cross-Entropy + Word Shuffle CTC"] -.->|End-to-End Joint Multi-Task Optimization| C
E -.->|End-to-End Joint Multi-Task Optimization| D
D --> F["Output: Closed-Set Transcriptions / Open-Set Out-of-Vocabulary Rejections"]
Key Designs¶
1. Script-Specific Normalization & Task Grouping: Resolving Inter-Script Conflicts and Gradient Degeneracy
Attempting to train radically different writing systems—such as Latin, Chinese logographs, and Brahmic Indic scripts—within a single monolithic network induces destructive negative transfer. Indic scripts incorporate multi-directional vowel matras and consonant ligatures, yielding aspect ratios and spatial stroke densities starkly divergent from square CJK ideographs. To preserve cross-script visual feature sharing without feature disruption, the framework maintains shared convolutional weights while assigning independent Batch Normalization parameters to each script group (e.g., CJK, core Indic, Punjabi, Tamil, English synthetic): $\(M = \text{ImEnc}(I^{dom}, BN^{dom})\)$ Crucially, empirical analysis revealed severe conflicts even within the Indic family itself. To mitigate this internal polarization, the authors isolate highly conflicting scripts (specifically Gurmukhi/Punjabi and Tamil) into separate task branches rather than pooling all Indic data together. This selective architectural decoupling prevents disparate distributional variances from destabilizing gradient trajectories.
2. CAM-CTC Alignment Pipeline: Combining Local Spatial Attention with Temporal Sequence Flexibility
In cross-family zero-shot recognition, the model lacks prior knowledge regarding character boundaries or ligature rules. Traditional position-wise attention decoders (such as OpenCCD) require rigid one-to-one character localization and sequence length estimation; any attention drift on unseen connected scripts (e.g., Bengali or Gujarati) precipitates catastrophic sequence-wide decoding failures. Conversely, 1D height-squeezing with CTC (such as Rosetta) loses crucial 2D spatial layout cues. This work integrates a Convolutional Attention Module (CAM) with Connectionist Temporal Classification (CTC): $\(A = S^{seq}_{dom}(M) = \text{CAM}(\text{SpatialEmb}(\text{detach}(M)))\)$ $\(F^{dom}[t] = \frac{A[t] \cdot M[2]}{\sum A[t]}\)$ The CAM module injects 2D spatial embeddings into the detached feature tower to construct a sequence of spatial RoI masks \(A\), weighting the highest-level feature map \(M[2]\) into character representations \(F^{dom}\). Importantly, the masks are no longer forced to delineate exact physical character boundaries; rather, the downstream CTC loss discovers the optimal unaligned alignment path. This grants the sampler spatial resolution benefits while affording temporal tolerance against overlapping or irregular characters in unseen scripts.
3. Dynamic Prototyping Module & Open-Set Classifier: Decoupling Character Registration and Unknown Rejection
To facilitate test-time open-set recognition without model retraining, novel character categories must be dynamically registered as class centers. The prototyping module (\(mkproto\)) takes isolated standard glyph renderings \(I^{tem}\) of unseen characters, processes them through a dedicated character RoI extractor and template BN layer to yield raw prototypes \(p_{raw}\), prepends a learnable CTC blank embedding \(E_{sp}^{sg}\), and applies L2 normalization:
$\(p = \text{normalize}([E_{sp}^{sg}, p_{raw}])\)$
Within the Open-Set Classifier (\(OSC\)), fixed linear classification heads are eliminated in favor of cosine similarity computation between feature representations \(F[t]\) and prototypes \(p[j]\). To reject characters outside the support set, the classifier incorporates a learnable scalar threshold \(th_{[-]}\) mapped through a hyperbolic tangent function:
$\(s^{class}[t, i] = \|F[t]\|_2 \max_{y_t[j]=i} [\text{cosine}(F[t], p[j]), \tanh(th_{[-]})]\)$
The L2 norm \(\|F[t]\|_2\) does not alter relative similarity rankings but functions as a dynamic temperature factor regulating the entropy of the softmax output. When an observed visual character matches no registered prototype, its maximum similarity remains below \(\tanh(th_{[-]})\), predictably emitting the [unk] rejection label and suppressing hallucinations.
4. Heterogeneous Synthetic Co-training: Bridging Fine-Grained Glyph Primitives and Word-Level Compositions
Real scene text datasets cover only a tiny fraction of global scripts and typographical varieties. This work constructs a decoupled synthetic data generation framework operating jointly across character and word levels. At the character level, all supported Unicode glyphs are rendered into binary masks; empty placeholder glyphs (detected as identical shapes mapped to more than five distinct codepoints) are pruned. In each training iteration, 64 character categories are sampled to synthesize perturbed glyph-sample pairs optimized via cross-entropy loss against prototypes, instilling universal shape comparison capability. At the word level, authentic word corpora are parsed into Unicode graphemes and randomly shuffled before rendering with fonts supporting the entire sequence. Word-level samples are merged onto synthetic backgrounds and trained via CTC loss. This heterogeneous pairing enables the backbone to capture atomic stroke geometries while simultaneously acquiring robust temporal composition modeling.
A Worked Example¶
Consider recognizing a synthetic 4-character word image in the rare Yi syllabary:
1. Prototype Registration: Prior to inference, 1,165 standardized modern Yi font glyph images are forwarded through the frozen prototyping module \(mkproto\), producing an offline cache of normalized \(d\)-dimensional prototype vectors \(p\).
2. Feature Extraction & Mask Sampling: A scene word crop (\(32 \times 128\)) is fed into the shared backbone with the CJK/general domain BN layers to yield feature tower \(M\). The CAM sampler generates sequential attention masks across the horizontal axis, aggregating \(M[2]\) into a feature sequence \(F[t]\) of length \(\max T\).
3. Similarity Matching & Rejection: At each timestep \(t\), the classifier calculates cosine similarities between \(F[t]\) and all 1,165 Yi prototypes against the unknown threshold \(\tanh(th_{[-]})\). If the first three characters match standard Yi syllables with high similarity (\(>0.85\), well above the threshold \(0.35\)), they are mapped to their respective Unicode labels. If the fourth character is severely corrupted or represents an unregistered archaic glyph, its maximum similarity falls below \(0.35\), allowing \(\tanh(th_{[-]} )\) to win and assign the [unk] rejection label.
4. CTC Beam Search Decoding: The time-step probability distributions are decoded using CTC beam search, collapsing repeated predictions and blank tokens [-]. The final transcription reliably yields the three valid characters while rejecting the corrupted glyph, avoiding the fictitious generation typical of large autoregressive models.
Loss & Training¶
The framework optimizes a composite multi-task objective comprising three distinct loss functions: 1. Synthetic Character-Level Cross-Entropy: $\(L_{char} = \text{CrossEntropy}(OSC(FE_{seq}(I_{char}, BN^{sg}, S_{char}^{sg}), mkproto(I_{tem})), y_{char})\)$ 2. Real Word-Level CTC Loss: $\(L_{word} = \text{CTC}(OSC(FE_{seq}(I_{real}, BN^{sg}, S_{seq}^{sg}), p_{batch}), y_{real})\)$ 3. Synthetic Word-Level CTC Loss: $\(L_{syn} = \text{CTC}(OSC(FE_{seq}(I_{syn}, BN_{syn}^{sg}, S_{seq}^{sg}), p_{batch}), y_{syn})\)$ The complete objective is a weighted combination across active task branches. All variants are trained for 140,000 iterations using the AdamW optimizer with joint multi-task gradient accumulation.
Key Experimental Results¶
Main Results¶
The framework is evaluated under the Generalized Zero-Shot Learning (GZSL) protocol measuring word-level accuracy (ACR). Evaluations cover Japanese (JPN), Korean (KR), the low-resource Yi script, and unseen Indic scripts (Bengali, Gujarati):
| Method | JPN (ACR) | KR (ACR) | Yi (ACR) | BEN-Val | GUJ-Val | BEN-Test | GUJ-Test | Note |
|---|---|---|---|---|---|---|---|---|
| OpenCCD (CVPR 2022) | 41.31 | - | - | - | - | - | - | Limited to CJK |
| SAVR (ICDAR 2023) | 42.58 | - | - | - | - | - | - | Visual reconstruction |
| CFOR (T-IP 2024) | 44.47 | 22.14 | - | - | - | - | - | Noto glyph pretraining |
| WnA-XL (ICDAR 2025) | 48.02 | 24.00 | - | - | - | - | - | Dynamic MoE with variable input |
| Ours-CJK (Chinese + Syn) | 45.92 | 23.32 | 16.10 | - | - | - | - | Monolithic 32×128 input |
| Ours-IND (Indic + Syn) | - | - | - | 8.04 | 7.52 | 6.82 | 8.68 | Indic transfer baseline |
| Ours-Uni (Unified Model) | 46.30 | 22.84 | 14.22 | 12.88 | 9.72 | 9.02 | 10.04 | Unified cross-family model |
| ChatGPT-5.2 (50-sample subset) | - | - | 0.00 | - | - | - | - | CER 148.44%, severe hallucination |
On the 50-sample modern Yi benchmark, commercial ChatGPT-5.2 (provided with the complete 1,165-character candidate list) scored 0% ACR with a Character Error Rate (CER) of 148.44% due to uncontrolled sequence hallucination. Conversely, Ours-Uni achieved 18.00% ACR and 48.67% CER, underscoring the necessity of specialized open-set metric models for minority scripts.
In the standard Open-Set Text Recognition (OSTR) benchmark containing out-of-set Japanese Kanji and Latin characters, Ours-Uni reached 75.38% ACR, 78.34% rejection recall, 94.59% rejection precision, and an 85.70% F-measure, confirming that cross-script scaling preserves high precision when discarding unseen characters.
Ablation Study¶
1. Word Pipeline Structure & Alignment Mechanism Ablation (Word Accuracy across zero-shot scripts):
| Pipeline Variant | Sampling Mechanism | Alignment Mechanism | ACR (JP) | ACR (KR) | ACR (Yi) | ACR (Bengali) | ACR (Gujarati) |
|---|---|---|---|---|---|---|---|
| OpenCCD-like | CAM (2D Attention) | Position-wise | 41.59 | 15.62 | 5.73 | 2.81 | 2.30 |
| Rosetta-like | Squeeze (1D Collapse) | CTC | 41.79 | 12.65 | 6.68 | 3.36 | 5.17 |
| Ours | CAM (2D Attention) | CTC | 42.36 | 19.18 | 8.51 | 4.74 | 5.98 |
2. Synthetic Co-training Strategy Ablation (Performance gains from heterogeneous synthetic data):
| Configuration | Word-Level Syn | Char-Level Syn | ACR (JP) | ACR (KR) | ACR (Yi) | Note |
|---|---|---|---|---|---|---|
| Baseline (No Co-training) | No | No | 42.36 | 19.18 | 8.51 | Real data only |
| 2X Batch Size | No | No | 43.39 | 16.60 | 7.49 | Matches sample count; novel scripts drop |
| Word Co-training (LSCT-Word) | Yes | No | 44.24 | 21.87 | 11.42 | Intra-word shuffle boosts context |
| Char Co-training (LSCT-Char) | No | Yes | 45.17 | 23.30 | 12.05 | Fine-grained glyph discriminability |
| Full Co-training (LSCT-Full) | Yes | Yes | 46.08 | 23.50 | 17.07 | Dual-level heterogeneous synergy |
Key Findings¶
- CAM and CTC Synergy is Crucial: Relying solely on position-wise attention decoders leads to severe failure on unfamiliar Indic writing (Bengali drops to 2.81%), whereas 1D height-squeezing loses subtle character contours. CAM provides localized spatial RoI masks while CTC absorbs temporal misalignments, yielding a +6.53% surge on Korean and +1.93% on Bengali.
- Co-training Benefits Exceed Mere Sample Scaling: Doubling the training batch size (2X Batch size) marginally improves seen Japanese (+1.03%) but degrades performance on completely unseen Korean (-2.58%) and Yi (-1.02%). In contrast, full heterogeneous co-training (LSCT-Full) doubles the Yi ACR from 8.51% to 17.07%, demonstrating that multi-granularity synthetic data enriches universal geometric representations.
- Inter-Script Interference within Indic Languages: Training all Indic languages as an undifferentiated single task induces acute polarization, causing recognition on Bengali, Gujarati, and Oriya to collapse near zero (ACR \(\le 0.04\%\)). Isolating Punjabi and Tamil into independent task branches restores balanced recognition across all scripts (Tamil recovers from 0.12% to 77.59%).
Highlights & Insights¶
- Breaking the CJK Boundary: Pioneers open-set text recognition across disparate script systems (CJK, Indic, Latin, and syllabic Yi), demonstrating the feasibility of metric learning on character glyph prototypes across unrelated writing systems.
- Hybrid CAM-CTC Architecture: Elegantly balances fine-grained 2D spatial attention with the flexible, alignment-free sequence modeling of CTC, sidestepping fragile character segmentation on unseen scripts.
- Addressing VLM Blind Spots: Achieves robust zero-shot recognition on minority scripts where frontier commercial models like ChatGPT-5.2 collapse to 0% accuracy and generate pure hallucinations, providing a crucial tool for linguistic inclusion.
Limitations & Future Work¶
- Sensitivity to Extreme Typographical Distortions: Severe blurring or heavy artistic styling causes the model to trigger unknown rejections; while preferable to hallucination, improving geometric spatial invariance remains necessary.
- Heuristic Task Partitioning: Resolving inter-script negative transfer currently depends on empirical task grouping. Developing automated Mixture-of-Experts (MoE) routing or soft cluster grouping will be critical for scaling to hundreds of languages.
- Script and Modality Scope: The framework has been validated on scene text and synthetic datasets, but has not yet been extended to complex historical handwritten manuscripts or cursive scripts like Arabic.
Related Work & Insights¶
- vs WnA (ICDAR 2025): WnA relies on a complex multi-orientation MoE routing mechanism and dynamic input resolutions to maximize CJK accuracy; this work achieves competitive performance using a fixed \(32 \times 128\) resolution while demonstrating broad zero-shot transfer across multiple script families.
- vs CFOR (T-IP 2024): CFOR focuses on context-free pretraining using isolated Noto font glyphs within CJK and Latin; this work introduces word-level grapheme shuffling and domain-specific BN layers, demonstrating far superior out-of-domain transfer to Indic and Yi writing.
- vs OpenCCD (CVPR 2022): OpenCCD enforces rigid position-wise decoding and character sequence length prediction; this paper replaces the position-wise decoder with CTC, successfully resolving character segmentation failures on unseen ligatures.
Rating¶
- Novelty: ⭐⭐⭐⭐ [First unified framework to extend open-set scene text recognition beyond CJK to Indic and Yi families with zero-shot generalization]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across Japanese, Korean, Yi, and multiple Indic scripts against state-of-the-art specialized models and commercial foundation VLMs]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear articulation of research questions, thorough architectural exposition, and transparent reporting of empirical conflicts and limitations]
- Value: ⭐⭐⭐⭐⭐ [Crucial benchmark and methodology for mitigating technological exclusion and VLM hallucination in minority low-resource scripts]