Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition¶
Conference: ECCV 2026
Paper: ECCV Official
Area: NLP Understanding
Keywords: Scene Text Recognition, Rotation Invariance, Cross-Attention, Multi-Oriented Text Recognition, Group Equivariant Convolution
TL;DR¶
To eliminate the error accumulation and inference overhead stemming from explicit orientation prediction and rotation data augmentation in multi-oriented scene text recognition, RISTER incorporates rotation equivariance into the encoder (via the RELG backbone) and rotation invariance into the decoder (via cross-attention in RITD), creating an end-to-end rotation-invariant recognition system with theoretical guarantees that outperforms prior SOTA by 4.0% on multi-oriented benchmarks.
Background & Motivation¶
Scene Text Recognition (STR) is a foundational capability for translating text in natural scene images into machine-readable digital representations, serving as a critical component in autonomous driving, environmental visual perception, and embodied intelligence. With advancements in language modeling and visual-language alignment, the recognition accuracy of conventional horizontal text arranged left-to-right has approached saturation. However, text in real-world scenes exhibits arbitrary spatial orientations. Even though contemporary text detectors regress tightly bounded bounding boxes or polygonal contours, cropped instances inevitably manifest canonical orientation offsets (specifically 0ยฐ, 90ยฐ, 180ยฐ, and 270ยฐ). Mainstream recognizers strictly assume horizontal reading order, leading to catastrophic degradation when presented with multi-oriented instances.
Existing approaches addressing multi-oriented text follow two paradigms: explicitly predicting orientation and rectifying text via spatial transformation modules (such as TPS-based ASTER, ESIR, or auxiliary direction classifiers), or performing multi-directional feature extraction and heavy rotational data augmentation (such as AON). Nonetheless, both strategies face three fundamental limitations: first, they lack theoretical guarantees, meaning any failure in the orientation perception network causes irreversible error propagation downstream; second, auxiliary rectification networks or multi-branch encoders impose substantial inference computational overhead and latency; third, heavy reliance on rotational data augmentation forces the model to expend representational capacity fitting multi-angle patterns, harming feature utilization efficiency on standard horizontal text.
Consequently, moving away from heuristic external rectification networks to establish an intrinsic, mathematically provable rotation-invariant mapping directly within the architecture is a promising angle of attack. The core idea is to decouple rotational symmetry across the pipeline by embedding rotation equivariance in visual feature extraction and proving and exploiting the intrinsic rotation invariance of cross-attention during sequence decoding, yielding an end-to-end Rotation-Invariant Scene TExt Recognition network (RISTER) with strict theoretical guarantees.
Method¶
Overall Architecture¶
RISTER follows the classic encoder-decoder paradigm while restructuring the symmetry properties of the feature-to-text mapping. Given an input image \(x\), the Rotation-Equivariant Local-Global Extractor (RELG) extracts a visual representation \(F\) that preserves local geometric details while modeling long-range character dependencies; the Rotation-Invariant Text Decoder (RITD) then autoregressively generates the target character sequence. Under arbitrary spatial rotations of the input image, visual features and cross-attention maps transform equivariantly, while the attended output vectors remain strictly identical, realizing rigorous end-to-end rotation invariance.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input image x<br/>(128ร128, multi-oriented)"] --> Stage1["Stages 1-3: Rotation-Equivariant Local Extraction<br/>F-Conv Local Blocks (C4 group) + Progressive Merging"]
Stage1 --> Stage2["Stage 4: Rotation-Equivariant Global Dependency<br/>Multi-Head Self-Attention (MHSA) for character context"]
Stage2 --> Feat["Equivariant visual features F<br/>(cรh/8รw/8, strictly equivariant under C4)"]
Feat --> CrossAttn["Cross-attention intrinsic rotation-invariant mapping<br/>Query fixed, Key/Value equivariant, attended Q' invariant"]
FixedQ["Decoded character history y_{0:t-1}<br/>Fixed Query Processing (Embedding+MHSA+MLP)"] --> CrossAttn
CrossAttn --> Out["Output classification & sequence prediction<br/>MLP projection + argmax iterative decoding"]
Key Designs¶
1. Rotation-Equivariant Local-Global Extractor (RELG): Preserving geometric details while modeling character context
Existing rotation-equivariant backbones (such as purely convolutional G-CNN, PDO-eConv, or transformer-based LieTransformer) fail to meet the dual requirements of STR: purely convolutional designs lack global receptive fields to capture inter-character linguistic dependencies, whereas equivariant vision transformers rely on early patch downsampling, discarding fine-grained stroke geometries. RELG decouples extraction into local and global phases. In the local phase, Local Blocks built with Fourier-based group-equivariant convolutions (F-Conv) over the discrete rotation group \(C_4\) perform feature extraction across three stages, coupled with stride-2 Merging blocks for progressive spatial downsampling (to 1/2, 1/4, and 1/8) and channel expansion (96, 192, and 384). In the global phase, taking advantage of the intrinsic rotation equivariance of multi-head self-attention layers, Global Blocks perform full-context interaction over high-level visual features, simultaneously preserving fine stroke details and establishing character context.
2. Provable intrinsic rotation invariance of cross-attention: Decoupling geometric coordinates from feature aggregation
Conventional attention decoders often degrade under rotated inputs, leading to the misconception that attention mechanisms are geometrically vulnerable. This paper identifies that standard cross-attention operates fundamentally over an unordered feature set and is inherently invariant to spatial permutations and rotations. When the query vector \(Q\) remains fixed and the source visual representation \(F\) undergoes a spatial rotation \(R_\theta\), the linearly projected keys \(K\), values \(V\), and spatial attention map \(A\) undergo identical equivariant spatial shifts. In continuous attention summation, this transformation amounts to a substitution of coordinates, guaranteeing that the attended contextual representation \(Q'\) remains strictly unchanged: $\(\mathcal{C}(Q, R_\theta(F)) \equiv \mathcal{C}(Q, F)\)$ This mathematical identity holds exactly at canonical \(C_4\) orientations (0ยฐ, 90ยฐ, 180ยฐ, and 270ยฐ) and holds to a high approximation under arbitrary continuous rotations, providing a theoretical foundation for orientation-agnostic text decoding without external rectification.
3. Rotation-Invariant Text Decoder (RITD): Autoregressive sequence generation with orientation-free queries
To ensure that the theoretical invariance of cross-attention propagates through the full decoding loop, the query tokens entering cross-attention must remain completely invariant to input image rotation. RITD enforces two guiding design principles: cross-attention must be used for multimodal interaction, and query tokens must remain independent of spatial image transformations. At decoding step \(t\), the previously decoded character sequence \(y_{0:t-1}\) (prefixed by start token \(y_0 = \text{<s>}\)) passes through token embedding, fixed multi-head self-attention, and an MLP to extract a language-driven query vector: $\(Q_{0:t-1} = \text{MLP}(\text{MHSA}(\text{Embedding}(y_{0:t-1})))\)$ The query vector \(Q_{t-1}\) then attends to the equivariant visual representation \(F\) via cross-attention, followed by a linear classification head predicting the next character: \(y_t = \text{argmax}(\text{MLP}(\mathcal{C}(Q_{t-1}, F)))\). Because the language prefix contains no image spatial coordinates, the invariance of the query coupled with the equivariance of the visual keys guarantees that the autoregressive prediction trajectory is identical before and after image rotation.
Loss & Training¶
The network is optimized end-to-end using the standard autoregressive sequence cross-entropy loss: $\(\mathcal{L}(\mathcal{W}) = -\mathbb{E}_{(x, y)\sim\mathcal{D}} \sum_{t=1}^{L} \log p_{\mathcal{W}}(y_t \mid y_{<t}, x)\)$ where \(\mathcal{W}\) denotes all trainable parameters and \(L=25\) is the maximum sequence length. Due to the built-in rotation symmetry of the network, RISTER requires no rotation-based data augmentation during training. Training exclusively on standard upright images naturally generalizes to arbitrary orientations, improving optimization efficiency and parameter utilization. The model is trained using the AdamW optimizer with a weight decay of 0.05, batch size of 512, and OneCycleLR scheduler over 20 epochs.
Key Experimental Results¶
Main Results¶
The model was comprehensively evaluated across 14 standard and multi-oriented STR benchmarks. Training was conducted exclusively on the real-world Union14M-Filter dataset (3.2M images) without synthetic rotated augmentations. Evaluation is reported using Word Accuracy Ignore Cases (WAIC). The table below summarizes performance across Common Benchmarks (CoB) and Union14M-Benchmarks (U14M-B):
| Benchmark Category | Dataset | Ours (RISTER-L) | Ours (RISTER-S) | Prev. SOTA (SVTRv2 / IGTR) | Gain / Advantage |
|---|---|---|---|---|---|
| Common (CoB) | IC13 | 98.9% | 98.4% | 98.7% (SVTRv2) | +0.2% |
| Common (CoB) | SVT | 98.1% | 97.8% | 98.4% (IGTR) | -0.3% |
| Common (CoB) | IIIT5k | 99.0% | 98.6% | 99.2% (SVTRv2) | -0.2% |
| Common (CoB) | IC15 | 90.6% | 90.0% | 91.0% (SVTRv2) | -0.4% |
| Common (CoB) | SVTP | 95.7% | 95.5% | 94.7% (IGTR) | +1.0% |
| Common (CoB) | CUTE80 | 97.6% | 97.2% | 99.0% (SVTRv2) | -1.4% |
| Multi-Oriented | Multi-Oriented (MO) | 96.6% | 96.0% | 92.6% (IGTR) / 89.5% (SVTRv2) | +4.0% |
| Challenging (U14M-B) | Curved (CUR) | 94.1% | 92.5% | 91.0% (SVTRv2) | +3.1% |
| Challenging (U14M-B) | Artistic (ART) | 79.4% | 76.7% | 78.7% (SVTRv2) | +0.7% |
| Challenging (U14M-B) | Contextless (CTL) | 80.6% | 77.8% | 81.3% (SVTRv2) | -0.7% |
| Challenging (U14M-B) | Salient (SAL) | 90.3% | 87.8% | 85.9% (SVTRv2) | +4.4% |
| Challenging (U14M-B) | Multi-Words (MTW) | 87.4% | 84.6% | 85.1% (SVTRv2) | +2.3% |
| General (U14M-B) | General (GEN 400k) | 83.7% | 81.9% | 82.3% (SVTRv2) | +1.4% |
| Overall Average | 13 Benchmark Avg. | 92.46% | 90.37% | 90.30% (SVTRv2) | +2.16% |
On dedicated oriented benchmarks (ASOT with 3,000 images and U14M-Multi-Oriented), compared to dedicated oriented recognizers (AON, SLOAN, ASTER, SVTRv2), RISTER-S achieves 96.0% on MO and 93.3% on ASOT, yielding an average of 94.65%, outperforming the second-best method SLOAN (90.65%) by 4.00% and SVTRv2 (88.00%) by 6.65%.
Ablation Study¶
To verify the individual contribution of each architectural component, rotation perturbation experiments across four canonical angles (0ยฐ, 90ยฐ, 180ยฐ, 270ยฐ) and encoder comparisons were conducted on IC13, SVT, and U14M-MO.
Table 1: Ablation on encoder-decoder components and rotation augmentation (IC13 Accuracy %)
| Model Configuration | 0ยฐ Acc | 90ยฐ Acc | 180ยฐ Acc | 270ยฐ Acc | Avg Acc | Note |
|---|---|---|---|---|---|---|
| (a) w/o Rot-E Encoder (standard convolutions) | 98.4 | 97.9 | 95.6 | 96.0 | 96.98 | Equivariance lost; accuracy drops at rotated angles |
| (b) w/o Rot-I Decoder (CTC decoder) | 98.0 | 91.2 | 95.6 | 90.5 | 93.82 | CTC relies on spatial direction; drops sharply |
| (c) Variant (a) + random rotation augmentation | 97.5 | 96.9 | 97.7 | 96.9 | 97.25 | Augmentation helps rotation but hurts 0ยฐ baseline |
| (d) Variant (b) + random rotation augmentation | 97.3 | 93.6 | 97.9 | 93.1 | 95.48 | CTC remains unable to fully fit orthogonal rotations |
| (e) RISTER-S (full model, no rotation aug) | 98.4 | 98.4 | 98.4 | 98.4 | 98.40 | Strict symmetry; 100% identical outputs at all canonical angles |
Table 2: Comparison of visual encoder architectures (Accuracy % & Parameter Counts)
| Encoder Backbone | IC13 (0ยฐ) | IC13 (90ยฐ) | SVT (0ยฐ) | SVT (90ยฐ) | U14M-MO | Params (M) | Architectural Properties |
|---|---|---|---|---|---|---|---|
| ResNet45 | 97.1 | 94.5 | 95.5 | 89.8 | 90.8 | 14.4 | No equivariance; severe drops under rotation |
| ResNet45 + G-CNN | 97.2 | 97.2 | 94.9 | 94.9 | 90.9 | 14.4 | Equivariant, but lacks long-range dependency |
| ResNet45 + F-Conv | 97.3 | 97.3 | 95.7 | 95.7 | 94.9 | 14.4 | Fourier filter group conv; solid local features |
| ViT (Vision Transformer) | 98.5 | 98.1 | 97.4 | 96.6 | 93.9 | 26.9 | Strong global context, lacks rotation equivariance |
| ViT + Stand-Alone Self-Attention | 98.3 | 98.3 | 97.1 | 97.1 | 94.5 | 27.2 | Equivariant attention, but early downsampling loses strokes |
| RELGโ (all local blocks, w/o self-attention) | 98.0 | 98.0 | 96.4 | 96.4 | 95.0 | 28.4 | Lacks self-attention context; drops 1.6% on MO |
| RELG (Ours: F-Conv + MHSA) | 98.6 | 98.6 | 98.3 | 98.3 | 96.6 | 32.8 | Combines progressive local downsampling and global character context |
Key Findings¶
- Verification of strict rotation invariance: Logit cosine similarity matrices demonstrate that complete RISTER produces exactly 1.00 similarity between predictions across 0ยฐ, 90ยฐ, 180ยฐ, and 270ยฐ, proving that output probabilities and confidence scores are identically invariant under canonical rotations, whereas omitting either component lowers similarity to 0.67โ0.89.
- Side effects of rotational data augmentation: Under fixed model capacity, training with random rotations forces the network to disperse capacity across orientations, degrading standard horizontal (0ยฐ) accuracy from 98.4% to 97.5%. RISTER avoids this compromise entirely by embedding architectural invariance.
- Zero inference overhead: F-Conv group-equivariant convolution weights can be collapsed into standard convolutional kernels post-training via cyclic shifting, incurring identical FLOPs (12.99 GFLOPs) and inference speed (31.4 FPS) compared to standard convolution baselines.
Highlights & Insights¶
- Provable rotation invariance of cross-attention: Proves that cross-attention operates on unordered sets; fixed queries guarantee identical aggregation regardless of spatial orientation in keys and values.
- Equivariant encoding coupled with invariant decoding: Maintains geometric correspondence in visual feature maps while eliminating orientation sensitivity in sequence generation, creating an elegant architectural symmetry.
- Zero-cost deployment via weight folding: Group-equivariant filters fold into standard convolution kernels during inference, combining strong inductive bias during training with standard inference latency.
Limitations & Future Work¶
- Square input resolution requirement: Strict \(C_4\) equivariance and cross-attention invariance currently require square feature maps (\(128\times 128\)), which may lead to suboptimal resolution efficiency for extreme aspect ratio text lines or long strings.
- Approximation at non-canonical angles: Arbitrary continuous angles (e.g., 37ยฐ, 65ยฐ) experience small discretization and interpolation errors on discrete pixel grids, transitioning strict invariance into close approximation. Extending to continuous Lie groups is a promising future direction.
Related Work & Insights¶
- vs ASTER / ESIR / SVTRv2: These methods introduce auxiliary rectification modules (e.g., Thin-Plate Spline), adding 15%โ30% computational cost and suffering from error compounding when orientation estimation fails. RISTER guarantees invariance intrinsically without extra modules.
- vs AON / SLOAN: AON uses 4-directional feature fusion causing semantic conflict and computation redundancy; SLOAN uses polar coordinate transforms that induce severe geometric distortions near the origin. RISTER operates directly on Cartesian coordinates with provable group equivariance.
Rating¶
- Novelty: โญโญโญโญโญ First theoretical proof of cross-attention rotation invariance in 2D features; elegant equivariant-invariant architecture.
- Experimental Thoroughness: โญโญโญโญโญ Evaluated across 14 benchmarks with extensive logit similarity matrices, four-angle rotations, and encoder ablations.
- Writing Quality: โญโญโญโญโญ Clear organization, mathematically sound formulations, and intuitive visualizations.
- Value: โญโญโญโญโญ Provides a benchmark paradigm for multi-oriented text recognition without external rectification or augmentation overhead.