Skip to content

Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

Conference: NeurIPS2026
arXiv: 2609.35473
Area: Self-Supervised Learning
Keywords: handwritten text recognition, pixel-variance subspace, masked image modeling, frozen encoder, character error rate

TL;DR

Using pixel-PCA projections and frozen-encoder probes, this paper studies the distribution of handwritten text recognition signals: high-variance pixel directions support transcription better on six Latin-script benchmarks, while real-data-pretrained MAE achieves 5.5% mean probe CER and 4.5% mean CER after adding a language decoder and full fine-tuning.

Background & Motivation

Handwritten Text Recognition (HTR) converts arrangements of strokes into character sequences. Differences in fonts, ligatures, slant, paper, and historical period make labeled data insufficiently representative; rendering synthetic text is inexpensive but does not eliminate the synthetic-to-real handwriting gap. The issue is therefore not simply decoder capacity, but whether the encoder preserves strokes and character-position relationships that a readout can use. Existing text self-supervised learning methods use pixel reconstruction, contrastive learning, and sequence-specific pretext tasks, yet different datasets and probes across papers make it difficult to attribute improvements to the objective rather than the training recipe.

One explanation from natural-image classification holds that reconstruction losses emphasize high-variance pixels whose variations need not carry category semantics, thereby wasting representation capacity. Handwritten transcription may have the opposite structure: dark strokes against a light background account for major pixel changes, and character distinctions depend on those strokes. Evidence from classification cannot be generalized directly to every visual task; the input directions supporting recognition must first be examined before comparing pretraining objectives.

The paper consequently follows an analytical rather than a new-network approach: it first compares the recognizability of high- and low-variance PCA reconstructions, then independently trains six methods on a shared visual backbone and tests its prediction through frozen probes, label efficiency, and systems with a language decoder. Core idea: whether pixels should be reconstructed depends on which pixel-variance directions carry the task signal; for the handwritten line images studied here, preserving high-variance stroke content supports transcription better than ignoring those details.

Method

Overall Architecture

The study contains two complementary experimental tracks, not an architecture chaining MAE, JEPA, and MoCo. The first examines the input directly: training-image pixel PCA → high-/low-variance image reconstruction → separately trained matching BiLSTM–CTC probes → test-CER comparison. The second examines representations: MAE, SimMIM, I-JEPA, V-JEPA-2, MoCo-v3, and SigLIP are independently pretrained on synthetic or real handwriting, then frozen for readout training and measurement of their linear relationship with pixel subspaces.

The subsequent language-model experiment is a separate downstream evaluation, not another layer added to the probes. The shared visual encoder produces position features, a projector turns them into an image prefix for the language model, and the decoder generates the transcript autoregressively. The frozen configuration trains only downstream components; the full configuration also updates the visual encoder. No network diagram is included, to avoid presenting parallel comparisons and statistical analyses as a new architecture.

Key Designs

1. Matched-variance projection probes: locate usable recognition information first

Images are converted to grayscale and placed on a canvas of height 64 and width 1024, yielding 65,536 dimensions after flattening. The PCA mean and basis are estimated exclusively from training images; validation and test images do not enter the fit. With eigenvalues sorted in descending order, components are accumulated from either end until the smallest set covering at least the target fraction of total variance is reached. The comparison therefore matches variance budgets, not subspace dimensions. The high-variance side typically needs fewer components, whereas the low-variance side needs more.

\[ k(p)=\min\left\{K:\sum_{j=1}^{K}\lambda_j\ge p\sum_{j=1}^{d}\lambda_j\right\},\qquad k'(p)=\min\left\{K:\sum_{j=d-K+1}^{d}\lambda_j\ge p\sum_{j=1}^{d}\lambda_j\right\}. \]

“At least” matters here: discrete eigenvalues can make both sides overshoot the threshold, so the construction does not guarantee exactly equal energy. At large fractions, the two component sets can also overlap; top and bottom cannot always be interpreted as complementary, disjoint information partitions. The training mean is added back, reconstruction values are clipped to the valid grayscale range, and probes trained with real transcript supervision test recognizability.

\[ x_p^{\mathrm{top}}=\operatorname{clip}_{[0,1]}\left(\mu+\Pi_{V_{\mathrm{top}}(p)}(x-\mu)\right),\qquad x_p^{\mathrm{bot}}=\operatorname{clip}_{[0,1]}\left(\mu+\Pi_{V_{\mathrm{bot}}(p)}(x-\mu)\right). \]

The probe thus receives images after projection, mean restoration, and nonlinear clipping, rather than pure linear subspace coordinates. The experiment supports a difference in decodability under this reconstruction and training protocol, but does not prove that every low-variance direction contains no label information. The appendix uses a Gram matrix to reduce computation, yet when the training sample count is below the pixel dimension, the centered training matrix has rank at most the sample count minus one. Gram decomposition recovers pixel directions associated with the nonzero spectrum, not a unique complete null-space basis. The paper's claim to obtain a complete orthonormal basis lacks null-space completion details, which particularly need clarification when reproducing the bottom construction.

2. Parallel pretraining-objective comparisons: preserve pixels, predict representations, or learn invariance

The six methods share a ViT-B-scale backbone, reported as 113.55M parameters in the main text. Each input strip spans the full height and 4 pixels in width, producing 256 position tokens per line; the encoder uses 1D RoPE and no class token. MAE randomly masks 75% of strips, encodes only the visible subset, and predicts masked raw pixels through a small four-layer Transformer decoder using MSE. SimMIM masks 60%, retains mask tokens at the encoder input, and reconstructs pixels with a linear head and L1 loss. Both directly supervise stroke pixels, but their encoding costs, masking ratios, losses, and decoders differ.

I-JEPA samples four target spans along the strip sequence, each covering 0.15–0.25 of the line; context is the complement of the union of these spans. The online encoder and predictor use context to predict EMA-teacher target features, which are normalized across the embedding dimensions at each position before L1 matching. V-JEPA-2 uses the same static 1D strip input: it does not actually turn handwriting into video frames. It concatenates the teacher's last four layers, deepens the predictor, and adds a context-side auxiliary loss. Target normalization prevents collapse in the authors' experiments, but these modifications mean that the evaluated methods are the paper's HTR adaptations.

The appendix's MoCo-v3 implementation is not purely global image contrast: each of two augmented views is divided into four horizontal windows, corresponding windows form positives, and other batch windows form negatives for symmetric InfoNCE. SigLIP instead mean-pools visual features and contrasts them against transcripts encoded by a six-layer text encoder using a sigmoid image–text objective. It explicitly uses ground-truth transcript labels and therefore is not label-free visual self-supervised learning. The comparison evaluates these implementations, but does not show that all sequence-contrastive methods must lose to reconstruction or that differences arise solely from contact with pixels.

3. Frozen-encoder readouts: separate positional readability from readout compensation

Linear–CTC maps each frozen position feature directly to character logits, with no recurrent or attention-based readout. BiLSTM–CTC adds one bidirectional LSTM layer with hidden size 256 before the same character prediction, allowing integration of neighboring positions. Both probes share 94 characters plus a CTC blank, train with real training-set transcripts, and use greedy CTC decoding at test time. The rescue gap is defined as \(\Delta=\mathrm{CER}_{\mathrm{Linear}}-\mathrm{CER}_{\mathrm{BiLSTM}}\), measured in percentage points.

A smaller gap indicates that a shallow linear readout can use the features relatively easily, whereas a larger gap indicates stronger compensation from the sequence head; it is not direct proof that the encoder contains no character information. MAE's mean gap is 5.7 pp, substantially below JEPA and contrastive methods, but it is not zero and does not imply that sequence modeling is unnecessary. A linear readout has no sequence model of its own, but visual-encoder features may still contain cross-position context.

4. Subspace–representation alignment: connect geometry and recognition through linear predictability

For each frozen encoder and dataset, separate linear regressors predict image-level encoder features from high- and low-variance pixel projections, using the coefficient of determination to measure predictability. A positive difference means that high-variance directions explain these features more easily than low-variance directions do. The appendix's threshold-averaging description yields the following compact summary expression; the main text's single-threshold equation must be read together with the figure caption's aggregation description.

\[ \overline{G}(\mathcal E,\mathcal D)=\frac{1}{|P|}\sum_{p\in P}\left(R^2_{\mathrm{top},p}(\mathcal E,\mathcal D)-R^2_{\mathrm{bot},p}(\mathcal E,\mathcal D)\right),\qquad P=\{0.10,0.25,0.50,0.75,0.90\}. \]

Figure 4 compares different datasets within each SSL method and reports the relationship between this gap and CER. “Within a method” therefore means an observation across datasets for a fixed method, not an intervention controlling other factors on a fixed dataset. Dataset difficulty, training-set size, and pixel spectra can jointly affect both quantities. The finding is an informative empirical association, not causal proof that high-variance alignment lowers CER, and not a theorem of Bayes sufficiency. The cache does not expose directly readable numerical correlation coefficients, so no figure values are invented here.

Loss & Training

The shared pretraining recipe uses AdamW, 10,000 warmup steps, cosine decay, and 500 epochs of 2,000 steps, approximately one million updates, with effective base batch size 32. Peak learning rates are \(3\times10^{-4}\) for MIM and JEPA and \(6\times10^{-4}\) for SigLIP. MAE computes MSE only on unnormalized pixels at masked positions, SimMIM uses L1, JEPA uses L1 against teacher features, and the remaining methods use their contrastive objectives. A shared backbone and similar schedule do not imply exactly matched computation.

SSL checkpoints are selected through an inline CTC probe trained on labeled IAM features. Although most visual pretraining losses do not use transcripts, model selection still introduces labels, so the complete pipeline is not fully unsupervised. Downstream CTC probes use Adam with learning rate \(10^{-3}\), batch size 64, and 100 epochs, selecting the checkpoint with the best validation CER. Label-efficiency experiments use nested labeled subsets of 1%, 10%, 25%, 50%, and 100%.

The language decoder is not an existing generic giant LLM: it is a causal model pretrained from scratch on five-language CC100 text, described as approximately 200M parameters in the main text. Stage 1 freezes the visual encoder and language model and trains only the projector on synthetic image–transcript pairs. Stage 2 keeps the visual encoder frozen and trains the projector and language model on each real training set. Stage 3 unfreezes the entire system at lower learning rates. Results must therefore be interpreted as frozen configuration “1+2” and full configuration “1+2+3,” not as a frozen configuration containing only Stage 1.

The main-text overview abbreviates frozen and full configurations as Stage 1 and Stage 2, while the detailed method and appendix use the three-stage numbering above. The main text also specifies a linear projector and approximately 200M decoder, whereas Appendix F describes a two-layer MLP, approximately 7.3M projector, and approximately 218M decoder. These are differences in source configuration descriptions, not details that can be silently merged into a verified exact implementation. Language-model pretraining and synthetic transcript supervision introduce an additional language prior, so language-system results must be interpreted separately from visual probes.

Key Experimental Results

Main Results

The metric is character error rate (CER; lower is better). The appendix defines it as prediction–reference Levenshtein distance divided by reference length, averaged uniformly over test examples, without case or punctuation normalization. The following results are from main-text Table 1: real-pretrained encoders are frozen and the readout is BiLSTM–CTC. The last column is the mean CER difference between linear and BiLSTM readouts.

Method IAM Rimes Bentham LAM Rodrigo Parzival Mean CER (%) Mean rescue gap (pp)
MAE 9.9 6.5 7.5 4.4 2.0 2.5 5.5 5.7
SimMIM 14.5 11.1 12.0 7.0 3.0 4.3 8.7 10.3
I-JEPA 23.2 16.9 22.1 13.6 6.0 6.5 14.7 39.0
V-JEPA-2 17.9 13.2 15.5 8.9 3.8 4.1 10.6 26.6
SigLIP 25.1 16.7 20.8 11.3 5.6 7.3 14.5 29.5
MoCo-v3 32.9 26.1 26.0 15.0 9.0 5.8 19.2 59.1

The arithmetic mean of MAE's six displayed dataset values is approximately 5.47%, consistent with the rounded 5.5% entry rather than a mean-value conflict. MAE is best in all six columns, but saying that every pixel-family member beats every other family is too strong: on Parzival, V-JEPA-2's 4.1% beats SimMIM's 4.3%. Likewise, MAE does not have the smallest rescue gap on every dataset: on IAM its gap is 8.8 pp, compared with SimMIM's 8.3 pp.

The following results are from main-text Table 2. SSL rows use real-handwriting pretraining followed by the language-decoder full-fine-tuning configuration; supervised baselines use the authors' synthetic-training and real-fine-tuning schedule. This is a comparison under the paper's reproduced experimental setting, not a unified leaderboard covering all public systems.

Method / Config IAM Rimes Bentham LAM Rodrigo Parzival Mean CER (%)
CRNN 7.4 4.9 9.6 5.6 2.8 2.8 5.5
DTrOCR-B 7.6 4.6 8.3 4.2 2.6 6.6 5.6
TrOCR-B 6.3 3.8 6.9 3.5 2.2 5.5 4.7
No SSL, random initialization 7.0 5.1 7.8 6.7 2.4 5.1 5.7
MAE, full fine-tuning 6.9 4.2 5.6 4.0 1.9 4.6 4.5
SimMIM, full fine-tuning 12.7 6.8 10.8 6.7 3.5 8.2 8.1
V-JEPA-2, full fine-tuning 14.3 7.8 13.7 8.3 4.1 8.5 9.5
MoCo-v3, full fine-tuning 13.9 7.5 12.0 5.8 3.4 7.9 8.4

MAE's frozen configuration has 5.1% mean CER, decreasing to 4.5% with full fine-tuning, a gain of 0.6 pp. Its full mean is 0.2 pp below TrOCR-B's 4.7%, but it still loses to TrOCR-B on IAM, Rimes, and LAM, and to CRNN on Parzival. Table 2's fully fine-tuned MoCo-v3 mean of 8.4% also beats V-JEPA-2's 9.5%; the complete visual-probe ordering cannot be claimed to remain unchanged with a language decoder. TrOCR-B additionally starts from approximately 684M printed-text pretraining images, whereas the other supervised baselines start from scratch.

Ablation Study

The following table summarizes synthetic-to-real pretraining changes from Figure 3 and the main text. The CER difference is real minus synthetic; negative values indicate improvement. These compare pretraining distributions, rather than strictly ablating one module.

Method Real-pretraining mean CER (%) Synthetic→real CER difference (pp) Interpretation boundary
MAE 5.5 -2.9 Improvement for this reconstruction recipe
SimMIM 8.7 -3.8 Improvement for this reconstruction recipe
V-JEPA-2 10.6 +1.9 Degradation for this static-strip adaptation
SigLIP 14.5 +1.7 Image–text comparison using transcript supervision
I-JEPA 14.7 +4.6 Degradation for this 1D adaptation
MoCo-v3 19.2 +5.2 Degradation for this window-contrastive recipe

Key Findings

  • The PCA variance sweep uses 0.05, 0.10, 0.20, 0.30, 0.50, 0.70, and 0.90; qualitative reconstructions use 0.80. The figures support better recognizability of top reconstructions, but the cache does not expose every point numerically, so an exact CER table cannot be constructed from the descriptions.
  • With only 10% of labels, MAE achieves 8.8% mean CER, below the strongest non-pixel method's 10.6% with all labels. It wins 29 of 30 dataset–label-budget settings; the exception is Parzival at 1%, where I-JEPA achieves 93.8% and MAE 95.6%.
  • SimMIM ranks second on the mean from 10% labels onward, not on every individual dataset: V-JEPA-2 ranks second on Parzival at 25%, 50%, and 100% labels.
  • Large mean rescue-gap differences show that readout capacity changes method comparisons, but MAE's 5.7 pp is still a substantive gain; it does not eliminate the need for readout compensation.

Highlights & Insights

  • Testing a task's pixel-statistical structure before interpreting its pretraining results is more explanatory than selecting a method from a final leaderboard alone. The reusable contribution is the combination of post-projection recognition, frozen readouts, and representation regression, not an assumption that all text tasks share the same spectral structure.
  • Reporting both linear and sequence probes helps distinguish character-aligned representations from representations compensated by downstream networks. A strong decoder's result alone can obscure that distinction.
  • A language model should not be treated as a free repair mechanism for visual representations. Frozen versus full configurations provide practical guidance, but their gains still involve CC100 language priors and transcript supervision.

Limitations & Future Work

  • The authors mainly test one ViT-B scale and one random seed per experiment, with five languages but only Latin scripts. Arabic, Chinese, Devanagari, scene text, music notation, and mathematical expressions remain untested; small mean differences particularly need multi-seed confidence intervals.
  • PCA reconstructions are not dimension-matched, discrete thresholds do not guarantee exact energy matching, large budgets can produce overlapping subspaces, mean restoration and clipping modify the input, and null-space handling for low-rank training data is incompletely specified. Actual retained variance, dimension, overlap, and unclipped controls should be reported before making broader information claims.
  • Methods differ in masks, losses, prediction heads, teacher targets, and computational cost. Window-level MoCo is only one sequence recipe, and SigLIP uses transcripts. Cross-dataset geometric correlations cannot rule out dataset-difficulty confounding and do not establish causality.
  • Unresolved source conflicts affect data configuration: the main text and Appendix C specify 12.5M synthetic lines, 2.5M per language, while Appendix B specifies 10M, approximately 2M per language. The main text describes real pretraining on the union of six training sets, whereas Appendix B adds Saint-Gall and Washington.
  • Augmentation descriptions also conflict: an earlier Appendix B passage says synthetic images are not augmented, while a later passage says both pretraining distributions use identical augmentation. The main text gives approximately 3,000 fonts and Appendix C approximately one thousand; Figure 6 mentions CulturaX and Gutenberg, while the prose mainly specifies Gutenberg. Code or author clarification is needed rather than silently choosing a version for reproduction.
  • The main text calls V-JEPA-2 inputs frames, while the appendix explicitly denies video processing. JEPA EMA momentum appears as 0.9999 and 0.999 in different passages, and target spans are allowed to overlap despite a separate “disjoint” description. This note follows the explicit static-strip mechanism while retaining those implementation uncertainties.
  • vs Balestriero and LeCun's reconstruction analysis: The paper adopts pixel-subspace projections but observes a handwritten-transcription pattern different from natural-image classification. The lesson is that the relationship between reconstruction and semantics depends on the task distribution, not that one loss is intrinsically suitable or unsuitable for all vision tasks.
  • vs MaskOCR, Text-DIAE, and DualMAE: Earlier work already uses reconstruction for text; this paper primarily adds matched cross-objective comparisons and a geometric explanation. It does not reproduce these specialized methods individually, so its six representative implementations cannot establish a ranking for the entire text-SSL literature.
  • vs SeqCLR, ChaCo, and RCLSTR: These methods change contrastive units or textual relationships. Although the paper's MoCo already uses window positives, it does not cover character- and relation-level designs. Further experiments should compare contrastive granularity under fixed data and compute rather than declare contrastive learning universally ineffective.
  • vs TrOCR and DTrOCR: The paper connects SSL visual representations to a medium-scale language decoder pretrained from scratch to assess system-level gains. Reuse should account separately for unlabeled visual data, supervised synthetic data, model-selection labels, and language corpora rather than summarize the result as an unsupervised improvement.

Rating

  • Novelty: 4/5 — The connection between task-relevant pixel-variance structure and HTR pretraining choices offers a clear analytical perspective.
  • Experimental Thoroughness: 4/5 — Six methods, six benchmarks, multiple readouts, and label budgets provide broad coverage, but single-seed experiments and configuration inconsistencies weaken the conclusions.
  • Writing Quality: 3/5 — The main argument is clear, but data sizes, augmentation, stage numbering, and implementation details need clarification.
  • Value: 4/5 — Practical evidence for choosing handwriting representations, not a universal conclusion about every writing system or SSL family.