Natural Image Pretraining Improves Abstract Reasoning¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: Conference PDF
Code: https://github.com/facebookresearch/mae
Area: Multimodal VLM / LLM Reasoning
Keywords: Abstract Visual Reasoning, ARC Benchmark, Masked Autoencoder, Visual Self-Supervised Pretraining, Test-Time Training
TL;DR¶
This paper introduces Nat-ARC, demonstrating for the first time that visual self-supervised pretraining via Masked Autoencoders (MAE) on natural images (ImageNet) transfers effectively to vision-centric solvers for few-shot abstract grid reasoning on the ARC benchmark, achieving 63.4% pass@2 with a single model and 70.2% pass@2 via multi-regime prediction-level ensembling.
Background & Motivation¶
The Abstraction and Reasoning Corpus (ARC) is a prominent benchmark designed to evaluate fluid intelligence in artificial systems. Each ARC task offers only 2 to 5 demonstration input-output grid pairs (at most 30×30 in size with up to 10 discrete colors), requiring the solver to induce a latent, compositional transformation rule under extreme data scarcity and apply it to unseen test inputs. For years, the ARC leaderboard has been heavily dominated by large language model (LLM) pipelines. These systems serialize 2D grids into textual or symbolic arrays and leverage the massive scale and code-synthesis priors of pretrained LLMs, coupled with chain-of-thought induction and program search, to hypothesize and verify underlying transformation rules.
However, ARC tasks are inherently visual grids, and human reasoners solve them using immediate perceptual intuitions—such as objectness, geometric symmetries, spatial relations, and foreground-background connectivity—rather than mechanical text token parsing. Recent vision-centric frameworks, such as VARC and LoopViT, have reformulated ARC as a conditional image-to-image translation task, demonstrating that end-to-end visual models can challenge LLM hegemony. Despite this promise, an acute pretraining asymmetry persists: while language-based solvers build upon foundation models pretrained on trillions of tokens, vision-centric solvers have predominantly been trained from scratch on small collections of ARC synthetic grids. This leaves visual solvers vulnerable to severe overfitting and poor parameter scaling in low-data regimes.
The central angle of attack in this work is to bridge this pretraining gap by asking whether the core asset of modern computer vision—unsupervised visual representations acquired from large-scale natural images—can bridge the apparent semantic divide and transfer directly to discrete, artificial grid reasoning. Core idea: incorporate natural-image Masked Autoencoder (MAE) self-supervised pretraining into a vision-centric ARC solver, leveraging frozen visual representations combined with lightweight test-time LoRA adaptation to supply robust perceptual primitives and geometric inductive biases for few-shot abstract reasoning.
Method¶
Overall Architecture¶
Nat-ARC extends the VARC conditional image-to-image translation paradigm into an end-to-end three-stage learning pipeline: unsupervised visual pretraining, offline multi-task ARC training, and per-task parameter-efficient test-time training (TTT). Input grids are mapped onto a unified 64×64 discrete canvas and tokenized into 2×2 visual patches. The encoder backbone is initialized from an ImageNet-pretrained MAE to endow the model with generic object and boundary awareness. During offline training, an encoder-decoder Transformer learns cross-task conditional transformations conditioned on a learnable task token. At inference time, the backbone parameters remain frozen while task-specific LoRA adapters and task embeddings are optimized on the demonstration pairs, culminating in prediction-level multi-view ensembling and exact-match majority voting.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
Input["Input: Few-Shot Grid Demonstrations & Test Query<br/>D_demo and x_infer"] --> Pretrain["Natural Image Self-Supervised Pretraining: MAE Initialization<br/>ImageNet-1K masked autoencoding provides perceptual primitives"]
Pretrain --> OfflineTrans["Offline ARC Cross-Task Training: Unified Canvas & 2D RoPE<br/>64x64 canvas color embedding + 2x2 patches + task conditioning"]
OfflineTrans --> AdaptLoRA["Test-Time Parameter-Efficient Adaptation: Frozen Backbone & LoRA<br/>Fine-tuning attention/MLP adapters and task embeddings avoids overfitting"]
AdaptLoRA --> EnsembleVote["Prediction-Level Cross-Regime Ensembling & Multi-View Voting<br/>Aggregating candidates from scratch/ImageNet/ARC-grid models via exact match"]
EnsembleVote --> Output["Output: Exact Test Grid Prediction<br/>pass@2 top supported modes"]
Key Designs¶
1. Natural Image Self-Supervised Pretraining: Decoupling Perception from Rule Induction Training a deep Vision Transformer from scratch solely on several thousand ARC demonstration pairs causes the network to overfit to superficial pixel co-occurrences rather than developing generalizable topological representations. Nat-ARC resolves this by initializing the encoder from official ViT checkpoints pretrained on ImageNet-1K using the Masked Autoencoder (MAE) objective with a 75% masking ratio. Because ImageNet images and patch tokenizations (typically 224×224 images with 16×16 patches) differ structurally from the discrete ARC canvas, the original linear patch projection and absolute positional embeddings are discarded, retaining solely the bidirectional Transformer encoder blocks. This pretraining endows the encoder with emergent object segmentation and foreground-background sensitivity before any ARC exposure, cleanly decoupling visual perceptual priors from downstream abstract rule induction.
2. Unified Canvas and 2D Spatial Positional Encoding: Bridging Discrete Grids to Visual Features ARC tasks present varying grid dimensions (ranging from 1×1 to 30×30) and 10 categorical color values. To handle these uniformly, Nat-ARC places all grids centrally onto a fixed 64×64 canvas and maps discrete color indices \(\{0, \dots, 9\}\) into high-dimensional embedding vectors via a learnable lookup table. The canvas is partitioned into non-overlapping \(2\times 2\) patches, yielding a sequence of 1024 visual tokens prepended by a task-conditioning token. To preserve spatial coordinates and translational invariance across the canvas, Nat-ARC equips every self-attention layer in both the encoder and decoder with 2D Rotary Position Embeddings (2D RoPE):
The orthogonal 2D rotation matrix \(\mathbf{R}_{\Theta, m-n}^{\text{2D}}\) decomposes relative spatial displacements along horizontal and vertical axes independently, allowing attention heads to compute exact geometric relative offsets, reflections, and spatial bounds on the discrete canvas.
3. Frozen-Backbone Test-Time Adaptation: Lightweight LoRA and Task Embedding Optimization During evaluation on an unseen task \(T_{\text{eval}}\), only 2 to 4 demonstration pairs are provided; fully fine-tuning a model of hundreds of millions of parameters under such data scarcity inevitably results in catastrophic forgetting and severe overfitting. Nat-ARC adopts a parameter-efficient test-time adaptation strategy: the offline-trained encoder and decoder backbones are completely frozen. LoRA adapters are injected into all attention projections (\(W_q, W_k, W_v, W_o\)) and MLP layers, and optimized alongside freshly initialized task embeddings for the target task and its auxiliary symmetry transformations (rotations, reflections, color permutations). Restricting updates to a low-rank subspace allows the model to rapidly assimilate task-specific rules from demonstration examples while safeguarding the general visual inductive priors stored in the pretrained backbone.
4. Prediction-Level Multi-Regime Ensembling and Multi-View Voting: Robust Solution Aggregation Single-shot inference on discrete grids is susceptible to minor boundary translation slips or isolated color pixel misclassifications, causing failures under strict exact-match evaluation. Nat-ARC implements task-agnostic test-time data augmentations, sampling diverse canvas translations, scales, and rotations to produce a diverse pool of candidate grid predictions \(\{\hat{y}^{(v)}\}\) mapped back to the canonical frame. Rather than soft-voting in continuous logit space, the system performs exact-match majority voting across discrete grid outputs. Crucially, this mechanism seamlessly extends across different training regimes: pooling candidate outputs from randomly initialized models, ImageNet MAE pretrained models, and in-domain ARC-style grid MAE models combines complementary visual inductive biases, yielding significant coverage improvements on difficult reasoning instances.
Loss & Training¶
During offline training, the model is trained jointly across 400 ARC-1 training tasks and 400k synthetic examples from RE-ARC using a pixel-wise cross-entropy loss over the 10 discrete colors:
Training runs for 200 epochs with extensive resolution and color jitter augmentations. In the test-time training (TTT) stage, the identical objective is optimized over the few demonstration pairs \(D_{\text{demo}}\) and auxiliary transformed tasks for several hundred steps to converge the LoRA adapters and task embeddings.
Key Experimental Results¶
Main Results¶
The table below compares Nat-ARC with representative human benchmarks, general-purpose LLMs, specialized language-based ARC systems, and non-pretrained vision/neural solvers on the official ARC-1 and ARC-2 evaluation benchmarks under the pass@2 exact-match metric.
| System Family / Model | # Params | Pretraining Data Source | ARC-1 (pass@2, %) | ARC-2 (pass@2, %) |
|---|---|---|---|---|
| Human Average (Avg. Human) | – | Human visual life experience | 60.2 | – |
| Human Best (Best Human) | – | Expert competitive reasoning | 98.0 | 100.0 |
| DeepSeek-R1 | 671B | Language corpus + RL reasoning | 15.8 | 1.3 |
| GPT-5 | – | Massive multimodal/language web data | 44.0 | 1.9 |
| Gemini 3 Pro | – | Frontier proprietary multimodal data | 98.0 | 77.1 |
| TTT + BARC | 16B | Language pretraining + ARC synthesis | 62.8 | – |
| The ARChitects | 8B | Language pretraining + ensemble | 71.6 | – |
| HRM (Hierarchical Reasoning Model) | 27M | Trained from scratch (No pretrain) | 40.3 | 5.0 |
| TRM (Tiny Recursive Model) | 7M | Trained from scratch (No pretrain) | 44.6 | 7.8 |
| VARC (Vision Baseline) | 19M | Trained from scratch (No pretrain) | 54.0 | 8.3 |
| LoopViT (Looped Transformer) | 18M | Trained from scratch (No pretrain) | 65.8 | 14.2 |
| Loop-OWM (Composable World Model) | 10.6M | Trained from scratch (No pretrain) | 68.5 | 22.5 |
| Nat-ARC (Single Model, Ours) | 0.6B | ImageNet-1K MAE Pretraining | 63.4 ± 0.7 | 8.6 ± 0.8 |
| Nat-ARC (Ensemble, Ours) | 2.0B | ImageNet + ARC-Grid MAE | 70.2 ± 0.6 | 13.1 ± 0.7 |
Ablation Study¶
To isolate stochastic evaluation fluctuations from true representation gains, the authors conducted controlled ablations across model scales (Base, Large, Huge) and pretraining regimes, reporting statistics over 5 independent offline checkpoints × 2 independent TTT runs.
| Pretraining Regime | Model Scale | Encoder Specs (depth / dim / heads) | Total Params | Offline Optimization Benefit | ARC-1 (pass@2, %) |
|---|---|---|---|---|---|
| No Pretraining | Base | 12 / 768 / 12 | 114M | Slow convergence baseline | ~55.2 ± 0.9 |
| No Pretraining | Large | 24 / 1024 / 16 | 334M | Moderate improvement | ~58.4 ± 0.8 |
| No Pretraining | Huge | 32 / 1280 / 16 | 664M | Capacity overfitting degradation | 57.6 ± 0.8 |
| ImageNet MAE Pretrain | Base | 12 / 768 / 12 | 114M | High early-epoch accuracy | 58.7 ± 0.6 |
| ImageNet MAE Pretrain | Large | 24 / 1024 / 16 | 334M | Accelerated convergence | 61.2 ± 0.7 |
| ImageNet MAE Pretrain | Huge | 32 / 1280 / 16 | 664M | Consistent monotonic scaling | 63.4 ± 0.7 |
| ARC-Grid MAE Pretrain (Single) | Huge | 32 / 1280 / 16 | 664M | In-domain synthetic optimal | ~64.8 |
| Ensemble: 2× No Pretraining | Huge | 2× Independent seed models | ~1.3B | Limited checkpoint diversity | ~61.5 ± 0.7 |
| Ensemble: 2× ImageNet Pretrain | Huge | 2× Independent MAE models | ~1.3B | High-confidence overlap | ~66.8 ± 0.6 |
| Ensemble: Scratch + ImageNet | Huge | Pooled heterogeneous candidates | ~1.3B | Strong representation synergy | ~66.5 ± 0.5 |
| Full Ensemble (Nat-ARC) | Huge | Scratch + ImageNet + ARC-Grid | ~2.0B | Prediction-level voting | 70.2 ± 0.6 |
Key Findings¶
- Overcoming Low-Data Scaling Degradation: For randomly initialized visual models, increasing model capacity from Large to Huge leads to a performance drop (from ~58.4% down to 57.6%), indicating severe overfitting in data-constrained abstract reasoning. In contrast, ImageNet MAE pretraining unlocks reliable scaling, steadily rising from 58.7% (Base) to 61.2% (Large) and 63.4% (Huge).
- Emergent Foreground-Background Grouping in Attention: Layer-wise attention maps reveal that at initialization (Epoch 0 before any ARC training), randomly initialized models exhibit completely diffuse attention. Meanwhile, ImageNet MAE pretrained models immediately exhibit concentrated attention on foreground visual structures, accelerating convergence across all training stages.
- Task Types Most Benefited by Natural Pretraining: On tasks where ImageNet-pretrained models dramatically outperform scratch baselines (improving solve rates from \(\le 30\%\) to \(\ge 70\%\)), 12 of the 15 tasks fall squarely into two classic perceptual categories: 6 "match-and-copy" tasks (identifying an exemplar pattern and relocating it) and 6 "connected-component reasoning" tasks (boundary detection, region filling, and flood-filling).
Highlights & Insights¶
- Unified Perceptual Substrate across Natural and Abstract Domains: The paper provides concrete empirical evidence that representations learned from natural scene photography can transfer directly to discrete symbolic grid reasoning, demonstrating that visual primitives like object boundary segregation and patch matching are universal across visual domains.
- Cross-Domain ImageNet Transformation Proof-of-Concept: By training a discrete autoencoder that projects natural ImageNet images into a \(60\times 60\), 10-color ARC-like latent grid, the authors show that Nat-ARC can execute abstract operations (e.g., 180° rotation, boundary framing) directly on natural photo latents, decoding back into coherent rotated or framed photographs.
- Vision-Centric ARC Surpasses Human Average: Achieving 70.2% on ARC-1 demonstrates that purely vision-based models, when paired with self-supervised pretraining and test-time voting, can decisively outperform average human performance (60.2%) without relying on symbolic interpreters or LLMs.
Limitations & Future Work¶
- Significant Performance Gap on ARC-2: On the more complex ARC-2 benchmark featuring longer compositional logic and novel rules, Nat-ARC achieves only 8.6% (single model) and 13.1% (ensemble), lagging far behind human experts (100%) and top-tier frontier multimodal LLMs (Gemini 3 Pro at 77.1%).
- Heavily Dependent on Data Augmentation and Test-Time Fine-Tuning: The solver requires 400k synthetic offline grids and hundreds of gradient steps per test task, indicating that the visual representations do not yet support zero-shot in-context rule deduction purely via forward-pass attention.
- Future Directions: Fusing natural-image self-supervised pretraining with recurrent/looped visual architectures (e.g., LoopViT) to enable test-time iterative computation and multi-step latent state exploration.
Related Work & Insights¶
- vs VARC (Hu et al., 2025): VARC introduced the vision-centric conditional translation formulation but remained restricted to a 19M scratch-trained model; Nat-ARC resolves the pretraining asymmetry, scaling up to 664M parameters and elevating single-model accuracy from 54.0% to 63.4%.
- vs LoopViT (Shu et al., 2026) / Loop-OWM (Gao et al., 2026): While LoopViT and Loop-OWM focus on recurrence and world-model dynamics during inference, Nat-ARC focuses on foundational visual pretraining; combining recurrent architectures with ImageNet MAE initialization represents an immediate and promising next step.
- vs Language-Based ARC Solvers (BARC / The ARChitects): Language pipelines serialize 2D grids into code/text tokens and execute compute-heavy program search; Nat-ARC proves that a 0.6B vision transformer can achieve competitive results purely through visual feedforward representations.
Rating¶
- Novelty: ⭐⭐⭐⭐ [First systematic proof that natural-image MAE pretraining transfers effectively to abstract grid reasoning]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive 5-seed repetitions, multi-scale ablations, attention analyses, and cross-domain natural image experiments]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, rigorous prose, elegant visual taxonomy, and transparent reporting of results]
- Value: ⭐⭐⭐⭐⭐ [Provides a crucial foundation for vision-centric AGI research, proving that abstract reasoning does not belong exclusively to language models]