Skip to content

On Locality and Length-Generalization in Visual Reasoning

Conference: ECCV 2026
Paper: ECCV Official
PDF: EventHosts
Area: Multimodal VLM / LLM Reasoning
Keywords: Visual Reasoning, Length Generalization, Locality of Perception, Recurrent Attention, State Tracking

TL;DR

This paper demonstrates that contemporary vision models fail at out-of-distribution length generalization due to shortcut learning encouraged by single-shot global perception, and establishes that strictly local foveated perception coupled with nonlinear recurrence is both necessary and sufficient for robust visual state tracking.

Background & Motivation

Contemporary state-of-the-art vision foundations and vision-language models process visual inputs by ingesting an entire high-resolution image in a single feed-forward pass, tokenizing spatial patches into a flat sequence and deploying global self-attention across the whole context. This standard paradigm stands in stark contrast to the human visual system, which perceives complex scenes through a trajectory of discrete, localized foveal glimpses orchestrated via rapid saccadic eye movements. Recent investigations in autoregressive language models have uncovered severe failures in length generalizationโ€”such as evaluating parity over sequences longer than those seen during trainingโ€”primarily because global attention mechanisms encourage models to exploit superficial statistical shortcuts instead of executing true inductive, step-by-step symbolic computation.

This deficiency is substantially more pernicious in the two-dimensional visual domain. Unlike textual reasoning where discrete tokens arrive in an explicit linear order, visual reasoning requires aggregating information that is spatially distributed across an uncurated 2D canvas. The agent must concurrently solve two fundamentally intertwined challenges: gathering required local visual clues across arbitrary spatial layouts, and updating internal system states iteratively. While nonlinear recurrent architectures have demonstrated the capacity for algorithmic length extrapolation in discrete text, a critical question remains unanswered in vision: does recurrence alone suffice when an agent is exposed to the full 2D scene, or does global perception inherently trigger shortcut learning that collapses out-of-distribution?

Through a controlled suite of visual algorithmic benchmarks, this work demonstrates that providing recurrent models with global access still causes complete catastrophic failure under distributional shifts in task length and spatial resolution. The core tension is that global views inevitably furnish spurious cross-entity correlations, disincentivizing models from learning strict step-by-step traversal policies. Core idea: by constructing a recurrent foveated agent (FoveAgent-LSTM) that restricts perception to high-resolution local glimpses guided by low-resolution downsampled peripheral context, the model is physically prevented from learning global shortcuts, thereby recovering robust out-of-distribution length generalization across unseen problem lengths and visual resolutions.

Method

Overall Architecture

The framework introduces a test-bed of controlled visual algorithmic tasks alongside a dual-scale recurrent agent, FoveAgent-LSTM. Given an input image depicting an underlying dynamical system (e.g., Visual Parity, State Machine, or Finding Roots), the model executes active visual search across discrete time steps \(t\). At each step, conditioned on its current hidden state \(h_t\) and spatial coordinate \(x_t\), the agent ingests two concentric visual crops: a high-resolution central foveated glimpse \(G_l\) capturing fine-grained local state, and a peripheral glimpse \(G_p\) covering a \(4\times\) wider spatial area that is strictly downsampled to the same sensor resolution.

A convolutional visual encoder encodes both glimpses, concatenating their representations before feeding them into an LSTM recurrent backbone. The recurrent state updates the evolving task representation and predicts next-step spatial displacement vectors \(u_t\), auxiliary probe states \(P_t\), and a termination decision bit \(s_t\) via lightweight multi-layer perceptrons (MLPs). The agent continues its trajectory of active visual exploration until sufficient evidence is aggregated to emit the final answer.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["High-Resolution Input Canvas"] --> B["Dual-Scale Local Perception<br/>Foveated Glimpse + Peripheral Glimpse"]
    B --> C["Visual Feature Encoding & Concat<br/>ResNet Encodes Local & Peripheral Inputs"]
    C --> D["Recurrent State Tracking<br/>LSTM Nonlinear State Update"]
    D --> E["Action & Auxiliary Probing Heads<br/>Predict Displacement Vector & Local Probes"]
    E -->|Stop Bit False| B
    E -->|Stop Bit True| F["Global System State & Final Answer"]

Key Designs

1. Dual-scale local perception: blocking global shortcuts while preserving navigation cues Standard vision-language models feed full-resolution canvas features directly to attention layers, enabling them to latch onto global spatial shortcuts during in-distribution training. FoveAgent enforces an asymmetric information bottleneck: the central foveated glimpse \(G_l\) maintains native pixel resolution to discern detailed switch/marker states, whereas the peripheral crop \(G_p\) spans four times the spatial scale (\(s=4\)) but is bilinearly downsampled to match the fixed sensor resolution of \(G_l\). This severe resolution reduction provides sufficient coarse geometric outlines to guide navigation toward neighboring visual entities, but strips away micro-level visual details, preventing the network from inferring remote states without physically traversing to them.

2. Constrained displacement actions and auxiliary probe supervision: ensuring step-by-step causal exploration To prevent the policy from executing unconstrained random-access jumps that emulate global attention, the spatial displacement action \(u_t = (\theta, d)\) is parameterized in polar coordinates where angle \(\theta \in [0, 2\pi]\) and distance \(d \in [0, G_p]\). Restricting \(d\) within the perimeter of the current peripheral glimpse guarantees that the next fixation center \(x_{t+1} = x_t + u_t\) remains strictly within the currently visible field of view, enforcing topological continuity. Concurrently, an auxiliary probing head predicts the local semantic state \(P_t \in \{0, 1, \text{null}\}\) at each step. This intermediate supervisory signal expedites representation learning without leaking future global trajectory information.

3. External canvas marking and active visual search: decoupling external memory from state tracking In unconstrained visual reasoning scenarios lacking directional arrows, the agent must simultaneously perform active visual search, remember visited entities, and track system state. Encoding complete spatial visit histories within finite-dimensional recurrent hidden vectors causes memory saturation as entity counts scale. The architecture resolves this by introducing a physical external memory: upon evaluating an entity, the agent renders a permanent black dot onto the environment canvas, which remains visible in all subsequent glimpses. Augmented with discrete zoom-in and zoom-out actions, the agent dynamically adjusts peripheral coverage during active search without exposing high-resolution global shortcuts.

Loss & Training

The local perception policies are optimized via behavioral cloning (imitation learning) over expert trajectories generated by algorithmic oracle policies. The multi-task trajectory loss is formulated as: $$ \mathcal{L} = \mathcal{L}{\text{action}}(u_t, u_t^) + \lambda_1 \mathcal{L}_{\text{probe}}(P_t, P_t^) + \lambda_2 \mathcal{L}(s_t, s_t^*) $$ where displacement regression utilizes smooth }\(L_1\) loss, while probing classification and stop prediction use cross-entropy loss. To foster trajectory error recovery and robustness against drifting during autonomous inference, Gaussian coordinate perturbations are injected into the glimpse sampling pipeline during training.

Key Experimental Results

Main Results

Evaluations span four visual reasoning environments comparing in-distribution (InD) settings with out-of-distribution (OOD) complexity shifts. In Visual Parity and State Machine, models are trained on 2โ€“10 switches and evaluated on 11โ€“20 switches (length extrapolation) as well as varied canvas resolutions. FoveAgent-LSTM is benchmarked against fine-tuned Qwen2.5-VL-3B-Instruct, closed-source foundation models (GPT-5.4, Claude Sonnet 4.6), and alternative sequence backbones.

Task & Architecture Setup InD Accuracy (%) OOD Complexity (%) OOD Resolution (%)
Visual Parity (2-10 vs 11-20 switches)
Qwen2.5-VL-3B-Instruct (CoT Fine-tuned) ~92.0 ~45.0 ~38.0
GPT-5.4 (Zero-Shot) ~82.0 ~35.0 ~31.0
Claude Sonnet 4.6 (Zero-Shot) ~80.0 ~32.0 ~28.0
FoveAgent-LSTM (Ours) 99.5 98.2 97.6
State Machine (Dihedral Group Modulo-3)
Qwen2.5-VL-3B-Instruct (Fine-tuned) 95.8 32.4 26.5
Qwen3.5-VL-27B (Zero-Shot) 88.4 24.1 19.8
Mamba (Selective State-Space Model) 89.2 18.5 15.2
xLSTM (with M-LSTM Parallel Blocks) 91.0 22.3 17.6
FoveAgent-LSTM (Ours) 99.1 96.8 95.4

Performance on the real-world mathematical plot reasoning benchmark Finding Roots (Table 1 from the paper):

Evaluation Scenario Global Baseline G(1200,800) G(300,200)+L G(480,320)+L G(600,400)+L G(1200,800)+L (Full FoveAgent-Qwen)
In-distribution (1-6 roots, 1-6 subplots) 57.24 52.86 74.50 80.44 82.26
OOD-subplots (7-9 subplots) 50.12 43.94 57.49 68.78 77.24
OOD-numroots (7-10 roots) 32.63 57.81 61.62 65.85 67.12
OOD-(subplots + numroots) 35.53 49.25 50.92 59.70 57.46

Ablation Study

A systematic perception-mode ablation demonstrates the impact of visual input scope on recurrent LSTM backbones:

Perception Configuration Input Representation Mode InD Acc (%) OOD Length Acc (%) Failure Mechanism
Global Baseline Full high-resolution image 94.6 28.3 Exploits global pixel correlations; attention disperses at scale
Local + Global Hybrid High-res crop + full high-res image 95.2 31.5 Unconstrained global bypass leaks shortcuts, breaking extrapolation
FoveAgent-LSTM High-res fovea + downsampled peripheral 99.5 98.2 Enforces causal step-by-step traversal; near-lossless length extrapolation

Analysis of glimpse scale \(s\) and sensor resolution \(r\) using the exploration efficiency metric \(p_{rs} = A / L\) shows that oversized peripheral views (\(s=16\)) or overly sharp peripheral sensors (\(r=320\)) collapse efficiency below 0.05 due to shortcut re-emergence. Optimal performance (\(p_{rs} = 0.99\)) is strictly localized at \(s=4, r=80\).

Key Findings

  • Dual necessity of locality and recurrence: Equipping an LSTM with global perception (Global or Local+Global) fails catastrophically OOD, while feeding local glimpses to parallel backbones (Mamba, xLSTM, Transformers) also exhibits severe degradation. Only genuine nonlinear recurrence combined with strictly local glimpses achieves length generalization.
  • Dichotomy between state tracking and feature recall: On the RECALL benchmark requiring parallel visual binding rather than sequential state updates, global models (Qwen2.5-VL) outperform serial foveated agents. Foveated sequential perception specifically targets the inductive bottleneck of algorithmic state tracking.
  • Inefficacy of uniform compute scaling: Scaling global image resolution and vision token budgets by \(10\times\) yields an insignificant \(+3.8\%\) accuracy gain on Finding Roots, whereas reallocating compute to local foveated crops delivers a massive \(+29.0\%\) gain at equivalent token budgets.

Highlights & Insights

  • Diagnosing the architectural root of visual shortcuts: Proves that visual reasoning failures out-of-distribution stem from global single-shot perception mechanisms rather than parameter scale or sequence context limits.
  • Downsampled periphery as an inductive constraint: Deliberately degrading peripheral resolution functions as an effective inductive bias, providing coarse topological routing signals while physically prohibiting premature semantic extraction.
  • Externalized memory via physical marking: Alleviates internal recurrent state bottlenecks by writing traversal markers directly onto the input canvas, establishing an elegant separation between spatial coverage tracking and abstract algebraic state updates.

Limitations & Future Work

  • Reliance on oracle supervision: Exploration policies are trained via behavioral cloning on synthetic ground-truth trajectories, which are non-trivial to curate for open-domain natural image tasks.
  • Static glimpse discretization: Sensor window dimensions and saccadic jump distances are governed by fixed hyperparameters rather than dynamically adapted based on local visual entropy or semantic ambiguity.
  • Integration with web-scale VLM pre-training: Seamlessly incorporating foveated sequential policies into multi-billion parameter autoregressive or diffusion backbones poses significant throughput and token-caching challenges.
  • vs Length Generalization in Transformers (Abbe et al., NeurIPS 2024; Anil et al., NeurIPS 2022): While prior literature diagnosed shortcut learning in 1D symbolic strings, this work extends the theory to continuous 2D visual layouts where information acquisition and state tracking must be co-optimized.
  • vs Adaptive Recurrent Vision (Veerabadran et al., NeurIPS 2023): Prior work applied recurrence over full-image inputs and observed steep degradation under image scale shifts; this paper reveals that restricting perception to strict local glimpses is the essential missing ingredient.
  • vs Recurrent Attention Models (Mnih et al., NeurIPS 2014 RAM): Whereas classical foveation models were motivated by computational savings and biological plausibility, this study repositions foveated attention as a fundamental prerequisite for compositional length generalization.

Rating

  • Novelty: โญโญโญโญโญ Establishes the necessity and sufficiency of local foveated perception for visual state tracking and length generalization.
  • Experimental Thoroughness: โญโญโญโญโญ Rigorous validation across controlled algorithmic benchmarks and real-world mathematical plot reasoning tasks.
  • Writing Quality: โญโญโญโญโญ Lucid formulation connecting biological visual systems, algorithmic shortcuts, and empirical foundation model limits.
  • Value: โญโญโญโญโญ Provides profound architectural insights for designing the next generation of robust multimodal reasoning models.