Skip to content

NAPA: Natively Multimodal Autoregressive Perception Architecture

Conference: ECCV 2026
Paper: ECCV Paper
Code: https://github.com/tiiuae/falcon-perception
Area: Segmentation
Keywords: multimodal, autoregressive perception, open-vocabulary segmentation, hybrid attention, feature upsampling

TL;DR

Presents Falcon Perception (NAPA), a single early-fusion Transformer that eliminates the encoder-decoder split via a hybrid bidirectional-causal attention pattern and an autoregressive Chain-of-Perception interface, enabling high-resolution masks via parallel dot-product decoding.

Background & Motivation

Perception systems for open-vocabulary detection and dense instance segmentation have long relied on a modular encoder-decoder paradigm. Under this recipe, a standalone vision backbone extracts visual features, while a separate decoder or late-fusion module interacts with text queries via cross-attention to output bounding boxes or instance masks. Although effective for classic closed-vocabulary benchmarks, this separation artificially bifurcates the visual and semantic representation spaces, restricting cross-modal reasoning to late-stage fusion. With the rise of generalist vision-language models (VLMs), researchers have attempted to serialize detection and segmentation into discrete text tokens or polygon vertices. However, pure language modeling over dense outputs causes severe sequence length inflation, making autoregressive decoding prohibitively expensive when hundreds of instances are present.

The fundamental tension of this modular separation is that early-layer visual representations cannot interact bidirectionally with language semantics, leading to attention disconnects and hallucinations on compositional prompts involving fine-grained attributes, spatial constraints, and complex relational bindings. Conversely, naively deploying an autoregressive language model for pixel-level dense outputs leads to catastrophic computational bottlenecks. Existing early-fusion multimodal models frequently resort to modality-specific projection layers, specialized normalization, or Mixture-of-Experts (MoE) branches, which compromises architectural elegance while still struggling to deliver high-resolution pixel masks efficiently.

This paper questions whether dense grounding systems actually require an encoder-decoder split to perceive and predict. Core idea: unify image patches, text prompts, and task tokens into a single early-fusion Transformer parameter space with a hybrid bidirectional-causal attention mask, decoding objects via a coarse-to-fine "Chain-of-Perception" interface where discrete tokens anchor instances and continuous spatial outputs are generated via lightweight parallel feature dot products.

Method

Overall Architecture

Falcon Perception departs from the conventional encoder-decoder design by operating on a single unified dense Transformer backbone \(f_\theta\), supplemented by Fourier feature spatial encoders and lightweight task heads. The model input unifies visual patch embeddings \(V \in \mathbb{R}^{N \times d}\), text prompt embeddings \(T \in \mathbb{R}^{L \times d}\), and task tokens for \(K\) predicted objects into a single sequence \(X \in \mathbb{R}^{(N + L + 3K) \times d}\).

During autoregressive generation, the model predicts object properties in a structured coarse-to-fine "Chain-of-Perception" order: center coordinate Token <coord>, size Token <size>, and segmentation Token <seg>. Continuous bounding-box coordinates and dimensions are mapped via Fourier features and reinjected into the sequence as strong spatial conditioning, while the final hidden state of the <seg> token is directly dot-producted with high-resolution image-guided upsampled features to yield binary instance masks in parallel.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multimodal Input<br/>Native Image Patches + Text Prompts"] --> B["Single Early-Fusion Transformer<br/>Hybrid Mask: Image Bidirectional + Text/Task Causal"]
    B --> C["3D Rotary Positional Embeddings<br/>1D Sequence Causal + GGRoPE 2D Isotropic Rotation"]
    C --> D["Chain-of-Perception AR Decoding<br/>Raster Order: coord โ†’ size โ†’ seg"]
    D --> E["Fourier Feature Spatial Encoding<br/>Random Gaussian Mapping Mitigates Spectral Bias"]
    D --> F["Feature Upsampling & Dot-Product Masking<br/>AnyUp Cross-Attention Guidance + seg Dot Product"]
    E -->|Continuous Geometric Conditioning| D
    F --> G["Multitask Outputs<br/>Bounding Boxes + High-Resolution Masks"]

Key Designs

1. Hybrid Attention Mechanism: Reconciling 2D Visual Context with 1D Causal Generation Standard autoregressive causal masking is inherently suboptimal for 2D image patches because spatial recognition requires bidirectional context across all visual tokens; conversely, full bidirectional attention prevents autoregressive sequence decoding. Falcon Perception introduces a tailored hybrid attention mask: image tokens attend bidirectionally to all other image patches, acting as an integrated vision encoder, whereas text prompts and task tokens attend to all image tokens but only to preceding text and task tokens causally. This allows a single set of Transformer parameters to simultaneously serve as a bidirectional image feature extractor and an autoregressive generation engine, completely removing the need for an external cross-attention bridge.

2. 3D Rotary Positional Embeddings: Decoupling Sequential Order and 2D Spatial Geometry Flattening 2D images into linear token sequences destroys spatial topological relationships. To preserve geometric continuity, the multi-head attention head dimension \(d_{\text{head}}\) is decomposed into two orthogonal subspaces: the first half \(d_{\text{head}} / 2\) applies standard 1D RoPE to encode the sequential token index \(t\), maintaining causal ordering for language and task generation; the second half \(d_{\text{head}} / 2\) adopts Golden Gate RoPE (GGRoPE) to encode continuous 2D coordinates \(p = (x, y)\). GGRoPE rotates each dimension pair based on the spatial projection onto directional vectors \(\{u_j\}\) sampled uniformly on the unit circle via the Golden Ratio \(\phi\): $\(\theta_j = \omega_j (p \cdot u_j)\)$ This isotropic formulation eliminates the orthogonal grid bias of conventional Axial RoPE, providing robust spatial grounding across arbitrary orientations and aspect ratios. The 2D rotation is applied sparsely only to valid image patches, preventing cross-modal interference with text tokens.

3. Continuous Fourier Feature Projection and Raster Ordering: Mitigating Spectral Bias and Instance Ambiguity Directly predicting dense pixel coordinates autoregressively causes catastrophic sequence inflation. Falcon Perception factorizes each object into center coordinates \(c\), dimensions \(s\), and a mask token <seg>. To overcome the severe spectral bias and low-frequency preference of standard 1000-bin discrete coordinate tokenization, a random Gaussian mapping matrix \(B \in \mathbb{R}^{d/2 \times 2}\) (\(\sigma = 10.0\)) maps continuous coordinates to high-frequency harmonic manifolds: $\(\gamma(c) = [\cos(2\pi B c), \sin(2\pi B c)]\)$ These continuous embeddings replace static token embeddings in the input stream. For multi-object targets, instances are serialized using a canonical top-to-bottom, left-to-right raster ordering, which minimizes search ambiguity and accelerates coordinate head optimization compared to random or size-based sequences.

4. Lightweight Feature Upsampling and Parallel Dot-Product Masking: Eliminating Heavy Mask Decoders Because <seg> is predicted after center coordinates and box dimensions, instance identity and spatial extent are already established, and the <seg> hidden state \(h_{\text{seg}}\) retains full access to the early-fusion multimodal representation. To recover pixel-level boundary details without heavy mask decoders (e.g., Mask2Former), Falcon Perception leverages a frozen, image-guided feature upsampler (AnyUp). AnyUp applies cross-attention between high-resolution input pixels \(I \in \mathbb{R}^{H \times W \times 3}\) as queries and backbone features \(V_{\text{out}}\) as keys/values, producing an enhanced feature map \(V^* \in \mathbb{R}^{H \times W \times d}\). Binary instance masks are then computed directly via an instantaneous dot product: $\(\hat{m} = \sigma\left(V^* \cdot \text{Proj}(h_{\text{seg}})^T\right)\)$ This reduces pixel-level prediction to a closed-form parallel inner product, combining sequence-level modeling flexibility with high-throughput dense output generation.

Loss & Training

The model is trained via a two-stage curriculum: 1. Multi-Teacher Distillation: Transfers representations from frozen DINOv3 and SigLIP2 teachers into the unified dense backbone. A Gram matrix loss is applied to match student patch feature correlations with teacher representations, stabilizing early visual representation learning. 2. Chain-of-Perception Pretraining (~700 GT): - Stage 1 (450 GT): Full autoregressive training over all text and task tokens with inverse learning rate decay (\(4\times 10^{-4} \to 1\times 10^{-4}\)) and up to 100 masks per expression; - Stage 2 (225 GT): Task alignment with query-masked attention (preventing cross-query attention leakage) and prompt loss masking, decaying learning rate to \(4\times 10^{-6}\); - Stage 3 (10 GT): Long-context adaptation for crowded scenes supporting up to 600 masks per expression. 3. Optimization and Engineering: Optimized with Muon for hidden layer matrices using Kimi Moonlight scaling rules; native variable-resolution images are processed via scatter-and-pack sequence packing with global rank loss normalization to prevent gradient skew across data-parallel ranks.

Key Experimental Results

Main Results

On the SA-Co open-vocabulary segmentation benchmark and the diagnostic PBench benchmark, Falcon Perception (600M parameters) demonstrates superior segmentation quality and compositional reasoning over much larger generalist VLMs and specialized models.

Model Params SA-Co Avg Macro F1 SA-Co pmF1 SA-Co MCC PBench Avg Macro F1 RefCOCOm F1
gDino-T ~170M - 16.2 0.15 - -
OWLv2* ~430M - 42.0 0.57 - -
DINO-X - - 55.2 0.38 - -
Qwen3-VL-2B + SAM 2B 50.3 36.8 0.65 37.0 54.0
Qwen3-VL-8B + SAM 8B 56.4 41.1 0.79 49.0 64.7
Qwen3-VL-30B + SAM 30B 62.8 49.9 0.64 52.7 72.1
Moondream2 2B 62.5 53.5 0.43 44.7 62.0
Moondream3 9B 58.0 47.8 0.45 50.5 87.1
SAM 3 0.9B 62.3 66.1 0.82 44.4 36.4
Falcon Perception (Ours) 0.6B 68.0 62.1 0.64 57.0 77.3

On PBench's level-by-level breakdown across compositional capabilities, Falcon Perception demonstrates widening advantages as prompt complexity increases:

Model L0 (Simple Objects) L1 (Attributes) L2 (OCR-Guided) L3 (Spatial) L4 (Relations) Dense (Crowded Scenes)
SAM 3 64.3 24.6 54.4 31.6 33.3 58.4
Moondream3-9B 63.2 41.8 59.6 45.4 42.3 12.9
Qwen3-VL-8B + SAM 65.6 58.2 65.6 49.1 50.9 4.7
Qwen3-VL-30B + SAM 69.2 61.2 68.8 52.9 55.2 8.9
Falcon Perception (0.6B) 65.1 38.0 63.6 53.5 49.1 72.6

Ablation Study

Component-wise ablations on optimizer choice, feature regularization, and instance sequence ordering:

Ablation Dimension Configuration PBench Metric SA-Co Metric Key Mechanism & Observation
Optimizer AdamW 56.4 (Det F1) 49.0 (Det F1) Standard optimizer exhibits slower convergence across multi-loss heads
Optimizer Muon (Ours) 57.7 (Det F1) 53.8 (Det F1) Orthogonalized weight updates yield significantly lower loss on coord and LM heads
Regularization w/o Gram Loss 52.7 (Seg F1) 51.1 (Seg F1) Lacks structural regularization; finetuning degrades low-level image features
Regularization With Gram Loss 53.8 (Seg F1) 52.6 (Seg F1) Preserves spatial patch correlation structure from teacher model, boosting mask quality
Instance Ordering Random Ordering 52.2 (Det F1) 46.3 (Det F1) Random target permutation creates spatial ambiguity during autoregressive learning
Instance Ordering Size (Largest First) 57.7 (Det F1) 53.8 (Det F1) Captures salient objects early but lacks spatial scanning locality
Instance Ordering Raster Ordering 59.3 (Det F1) 56.2 (Det F1) Matches canonical 2D reading order, driving faster convergence on coordinate loss

Key Findings

  • Resolution Phase Transition in Dense Scenes: Scaling image resolution from \(448^2\) to \(1024^2\) yields a moderate 1.6ร— gain in presence recognition (MCC), but drives a dramatic 15ร— leap in dense localization (pmF1 rises from 3.9% to 61.0% on PBench Dense). This demonstrates that crowded-scene failure in VLMs is primarily spatial resolution deprivation rather than semantic incapacity.
  • Latent Distribution Recovery via Stochastic Sampling: Sampling from coordinate and size head distributions (Pass@8) increases SA-Co cgF1 from 34.7 to 54.3 (+19.6 points), matching SAM 3 and doubling accuracy on the difficult Wiki-Common subset (19.3 to 45.0), proving that autoregressive likelihood retains rich candidate modes.
  • Catastrophic Divergence without Distillation: Initializing the dense perception model from scratch caused immediate divergence during segmentation pretraining due to high gradient instability on visual tokens, whereas multi-teacher distillation ensured stable convergence and outperformed random initialization trained for 230K steps.

Highlights & Insights

  • Unified Parameter Minimalism: Collapses the traditional visual backbone and heavy cross-attention mask decoder into a single Transformer, achieving end-to-end open-vocabulary segmentation and OCR with just 600M parameters.
  • Isotropic 2D Positioning via GGRoPE: Projecting positional coordinates along Golden Ratio-spaced angles on the unit circle avoids axis-aligned inductive bias, maintaining robust spatial grounding under arbitrary aspect ratios.
  • Reusable Parallel Dense Projection Head: Combining an autoregressively generated <seg> token with a frozen, cross-attention-guided feature upsampler (AnyUp) simplifies dense mask prediction into a fast matrix dot product.

Limitations & Future Work

  • Presence Calibration on Hard Negatives: Image-Level MCC (0.64 vs. 0.82 on SA-Co) lags behind SAM 3 on sets with abundant absent-target queries, indicating a tendency toward over-prediction.
  • Extreme Small Text Degradation in OCR: While the 300M Falcon-OCR variant reaches 80.3% on olmOCR, performance plateaus on severely degraded scans and ultra-tiny fonts.
  • Future Directions: Applying preference optimization (DPO/RLHF) to calibrate object presence decisions, and extending the unified architecture to video object grounding and interactive robotics control.
  • vs SAM 3: SAM 3 relies on prompt engineering and a dedicated mask decoder, excelling at generic object presence (MCC 0.82) but failing on compositional queries requiring spatial layout (Level 3: 31.6 vs. 53.5) and crowded scenes; Falcon Perception bridges this semantic-geometric divide through native early fusion.
  • vs Generalist VLMs (Qwen3-VL / Moondream): Standard VLMs predict bounding boxes as text coordinates, which suffers severe context exhaustion and hallucinations in crowded environments (Qwen3-VL-30B scores only 8.9 on Dense split); Falcon Perception's rasterized Chain-of-Perception maintains robust localization up to 600 instances (scoring 72.6).

Rating

  • Novelty: โญโญโญโญโญ Completely removes the encoder-decoder split, successfully demonstrating native early-fusion autoregressive dense segmentation at 600M scale.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluation across the 5-level PBench, SA-Co, RefCOCO, resolution scaling, Pass@k sampling, and document OCR.
  • Writing Quality: โญโญโญโญโญ Exceptionally clear narrative, disciplined mathematical formulation, and solid empirical ablations.
  • Value: โญโญโญโญโญ Offers a blueprint for lightweight, unified edge perception models and natively multimodal vision architectures.