MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention¶
Conference: ECCV 2026
Paper: ECCV 2026 Paper Page
Code: To be released
Area: Image Generation
Keywords: Masked autoregressive models, KV cache acceleration, cache-aware attention, selective KV refresh, training-free acceleration
TL;DR¶
To eliminate the severe computational redundancy in masked autoregressive image generation where full bidirectional attention and feed-forward networks are recomputed every step, MARché proposes a training-free framework featuring decoupled cache-aware attention, attention-guided selective KV refresh, and periodic cache calibration, attaining up to a 1.79× speedup with negligible fidelity loss.
Background & Motivation¶
Masked autoregressive (MAR) models combine the expressive power of autoregressive generative models with the efficiency of masked prediction by predicting continuous tokens in a permuted sequence order using bidirectional attention and diffusion-based losses. As a consequence, MAR has established itself as a leading paradigm for high-fidelity visual synthesis. Nonetheless, during inference, MAR models undergo multi-step decoding (e.g., 64 steps by default) where every single step re-evaluates full bidirectional self-attention and feed-forward network (FFN) projections across all tokens—both ungenerated masked positions and previously generated visible tokens. This unmitigated per-step workload results in substantial computational overhead and latency, hindering deployment and scalability to larger models and higher resolutions.
An empirical analysis of intermediate token representations in MAR reveals temporal locality: between successive decoding steps, key and value projections for the vast majority of tokens remain exceptionally stable, frequently exhibiting cosine similarities exceeding 0.95. However, standard causal KV caching methods common in large language models cannot be directly ported to MAR. Because MAR relies on bidirectional attention, newly predicted tokens dynamically reshape the context for existing visible tokens; blindly reusing historical KV representations indefinitely induces severe value drift and completely degrades image generation fidelity.
The solution lies in decoupling stable context from dynamic contextual shifts. Core idea: exploit the cross-step attention scores from newly generated tokens to locate contextually sensitive tokens, establishing a training-free decoding framework driven by decoupled cache-aware attention and selective KV refresh that computes only active subsets and bypasses redundant FFN passes while preserving full bidirectional context.
Method¶
Overall Architecture¶
MARché restructures the per-step decoding pipeline without altering the underlying Transformer architecture or requiring any retraining. At each generation step, the full token sequence is partitioned into an active set (comprising generating tokens for the current step, caching tokens from the immediately prior step, and contextually influenced refreshing tokens) and a cached set containing all remaining stable tokens. In the first two decoder layers, standard full attention is executed to capture fresh representations and calculate head-averaged cross-token attention scores for active token budget allocation. From Layer 3 onward, a fused cache-aware attention operator is invoked: queries from active tokens independently attend to active KV pairs and cached KV pairs, which are smoothly merged via safe online softmax without physical concatenation, while feed-forward execution for cached tokens is skipped entirely. Periodic full cache refreshes ensure long-step drift is eliminated.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Sequence<br/>Masked and visible token embeddings"] --> B["Selective KV Refresh via Generating Attention<br/>Layers 1-2 compute full attention, aggregate generating query scores to select Top-K refreshing tokens"]
B --> C{"Trigger Periodic Full Refresh?<br/>(step mod 3 == 0 or initial step)"}
C -->|Yes| D["Periodic Full Refresh & Calibration<br/>Recompute all layers and update entire KV cache"]
C -->|No| E["Dual-Path Decoupled Attention & Online Merging<br/>Active queries interact with active & cached KV via Safe Online Softmax, skipping cached FFN"]
D --> F["Decoded Context Vectors<br/>Fed to diffusion sampler for next-token prediction"]
E --> F
Key Designs¶
1. Selective KV Refresh via Generating Attention: Targeting Contextually Sensitive Subsets
Conventional KV cache compression or eviction strategies in LLMs rely on discarding unimportant past keys, whereas MARché leverages bidirectional attention dynamics to identify which cached tokens must be recomputed. In masked image decoding, newly unmasked patches alter the local visual semantics of their correlated regions (e.g., boundaries of the same object or texture-consistent regions). To accurately capture these shifts without excessive overhead, MARché identifies Layer 2 as the optimal decision anchor—balancing shallow-layer context quality against deep-layer alignment. During Layer 2 full attention, attention weights from all generating tokens to every other token in the sequence are averaged across all attention heads to form global relevance scores. The Top-\(K\) tokens with the highest scores are selected as refreshing tokens. Grouped with current generating tokens and previous caching tokens into a fixed active budget (e.g., 64 tokens), this mechanism maintains high contextual fidelity at bounded computational cost.
2. Dual-Path Decoupled Attention & Online Merging: Avoiding Concatenation and Skipping Redundant FFN
Directly concatenating active keys and values with cached key/value buffers before calling standard attention kernels creates substantial non-contiguous memory transfers and kernel latency bottlenecks. MARché explicitly separates attention into two paths: active tokens project fresh query vectors \(q_i\), key vectors \(K_A\), and value vectors \(V_A\), while cached tokens supply pre-stored key/value tensors \((K_C, V_C)\) without computing queries or passing through the FFN. For every active query \(q_i\), attention is computed over the active and cached subsets independently and fused on the fly:
The safe online softmax tracks local normalization coefficients \(\ell_i\), yielding mathematical equivalence to standard bidirectional attention while ensuring high cache locality and kernel fusion. Crucially, following attention, only the active transformed vectors \(z_i\) undergo feed-forward network processing; all cached tokens bypass the FFN completely, saving roughly two-thirds of the model's linear projection workload.
3. Periodic Full Refresh & Calibration: Triple Guards Against Multi-Step Error Drift
Although selective refresh captures the most salient tokens, minor numerical discrepancies across continuous residual connections and normalization layers can accumulate across dozens of decoding iterations, leading to subtle edge artifacts in late denoising stages. MARché introduces three structural safeguards: first, at Step 0, full attention is executed across all layers to initialize the cache; second, at every decoding step, Layers 1 and 2 perform full attention to ensure refresh decisions originate from clean representations; third, a periodic full refresh executes every 3 steps (step \(\pmod 3 == 0\)), recalculating full KV projections across the entire sequence. This amortizes calibration overhead while guaranteeing that cumulative representation drift remains strictly bounded throughout all 64 steps.
Key Experimental Results¶
Main Results¶
On class-conditional ImageNet \(256 \times 256\) synthesis, MARché was evaluated across MAR-B, MAR-L, and MAR-H on a single NVIDIA H100 GPU under the standard 64-step schedule against leading diffusion, masked, and autoregressive models.
| Method | Latency (s/im) ↓ | FID ↓ | Inception Score ↑ | Param | Speedup ↑ |
|---|---|---|---|---|---|
| MaskGIT | 0.440 | 6.18 | 182.1 | 227M | - |
| DiT-XL/2 | 0.787 | 2.27 | 278.2 | 675M | - |
| LlamaGen-XXL | 0.897 | 3.09 | 253.6 | 1.4B | - |
| LlamaGen-3B | 1.011 | 3.05 | 222.3 | 3.1B | - |
| MAR-B (Baseline) | 0.104 | 2.35 | 281.1 | 208M | 1.00× |
| LazyMAR-B | 0.074 | 5.32 | 235.1 | 208M | 1.41× |
| MARché-B (Ours) | 0.064 | 2.56 | 270.3 | 208M | 1.57× |
| MAR-L (Baseline) | 0.193 | 1.84 | 296.3 | 479M | 1.00× |
| LazyMAR-L | 0.132 | 4.32 | 246.6 | 479M | 1.46× |
| MARché-L (Ours) | 0.115 | 2.16 | 278.6 | 479M | 1.68× |
| MAR-H (Baseline) | 0.336 | 1.62 | 298.6 | 943M | 1.00× |
| LazyMAR-H | 0.220 | 4.00 | 251.9 | 943M | 1.52× |
| MARché-H (Ours) | 0.188 | 2.02 | 281.4 | 943M | 1.79× |
Ablation Study¶
Ablation on Active Set Token Construction Evaluating the necessity of distinct token types within the active set on MAR-B.
| Active Set Strategy | FID ↓ | Note |
|---|---|---|
| MARché (Full model) | 2.56 | Incorporates generating, caching, and high-attention refreshing tokens |
| MARché w/o caching tokens | 2.70 | Drops caching tokens; small degradation without computational benefit |
| MARché w/o generating tokens | 504.82 | Omits current generation targets; generation completely fails |
| Random selection | 564.61 | Randomly samples tokens to recompute without attention guidance |
Speedup Component Breakdown Isolating the speedup contributions of cache-aware attention and FFN skipping on MAR-B.
| Method | Implementation | Latency (s) | Speedup |
|---|---|---|---|
| MAR | Standard attention + full FFN | 0.104 | 1.00× |
| MAR + selected FFN | Standard attention + FFN skipping | 0.092 | 1.12× |
| MAR + cache-aware Attn | Cache-aware attention + full FFN | 0.067 | 1.55× |
| MARché (Full model) | Cache-aware attention + FFN skipping | 0.064 | 1.63× |
Key Findings¶
- Generating tokens are non-negotiable anchors: Excluding generating tokens from recomputation leads to an catastrophic collapse in generation (FID 504.82). Meanwhile, selecting refreshing tokens by high attention scores achieves an FID of 2.56, substantially outperforming low-attention selection (FID 3.80) and random selection (FID 3.01), verifying that attention magnitude reliably measures contextual dependency.
- Cache-aware attention drives primary speedup gains: As shown in the ablation table, cache-aware attention alone accounts for the bulk of acceleration, boosting throughput by 1.55× (latency from 0.104s to 0.067s). FFN skipping provides an additional 1.12× improvement, collectively yielding a 1.63× speedup in the benchmark configuration.
- Layer 2 provides optimal decision trade-offs: Layer 2 attention scores achieve a 66.8% selection overlap with subsequent layers, balancing context accuracy and overhead. Layer 1 offers lower latency (0.154s) but suffers from 49.3% overlap and poorer FID (2.62); Layer 4 improves FID to 2.49 but increases latency to 0.1613s due to four full-attention layers.
Highlights & Insights¶
- Zero Retraining and Universal Plug-and-Play: MARché achieves up to 1.79× speedup purely at inference time without modifying model weights, architecture, or the original diffusion-based loss framework.
- Inverting Attention Eviction Heuristics: While LLM inference uses attention to evict unimportant historical tokens, MARché repurposes attention in bidirectional models to pinpoint which cached tokens must be recomputed, turning a compression heuristic into an active context selector.
- Online Softmax Streamlining: Merging separate attention matrices on-chip via Safe Online Softmax eliminates off-chip concatenation overhead and improves memory locality by 16.2%–27.3%.
Limitations & Future Work¶
- Static Budgeting: The active token set budget is currently fixed to a constant size (e.g., 64 tokens) and may not adapt dynamically to varying scene complexity or token sparseness across steps.
- Fixed Layer Anchor: Layer 2 is uniformly utilized for refreshing token selection across all steps and architectures; future investigations could explore dynamic, scale-adaptive layer selection tailored to progressive generation stages.
Related Work & Insights¶
- vs. LazyMAR (Yan et al., 2025): LazyMAR reuses hidden features and alters the generation order via similarity heuristics, leading to severe visual degradation (FID increases to 4.00–5.32 across scales). MARché strictly adheres to the official generation trajectory, achieving superior speedups (1.57×–1.79×) while maintaining competitive FID (2.02–2.56).
- vs. ENAT (Ni et al., NeurIPS 2024): ENAT requires training specialized spatial-temporal attention modules, whereas MARché is completely training-free and directly compatible with pre-trained weights.
- vs. LLM KV Caching (Keyformer / H2O): While LLM caching exploits unidirectional causal masks where past tokens are immutable, MARché addresses the challenging 2D bidirectional setting where unmasked tokens interact globally across all directions.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ An elegant formulation adapting KV caching and online attention fusion to bidirectional masked visual transformers.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks across three model scales (B/L/H), detailed ablation on active set token types, and granular kernel profiling.
- Writing Quality: ⭐⭐⭐⭐⭐ Well-structured, mathematically rigorous, with clear empirical motivations and transparent trade-offs.
- Value: ⭐⭐⭐⭐⭐ Outstanding practical utility for accelerating visual autoregressive models in production environments without retraining costs.