Skip to content

Test-Time Registers as Global Priors for Tokenized Image Generation

Conference: ECCV 2026
Paper: ECCV Official
Project: https://y-research-sbu.github.io/RegToken
Area: Image Generation
Keywords: Test-Time Registers / Tokenized Image Generation / Attention Sinks / Global Priors / Vision Transformers

TL;DR

This paper proposes RegToken, a training-free framework that localizes ViT attention sinks using Normalized Feature Norms (NFN), extracts register subspaces via TokenRank, and applies a projection-and-conservation update to repurpose high-norm attention artifacts into global low-frequency priors, significantly improving generation fidelity and alignment for compact 1D token decoders under completely frozen weights.

Background & Motivation

Vision Transformers (ViTs) and multimodal foundation models consistently exhibit the "attention sink" phenomenon in their deep layers: a tiny subset of tokens draws an overwhelmingly disproportionate share of attention mass while accumulating massive activation norms. Early studies predominantly treated these outliers as numerical artifacts stemming from the lack of an internal garbage-collection mechanism, introducing explicit learnable register tokens during pretraining to absorb redundant activations and stabilize dense visual representations. Subsequent investigations revealed that similar outlier suppression could be achieved without retraining by identifying and injecting test-time register neurons into frozen models. However, prior research has largely confined registers to interpretability diagnoses, attention map smoothing, or linear probe benchmarks, leaving open a fundamental question: do these high-norm representations pooled within attention sinks encode structured information that can be operationalized for generative downstream tasks?

A spectral analysis of these representations in feature space indicates that outlier tokens in attention sinks are not unstructured numerical noise. Instead, they systematically concentrate substantial energy in smooth, global, low-frequency scene statistics, such as color tone, ambient illumination, and coarse layout. In compact 1D discrete visual token generation frameworks (such as TiTok and HCT), models reconstruct high-resolution images from merely dozens of discrete tokens, frequently lacking lightweight, plug-and-play global conditioning signals. Conventional approaches either lean heavily on discriminative [CLS] readouts or undergo computationally demanding test-time optimization across all tokens, which can easily drift into mode collapse or compromise global scene harmony within tight optimization budgets.

The central hurdle in converting register structures from frozen backbones into plug-and-play generative priors lies in bypassing brittle manual heuristics to achieve automated layer localization, subspace projection, and discrete codebook alignment without training. Core idea: RegToken leverages Normalized Feature Norms (NFN) and TokenRank to automatically localize layers and feature subspaces dominated by attention sinks in frozen ViTs, distilling outlier features into compact global register priors via projection-and-conservation interpolation to guide discrete tokenized image generation under completely frozen backbones and decoders.

Method

Overall Architecture

The RegToken pipeline comprises two primary stages: register subspace localization and extraction from frozen backbones, followed by global prior injection into discrete generation pipelines. During the localization stage, the model performs forward inference on a frozen ViT backbone, computing Normalized Feature Norms (NFN) to identify candidate layers where attention sinks emerge, isolates persistent outlier patch tokens within adjacent temporal windows, and ranks dominant channels using Markov-chain-based TokenRank. In the feature consolidation stage, an orthogonal projection-and-conservation update transfers register-subspace energy from outlier patches to newly introduced test-time register tokens. Finally, during compact 1D tokenized image generation, the synthesized register features are aligned and mapped to the discrete codebook, acting as an initialization prior or soft bias directly injected into the frozen decoder to steer image generation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image & Frozen ViT Backbone"] --> B["NFN Layer & Neuron Localization<br/>Forward NFN ratio identifies sink layers"]
    B --> C["TokenRank Subspace Importance Metric<br/>Markov transition matrix quantifies centrality"]
    C --> D["Projection & Conservation Feature Interpolation<br/>Orthogonal transfer from outlier patches to register token"]
    D --> E["Discrete Codebook Alignment & Prior Injection<br/>Project into 1D token sequence z0"]
    E --> F["Frozen Decoder Generates Output Image"]

Key Designs

1. NFN-Guided Layer and Neuron Subspace Localization: Eliminating Manual Heuristics to Pinpoint Attention Sinks

Prior test-time register approaches rely heavily on hand-crafted layer windows and empirical thresholds, resulting in poor cross-architecture transferability. To establish a generalized and automated selection process, RegToken adopts Normalized Feature Norms (NFN). For any projection weight \(W\) and input feature \(z_{\text{in}}(x)\) in the backbone, NFN gauges the degree of directional activation amplification by comparing real feature responses against Gaussian random inputs \(\tilde{z}_{\text{in}}(x)\) of matching norm:

\[NFN(W; x) = \frac{\|W z_{\text{in}}(x)\|_2}{\|W \tilde{z}_{\text{in}}(x)\|_2}\]

By computing the median across attention heads on a validation subset and maximizing the submodule score \(S_\ell = \max_m NFN_{\ell, m}\), the system automatically isolates the top \(\kappa\) layers displaying the strongest amplification as candidate layers \(\mathcal{L}_{\text{cand}}\). Across adjacent layers, spatial patch positions \(\Omega_\ell\) consistently exhibiting outlier norms are tracked, and channels are ranked by their accumulated absolute activations across preceding layers to yield the top-\(k\) register neurons \(\mathcal{R}_\ell = \text{top-}k(\{r_d\})\).

2. Value-TokenRank Attention Centrality Metric: Disentangling True Information Flow from Weight Accumulation

Raw attention probabilities can be deceptive: a token may draw massive attention weights despite having near-zero value vector norms, thereby carrying negligible effective content. To resolve this discrepancy, the method treats the attention matrix \(A\) as the transition probability matrix of a discrete-time Markov chain, computing its stationary distribution \(\pi^\top = \pi^\top A\) to quantify each token's multi-hop TokenRank centrality. Concurrently, it measures the physical content transfer via write mass \(WriteMass(t) = \sum_i A_{i,t} \|V_t\|_2^2\). Monitoring the Markov chain's second-largest eigenvalue \(\lambda_2\) (where a narrower spectral gap \(1-\lambda_2\) signifies sluggish mixing and persistent attention entrapment) enables the model to accurately assess the global importance of each channel, providing a sound theoretical foundation for multi-layer register fusion.

3. Projection-and-Conservation Feature Reallocation: Losslessly Absorbing Low-Frequency Energy While Suppressing Spatial Artifacts

Once the register subspace \(\mathcal{R}_\ell\) and outlier patch set \(\mathcal{P}\) are determined, naively truncating outliers induces severe feature instability, while simple duplication fails to clean up local representations. The method constructs an orthogonal projector \(P_{\mathcal{R}}\) onto the register subspace and calculates the mean register activation across outlier tokens \(\bar{v}_{\mathcal{R}} = \frac{1}{|\mathcal{P}|} \sum_{p \in \mathcal{P}} P_{\mathcal{R}} V_p\). It then updates activations via scale parameter \(s\) and conservation factor \(\alpha\):

\[\begin{aligned} V_{t_{\mathrm{reg}}} &\leftarrow V_{t_{\mathrm{reg}}} + s\,\bar{v}_{\mathcal{R}} \\ V_p &\leftarrow V_p - \alpha P_{\mathcal{R}} V_p, \quad \forall p \in \mathcal{P} \end{aligned}\]

This formulation orthogonally decouples the global outlier components that previously clogged local attention, reallocating them into the newly introduced test-time register token \(t_{\text{reg}}\) while preserving first-order projection magnitude. This simultaneous smoothing of regular patch features and distillation of clean global scene statistics yields an optimal global baseline.

4. Discrete Codebook Alignment and Lightweight Global Prior Injection: Plug-and-Play Conditioning for 1D Token Decoders

To interface continuous register representations with frozen discrete visual tokenizers (such as TiTok and HCT), RegToken employs whitening followed by cosine-similarity nearest-neighbor codebook lookup (or lightweight Procrustes alignment) to project fused register features \(\tilde{r}\) into discrete token space. During decoding, two flexible conditioning schemes are supported: Init-prior directly replaces the sequence's leading token with the interpolated code \(z_0 \leftarrow (1-\gamma)z_0 + \gamma \tilde{r}\), whereas Soft-bias modulates the multinomial sampling probabilities via exponential cosine scaling \(p'(c) \propto p(c) \cdot \exp(\beta \cos(e_c, \hat{u}))\). During optional test-time optimization, gradient updates are restricted strictly to the single injected global token \(z_0\) while all other token indices and decoder parameters remain frozen, achieving rapid and stable global semantic alignment.

Loss & Training

The core methodology is entirely training-free, requiring only a single forward pass over a 1,000-image ImageNet validation subset offline to establish NFN statistics and fix layer/channel indices. When test-time optimization (token opt.) is enabled during decoding, pretrained HCT generative decoder weights and subsequent sequence tokens \(z_{1:T}\) are held completely frozen. Only the injected global token \(z_0\) is optimized for a few steps (e.g., 50 iterations) using a frozen CLIP text-image similarity alignment objective, utilizing small learning rates to prevent distortion of pretrained discrete token distributions.

Key Experimental Results

Main Results

Main evaluations are conducted on ImageNet-1k using a frozen HCT decoder with a VQ-LL-32 compact codebook. Under identical decoding hyperparameters, different global prior sources are benchmarked on visual generation fidelity (FID-5k, IS) and text-image alignment metrics (CLIPScore, SigLIP).

Backbone Injected Prior Method FID-5k (↓) IS (↑) CLIP (↑) SigLIP (↑)
DINOv2-L/14 HCT w/o prior (Beyer et al.) 21.2 281 0.40 3.5
DINOv2-L/14 HCT w/ Random prior 22.5 278 0.39 3.4
DINOv2-L/14 HCT w/ [CLS] prior 20.5 281 0.39 3.6
DINOv2-L/14 HCT w/ Test-time register prior (TTR) 21.5 283 0.40 3.6
DINOv2-L/14 HCT w/ Trained register prior (Darcet et al.) 20.4 285 0.41 3.8
DINOv2-L/14 HCT w/ RegToken (Ours) 20.1 289 0.45 3.9
OpenCLIP-L/14 HCT w/ [CLS] prior 20.8 283 0.40 3.5
OpenCLIP-L/14 HCT w/ Test-time register prior (TTR) 21.1 284 0.41 3.7
OpenCLIP-L/14 HCT w/ RegToken (Ours) 20.3 287 0.42 3.8

Ablation Study

Ablation experiments evaluate the feature norm separation achieved across different register neuron budgets and examine prior transferability across varying image source conditions.

Table 1: Feature norm and attention separation (validating the efficacy of register subspaces in absorbing outlier activations):

Backbone Register Localization Config Norm Gap Attention Gap
OpenCLIP Vanilla baseline (50 neurons) 62.03 0.58
OpenCLIP NFN-guided (8 neurons) 61.36 0.52
OpenCLIP NFN-guided (16 neurons) 64.42 0.55
OpenCLIP NFN-guided (50 neurons) 70.09 0.58
DINOv2 Vanilla baseline (50 neurons) 468.83 0.60
DINOv2 NFN-guided (8 neurons) 485.23 0.59
DINOv2 NFN-guided (16 neurons) 494.63 0.63
DINOv2 NFN-guided (50 neurons) 633.71 0.66

Table 2: Impact of prior source on generation performance (demonstrating that registers convey genuine class- and instance-level global semantics):

Prior Input Source FID-5k (↓) IS (↑) CLIP (↑) SigLIP (↑)
Same-image register prior 20.1 289 0.45 3.9
Cross-image prior (same class) 20.6 285 0.41 3.8
Random-image prior (any class) 21.1 283 0.40 3.7

Key Findings

  • Substantial Low-Frequency Concentration: 1D FFT spectral analyses on DINOv2 and OpenCLIP demonstrate that test-time register features exhibit significantly higher low-frequency energy ratios (\(\rho_{1:15}\) reaching 0.507 and 0.524, respectively) and superior low-band spectral flatness (\(\bar{\text{dB}}_{\text{low}}\)) compared to both [CLS] tokens and patch-mean embeddings (\(p < 10^{-16}\)), confirming that registers capture macro-level scene foundations rather than fine high-frequency noise.
  • Outlier Suppression and Accelerated Mixing: NFN-localized neurons achieve a more substantial reduction in outlier patch proportions (a median drop of -0.0312 on DINOv2 compared to -0.0215 for vanilla test-time registers) and a marked decrease in the Markov chain second eigenvalue (\(\Delta\lambda_2 = -0.0946\)), facilitating thorough attention mixing and preventing attention collapse into local singular points.
  • Enhanced Test-Time Optimization Efficiency: In HCT decoding optimization, RegToken slashes the iterations required to hit the target CLIP similarity threshold from 74 down to 52 steps (nearly 30% speedup), exhibiting markedly lower variance across 1,000 random seeds and validating its role as a robust global initialization prior.

Highlights & Insights

  • Repurposing Computational Artifacts into Generative Assets: Rather than viewing attention sinks solely as nuisance artifacts that require suppression or smoothing, this work unveils their role as low-frequency energy reservoirs and harnesses them as compact global control signals for frozen decoders.
  • Training-Free Orthogonal Conservation Formulation: The projection-and-conservation update elegantly maintains first-order energy conservation between local outlier patches and global register tokens, resolving spatial attention collapse without losing scene context.
  • Seamless Plug-and-Play Versatility: Requiring zero modifications to pretrained backbone weights and no retraining of generative decoders, RegToken transfers reliably across self-supervised ViT families (DINOv2, OpenCLIP) and diverse generative frameworks (HCT, MaskGIT-VQGAN).

Limitations & Future Work

  • Dependence on Precomputed Reference Sets: Automated layer identification using NFN requires forward inference across a small reference subset (~1,000 validation images) to establish median benchmarks, preventing fully autonomous on-the-fly layer discovery for single out-of-domain instances.
  • Information Capacity Limits of a Single Token: Currently, register representations are condensed into a single global token (\(z_0\)). For complex scenes featuring multiple foreground objects and divergent lighting sources, a single token struggles to encode spatially heterogeneous priors.
  • Future Directions: Exploring multi-token register representations with spatial disentanglement, as well as extending training-free register priors to latent attention sinks in continuous Diffusion Transformers (DiT).
  • vs Darcet et al. (ICLR 2024, Vision Transformers Need Registers): Darcet et al. introduce learnable register tokens during pretraining to soak up attention outliers; in contrast, RegToken is entirely training-free, extracting latent sink features from frozen backbones and repurposing them as generative priors.
  • vs Jiang et al. (NeurIPS 2025, Vision Transformers Don't Need Trained Registers): Jiang et al. propose heuristic test-time registers (TTR) primarily for feature smoothing and linear probing; RegToken introduces NFN and TokenRank for principled localization and unlocks their utility in generative downstream tasks.
  • vs Beyer et al. (ICML 2025, HCT): HCT generates images using compact 1D discrete tokens but suffers from slow, high-variance test-time optimization due to weak global priors; RegToken supplies a well-matched low-frequency initialization that accelerates convergence and elevates generative fidelity.

Rating

  • Novelty: ⭐⭐⭐⭐ [Pioneering reuse of attention sinks and test-time registers as generative global priors]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive spectral diagnostics, norm separation ablations, and ImageNet generative evaluations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Rigorous problem formulation bridging empirical observation, theoretical metrics, and practical generation]
  • Value: ⭐⭐⭐⭐ [Provides an effective, training-free paradigm for feature repurposing in compact generative architectures]