Skip to content

Tokenizer-Generator Coupling in Medical Image Generation

Conference: NeurIPS2026
arXiv: 2608.07713
Code: https://github.com/liamchalcroft/tokenizer-generator-coupling
Area: Medical Imaging / Image Generation
Keywords: image tokenizers, discrete quantization, generator coupling, sampling budget, modelability

TL;DR

This controlled factorial study treats the tokenizer–generator–sampler triple as its experimental unit: across 70 generation cells on ChestMNIST-64, reconstruction quality does not reliably predict generation quality, quantizer rankings depend on the generator, and validation-selected reduced-step sampling lowers LFQ-1024 + D3PM FID-192 from 0.44 to 0.09.

Background & Motivation

Latent medical image generation typically trains a VQ or KL-regularized autoencoder first, freezes it, and then trains a generator on its latents. Tokenizers are often selected using reconstruction PSNR, SSIM, or LPIPS before diffusion models are compared. This workflow assumes that the representation preserving image information best is also the easiest for the downstream generator to learn. Natural-image studies already suggest that these properties can diverge, but low-resolution medical images have more concentrated anatomical structure; whether the same relationship holds requires direct experiments rather than transferring the conclusion unchanged.

A learned VQ codebook, LFQ binary sign codes, and FSQ scalar grids can produce similar reconstructions while inducing different categorical token distributions. Autoregressive prediction, confidence-based unmasking, and absorbing-state diffusion learn those distributions differently. Changing only the tokenizer while fixing the generator, or vice versa, can mistake a favorable pairing for a component's universal advantage. Even with a fixed checkpoint, sampling temperature and step count can change the ranking, so claims that one model outperforms another must specify the inference budget and sampling rule.

The paper does not introduce a single new unified network. Instead, it places these choices within a common experimental protocol: match discrete vocabulary sizes and spatial grids, compare generator training recipes, and select sampler settings on a separate validation split. Core Idea: a tokenizer's value for generation depends on its combination with the generator, sampler, and inference budget; reconstruction fidelity or a single token statistic cannot replace empirical evaluation of that triple.

Method

Overall Architecture

The input is a \(64\times64\) single-channel image. A tokenizer compresses it into an \(8\times8\) latent grid, a generator trains on representations from the frozen tokenizer, and sampled latents are reconstructed by the corresponding decoder. The main discrete panel crosses VQ/LFQ/FSQ, three vocabulary sizes of 1,024/2,048/4,096, and six generator families, yielding 54 cells. Eight distinct VAE/AE settings paired with LDM/RF provide 16 continuous reference cells, for 70 cells in total.

The continuous channel and KL sweeps share the \(c=4,\lambda_{\mathrm{KL}}=10^{-6}\) corner, which counts as one setting rather than two. Every cell in the full grid has only one training seed. Only the vocabulary-1,024 block crossing the three quantizers with AR/MaskGIT/D3PM, plus LFQ + SEDD, retrains generators at seeds 42/43/44 while keeping the tokenizer frozen.

The diagram represents the study's training, sampling, and evaluation workflow, not a newly proposed network architecture. Dashed edges indicate training objectives or validation selection constraining later stages; solid edges show the flow of latent representations and generated samples.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    X["Real training images"] --> A["Matched Quantization Panel"]
    L["Reconstruction and regularization objectives"] -.-> A
    A -->|Freeze encoder and decoder| B["Crossed Generator Recipes"]
    O["Family-specific generation objectives"] -.-> B
    B -->|Fixed checkpoint| C["Validation-Based Sampler Selection"]
    V["Separate validation split"] -.-> C
    C -->|Sample and decode| D["Layered Evidence Evaluation"]
    D --> Y["Triple-level comparison"]

Key Designs

1. Matched Quantization Panel: separate quantization mechanisms before interpreting capacity

All tokenizers use a common convolutional encoder–decoder backbone with channels 64/128/256, two residual blocks per resolution, single-headed self-attention at \(16\times16\), and spatial compression by a factor of 8. The shared backbone has approximately 8.5M parameters. What is controlled is the backbone and spatial grid, not equal information content across all latent representations. Each position holds one discrete token, so residual quantization is excluded from the main comparison: dependencies between multiple codes would change the downstream modeling problem.

VQ selects entries from a learned codebook using cosine similarity, updates the codebook by EMA, passes gradients through a straight-through estimator, and constrains encoder outputs with a commitment loss. Code vectors are L2-normalized before decoding. LFQ converts dimension-wise signs into a categorical token without a learned codebook; its regularizer lowers per-position assignment entropy while raising batch-averaged entropy, encouraging confident assignments and broader code coverage. FSQ bounds each dimension and independently rounds it to fixed levels, with their joint configuration determining the token. It has no auxiliary quantization loss, but its code frequencies are not constrained by LFQ-style entropy regularization.

At vocabulary 1,024, LFQ uses 10 binary dimensions and FSQ uses five four-level scalars. FSQ uses level vectors (8,4,4,4,4)/(8,8,4,4,4) at 2,048/4,096. The matched vocabulary sizes control raw discrete token bits rather than actual entropy-coded bitrate:

\[ R_{\mathrm{raw}}=64\log_2 K\in\{640,704,768\}\ \text{bits/image}. \]

The continuous references sweep \(c\in\{1,2,4,8,16\}\) at fixed \(\lambda_{\mathrm{KL}}=10^{-6}\), and, at fixed \(c=4\), KL weights of \(10^{-5}/10^{-6}/10^{-7}/0\), with zero corresponding to AE. Their bit counts use \(c\times8\times8\times16\) float16 storage as a proxy. This is not rate-matched to the discrete representations and does not measure clinical information content.

This control supports asking which quantizer suits a given generator at equal vocabulary size, but not declaring discrete representations more efficient than continuous ones. All three discrete methods achieve 100% code utilization at vocabulary 1,024, with normalized marginal entropies of 0.997 for VQ, 0.999 for LFQ, and 0.959 for FSQ. Thus, the comparison is not simply one against a collapsed VQ codebook. Every discrete generator still operates on flattened categorical IDs, without explicitly exploiting LFQ bit structure or FSQ scalar factorization.

2. Crossed Generator Recipes: expose the same latent distribution to different generation mechanisms

The common DiT-style backbone has 12 layers, width 512, eight attention heads, RoPE, and SiLU, but this is not an equal-parameter or equal-training-budget experiment. AR/MaskGIT omit timestep modulation and have 38.9M parameters at vocabulary 1,024. DFM/D3PM/SEDD/BFN include adaLN-Zero time conditioning and have 57–58M; continuous LDM/RF have 57.1M. Learning rates and training lengths also vary by family, so cross-generator differences describe full recipes rather than a pure causal effect of architecture.

AR performs next-token prediction in raster order, requiring 64 forward calls for 64 positions while reusing prefix computation through KV caching. MaskGIT starts from an all-mask grid, predicts unresolved positions in parallel, and commits tokens according to confidence perturbed by Gumbel noise, using 12 rounds by default. Its conditional marginal predictions differ from AR's strictly ordered joint factorization. The comparison therefore contrasts sequential modeling with bidirectional, confidence-driven parallel commitment.

D3PM progressively absorbs tokens into MASK under a cosine schedule, trains the network to predict original categories, and executes reverse transitions using a closed-form posterior. Here, SEDD shares D3PM's forward process, backbone, and reverse sampler but uses a simplified score-entropy objective with a cross-entropy stabilizer weighted 0.001. It is not canonical SEDD with tau-leaping, so similar results here cannot establish equivalence between the original algorithms.

DFM trains discrete jump velocities along a continuous-time MASK-to-data probability path, using a cubic schedule and 100 Euler steps by default. BFN instead maintains a categorical probability distribution at each position, predicts a receiver distribution, and updates the state by Bayesian updates. Its accuracy-schedule scale is set to \(\beta=\sqrt{2\ln K}\), with 1,000 default steps. The panel therefore compares different state representations and update dynamics, not merely different losses.

Continuous LDM predicts Gaussian noise in VAE/AE latents and uses 1,000 stochastic DDPM reverse steps by default. RF predicts velocity along straight noise-to-data paths and integrates 100 ODE steps by default. Training latents are extracted using VAE posterior means rather than fresh posterior draws each epoch. AE and VAE-c1 use latent normalization estimated from the training split, followed by inverse normalization before decoding; other continuous rows use raw latents. This nonuniform normalization also limits attribution to the training objective alone.

3. Validation-Based Sampler Selection: default step counts are not model ceilings

The main table uses temperature 0.9 and each generator's default step budget. With trained checkpoints fixed, the authors separately sweep temperature, step count, top-k, or mask schedule for the three vocabulary-1,024 quantizers. Each candidate generates 2K samples and is evaluated against the 11,219-image official validation split. Only after selection is the configuration evaluated once using 10K generated test samples, so the reported selection does not directly optimize the test score.

The number of candidates per tokenizer is 18 for AR, 60 for MaskGIT, 25 for DFM, 20 each for D3PM and SEDD, and 30 for BFN. Search sizes are unequal, but the large D3PM/SEDD gains do not come from larger grids. The real-vs-real FID-192 noise floor for the 2K validation screen is 0.010, leaving small improvements potentially unstable. Validation selection does not remove finite-sample or model-selection uncertainty.

For reduced-step D3PM/SEDD sampling, a fresh \(N\)-step absorbing process is constructed within the same cosine schedule family. This is neither arbitrary respacing of the original 1,000-step chain nor retraining. Already unmasked positions retain their categories; at a masked position, the predicted clean-category distribution is multiplied by the following unmasking probability, with the remaining mass assigned to MASK:

\[ p(x_{n-1}=j\mid x_n=\mathrm{MASK})= p_\theta(x_0=j\mid x_n,n)\frac{\beta_n\bar\alpha_{n-1}}{1-\bar\alpha_n},\qquad j<K. \]

Here, \(\bar\alpha_n\) is the cumulative retention probability, and SEDD receives normalized time \(n/N\). This form also avoids storing dense D3PM transition matrices at every timestep: transition-matrix storage falls from \(O(TK^2)\) to \(O(T)\), but each prediction still involves \(O(BLK)\) categorical computation. Logits, embeddings, activations, and optimizer state remain vocabulary-dependent. Matrix-free does not mean vocabulary-independent total compute or memory.

Validation-selected LFQ D3PM uses temperature 0.7 and 100 steps, while SEDD uses temperature 0.7 and 500 steps. Both FSQ variants select temperature 1.0 and 250 steps. VQ selects temperature 0.5 and 500 steps for D3PM, and temperature 0.7 and 250 steps for SEDD. The dependence of the best step budget on the tokenizer directly supports treating the triple, rather than an isolated component, as the comparison unit.

4. Layered Evidence Evaluation: separate estimator noise, training variance, and clinical meaning

Test evaluation uses EMA weights, generates 10K images per cell, and compares them with 22,433 ChestMNIST test images. FID-192 uses globally averaged 192-dimensional features from the second max-pooling block of an ImageNet-pretrained InceptionV3. It is an internal ranking metric, not clinical accuracy, and its absolute values are not directly comparable with conventional FID-2048 literature.

At 10K samples, the real-vs-real FID-192 bootstrap has mean 0.002 and a 95% interval of [0.001,0.004]. This measures estimator noise, not training-seed variance. A descriptive sum-of-squares decomposition of log(FID-192) over the 54 default discrete cells assigns 59.5% to generator family, 10.1% to quantizer, 8.4% to quantizer–generator interaction, and 15.0% to vocabulary–generator interaction. With only one training seed per cell, this is not factorial inference with an experimental-error term and significance tests; residual interactions cannot be treated as random noise.

Actual training replication covers only the vocabulary-1,024 block described above. Four of nine pairwise quantizer comparisons exceed three pooled seed standard deviations, supporting quantizer–generator non-separability; the other five do not receive reliable rankings. FID-2048 computed on the same images preserves LFQ's D3PM advantage but can change the AR and MaskGIT winners. The appropriate conclusion is that interaction exists, not that one global quantizer leaderboard is established.

Additional checks use domain-FID, a label-free classifier two-sample test, and nearest-neighbor screening. Domain-FID covers only six representative cells, with correlation 0.71 and exact permutation \(p=0.14\); it cannot independently validate fine-grained rankings and does not validate all tuned results. The nearest-neighbor ratio divides median generated-to-training nearest-neighbor distance by median within-training nearest-neighbor distance. AuthPct measures the fraction of generated images closer to their nearest training neighbor than that neighbor is to its own nearest neighbor. These measures screen for copying in particular distance spaces; they do not certify privacy.

A Worked Example

For LFQ-1024 paired with D3PM, the frozen tokenizer converts real training images into 64 discrete positions, and the generator learns their token distribution. Validation compares temperatures and reverse-process step counts without retraining. After selecting temperature 0.7 and 100 steps, the pipeline generates and decodes 10K test images for evaluation.

Table 7 reports FID-192 decreasing from the default 0.44 to 0.09, illustrating that default sampling is not this checkpoint's performance ceiling. This is the data flow of a reported configuration, not diagnosis of a particular chest image; lower FID does not certify clinical structure or patient privacy.

Loss & Training

Tokenizers train for 50 epochs with L1 reconstruction loss weighted 4.0, VGG-16 perceptual loss weighted 0.5, a multi-scale PatchGAN weighted 0.05 starting at epoch 10, and LeCam regularization weighted 0.001. Applicable family-specific terms are added; FSQ has no auxiliary quantization loss. VAE KL weights receive a 10-epoch linear warmup.

Inputs are min–max normalized to [0,1] without data augmentation. Tokenizers use AdamW at learning rate \(10^{-4}\), batch size 128, and checkpoint selection by total validation loss. Generators then read token files or float16 continuous latent files from the frozen tokenizer, without jointly updating it.

Discrete generators train for 100 epochs and continuous generators for 200. AR/MaskGIT/DFM use learning rate \(3\times10^{-4}\); the others use \(10^{-4}\). Common settings include AdamW, cosine learning-rate decay, 5% warmup, EMA 0.9999, gradient clipping at 1.0, and BFloat16. The main grid trains on one 48GB A6000, with some replication experiments on an L40S. These shared settings do not eliminate parameter and recipe differences.

Key Experimental Results

Main Results

The following selection from paper Table 5 uses ChestMNIST-64, 10K generated test samples, default sampling, and one training seed. FID-192 is lower-is-better. Continuous references are comparable only at the full-recipe level.

Tokenizer AR MaskGIT DFM D3PM SEDD BFN LDM RF
VQ-1024 0.52 1.86 4.42 1.30 1.31 5.98 — —
LFQ-1024 0.33 1.91 1.77 0.44 0.41 2.27 — —
FSQ-1024 0.39 1.15 3.02 1.21 1.05 7.60 — —
VQ-4096 0.31 2.03 3.39 3.69 3.54 4.22 — —
VAE-c8 — — — — — — 1.61 0.32
AE — — — — — — 0.09 0.07

For reconstruction, LFQ-1024 has PSNR 28.0 dB and FSQ-1024 has 29.4 dB, yet this does not imply that FSQ is better under D3PM. Continuous VAE-c8 reaches PSNR 34.4 dB but is the worst generator setting in the channel sweep. Spearman correlation between PSNR and each tokenizer's best default generation FID is −0.03 for the nine discrete tokenizers and +0.10 for the continuous channel sweep.

Paper §4.2 calls LFQ-1024 + AR at 0.33 the best discrete result, although Table 5 and §4.1/§6.1 explicitly report VQ-4096 + AR at 0.31. Both table values are retained here rather than silently reconciled. The authors also note that their small difference does not exceed observed AR seed variability, preventing a reliable ranking.

Ablation Study

These ablations mainly change sampler settings at fixed checkpoints rather than remove a proposed module. Paper Table 7 reports 10K-sample test scores, with candidate settings first screened using 2K validation samples.

Generator LFQ-1024: default → tuned FSQ-1024: default → tuned VQ-1024: default → tuned
AR 0.33 → 0.33 0.39 → 0.29 0.52 → 0.46
MaskGIT 1.91 → 1.89 1.15 → 0.88 1.86 → 1.42
DFM 1.77 → 1.58 3.02 → 1.47 4.42 → 3.13
D3PM 0.44 → 0.09 1.21 → 0.13 1.30 → 0.15
SEDD 0.41 → 0.10 1.05 → 0.10 1.31 → 0.13
BFN 2.27 → 2.13 7.60 → 5.80 5.98 → 4.10

The next table selects three-seed default-sampling results from paper Table 9, retaining the source precision for means and standard deviations. These are not the same evaluation pass as the preceding tables, so recomputed values should not be substituted into the main table.

Vocabulary-1,024 cell Three-seed mean FID-192 Seed standard deviation
VQ + MaskGIT 2.16 0.214
LFQ + MaskGIT 1.92 0.052
FSQ + MaskGIT 1.15 0.012
VQ + D3PM 1.32 0.043
LFQ + D3PM 0.39 0.022
FSQ + D3PM 1.20 0.042

Key Findings

  • Under FID-192, MaskGIT favors FSQ and D3PM favors LFQ; this reversal is supported within the three-seed block. However, FID-2048 changes the MaskGIT winner. The robust conclusion is coupling, with LFQ's D3PM advantage preserved across metrics.
  • Tuned LFQ + D3PM achieves FIDs of 0.09/0.15/0.10 across three seeds, and SEDD achieves 0.10/0.11/0.13; both average 0.11. Each seed is tuned independently. Applying seed 42's D3PM setting unchanged to seed 43 gives 0.34, so specific hyperparameters cannot be claimed to transfer without reselection.
  • Tuned LFQ D3PM reaches 33.2 samples/s versus 3.67 by default, while tuned SEDD reaches 6.60 versus 3.67. Cached AR remains at 146, illustrating that fewer NFE do not imply equal wall-clock cost. Timings include sampling and decoding but exclude metric computation; default and tuned batch sizes are 128/64.
  • Replicating six representative cells on two additional datasets gives AE + RF FIDs of 0.07/1.05/0.94 for Chest/Pneumonia/OrganAMNIST. On Pneumonia, VQ + AR at 2.09 beats LFQ + AR at 2.73, yet subsampling Chest training data to 4,700 still favors LFQ at 0.95 over VQ at 1.11, rejecting a sample-size-only explanation for the reversal.

Highlights & Insights

  • Including the sampler in the experimental unit prevents an unsuitable default 1,000-step budget from being mistaken for an intrinsic weakness of discrete diffusion. Quality and cost can change substantially without retraining, making sampler evaluation worthwhile before increasing training compute.
  • Shared backbones and matched vocabularies improve interpretability without being presented as strictly compute-matched experiments. The central value is identifying pairing effects rather than declaring a universally best model.
  • Marginal entropy, per-position entropy dispersion, and left/upper-neighbor conditional predictive gains do not reliably predict the best generation FID. Generator-free compression statistics are insufficient proxies for modelability under a particular generator and budget.

Limitations & Future Work

  • Most evidence concerns unconditional \(64\times64\) ChestMNIST, and replication on the other two MedMNIST datasets covers only six cells. Library support for high resolution and 3D is not validation in those settings, nor evidence of diagnostic fidelity in real clinical imaging.
  • The entire 70-cell panel is not a three-seed experiment, and the continuous-channel anomaly may reflect training variability. Further work should expand seed replication and standardize continuous-latent normalization before separating capacity, objective, and noise-schedule effects.
  • Tuned discrete diffusion at 0.09/0.10 approaches default continuous references, but continuous models receive no equivalent sampler search and train for different numbers of epochs. This is not a final conclusion that discrete generation matches continuous generation under fair budgets.
  • FID-192 correlates with FID-2048 at 0.89 across 12 cells and with label-free two-sample discriminator rankings at 0.86; the six-cell domain-FID check is weaker. Radiologist assessment, pathology-conditional generation, downstream utility, and more medical-specific feature metrics remain necessary.
  • Feature-space nearest-neighbor screening gives a minimum ratio of 0.975 and maximum AuthPct of 0.107 without clearly detecting copying. However, AE + LDM has raw pixel/LPIPS ratios of 0.873/0.821, motivating stronger audits. Without membership inference or patient-level duplication checks, these results cannot certify privacy.
  • Counterfactual inpainting with a known region is demonstrated on six synthetic bright-ellipsoid examples with IoU 0.31–0.96. The region is supplied in advance: this is oracle-mask inpainting, not blind anomaly localization or clinical anomaly-detection validation.
  • vs VQ-VAE / MAGVIT-v2 / FSQ: These methods supply different quantization mechanisms; this paper contributes matched vocabularies crossed with multiple generators rather than a new quantization head. It does not test whether bitwise LFQ or factorized FSQ output heads better exploit latent geometry.
  • vs canonical SEDD / D3PM: The evaluated SEDD is a specific hybrid score-entropy loss paired with a D3PM sampler. Matrix-free absorbing-state computation is reusable, but conclusions about training objectives and reduced steps remain bounded by this interface and schedule.
  • vs rate-distortion-perception / modelability: The study supports the empirical interpretation that reconstruction alone cannot select a tokenizer for generation. It does not prove a universal law or a capacity threshold for medical images. Direct extensions include reconstruction FID in a common feature space, generator-to-reconstruction FID gaps, and generator-conditional prediction errors.
  • vs MedVAE / MedITok / CheXGenBench: The first two emphasize medical representations, while the latter emphasizes fidelity, privacy, and utility in end-to-end synthesis. This paper provides a component-level study protocol, not a replacement for clinical or application-oriented evaluation axes.
  • Implementations are available in medtokenizers and medlatents. The reusable contribution is the interchangeable tokenizer–generator workflow and validation-based sampler selection, not direct deployment of these low-resolution unconditional models.

Rating

  • Novelty: 4/5 — Studies tokenizer, generator, and sampler coupling in controlled medical-image experiments; the contribution is primarily experimental perspective rather than a new architecture.
  • Experimental Thoroughness: 4/5 — Includes 70 cells, partial three-seed replication, and multiple robustness checks, but lacks complete replication, unified budgets, and clinical validation.
  • Writing Quality: 4/5 — Clearly states recipe, seed, and metric boundaries, with a remaining inconsistency in the description of the best default discrete result.
  • Value: 4/5 — Offers a practical selection protocol and reusable implementations for latent generation, with conclusions limited to this low-resolution study setting.