Phaedra: Learning High-Fidelity Discrete Tokenization for the Physical Sciences¶
Conference: NeurIPS2026 (queue assignment; full text: arXiv v2, 2026-09-29)
arXiv: 2602.03915
Area: Physics / Scientific Computing
Keywords: discrete tokenizer, amplitude–morphology factorization, finite scalar quantization, partial differential equations, scientific data compression
TL;DR¶
Phaedra separates physical-field latents into multidimensional morphology tokens and high-precision one-dimensional amplitude tokens, then learns to recombine and decode them, reducing in-distribution PDE reconstruction nMAE from FSQ's 2.603 to 1.522 and compressing 65.0 GB of scientific data to 3.44 GB.
Background & Motivation¶
Representing PDE solutions as discrete tokens enables cross-entropy training, masked prediction, and Transformer sequence modeling while reducing the storage and transfer costs of training data. Physical fields, however, are not ordinary grayscale images: the geometry of vortices and shocks matters, but so do the absolute magnitudes of velocity, density, and pressure. Perceptual losses for natural images permit textures that merely look plausible, whereas scientific computing requires accurate errors, gradients, and spectra; both hallucinated high-frequency details and excessive smoothing can compromise subsequent dynamics prediction.
Scientific data also retain broad dynamic ranges and heavy-tailed distributions after standardization. The same local shape can occur at different intensities, so a single discrete code representing both shape and amplitude must allocate vocabulary capacity to many same-shape, different-intensity combinations. Adding a multiscale hierarchy does not directly resolve this bottleneck: it separates coarse and fine scales without specifically providing a dense coordinate for continuously varying physical amplitudes.
Phaedra draws inspiration from classical shape-gain quantization but neither specifies analytical basis functions nor forces shape and gain to multiply. Its quantization structure induces complementary roles in two branches, and a convolutional decoder learns their interaction. Core Idea: preserve reusable local morphology with a multidimensional discrete vocabulary and amplitude with a dense one-dimensional scalar vocabulary, preventing geometric diversity and numerical precision from competing for the capacity of a single code.
Method¶
Overall Architecture¶
The input is one physical variable on a two-dimensional grid; multivariable fields are processed independently by channel, so the tokenizer always receives 1 input channel. A shared convolutional encoder reduces spatial resolution, followed by dual-latent factorization, split-FSQ, and learned recombination; a convolutional decoder restores the original grid. The standard configuration maps a \(128\times128\) field to a \(32\times32\) latent grid, producing one morphology token and one amplitude token per location, or 2048 discrete codes in total.
The morphology branch has 8 latent channels and the amplitude branch has 1; this splits encoded latent features by channel rather than dividing the raw input into two images. Both branches are quantized in parallel. During decoding, token lookup retrieves the corresponding quantized vectors, which are concatenated and recombined by a channel-mixing convolution. Reconstruction supervision is used only during training; reconstruction at inference time requires neither ground-truth field labels nor a PDE solve.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
X["Single-variable physical field"] --> E["Convolutional encoder"]
E --> F["Dual-latent factorization"]
F --> Q["Split-FSQ<br/>Morphology 8D / amplitude 1D"]
Q --> R["Learned recombination"]
R --> D["Convolutional decoder / reconstructed field"]
X -.->|Training only: L1 reconstruction supervision| D
F -.->|Training only: latent commitment constraint| Q
Key Designs¶
1. Dual-latent factorization: give amplitude variation an independent representation path
Shock boundaries can have similar positions and shapes but different pressure intensities; placing every variation in the same multidimensional code makes quantization damage both boundary structure and amplitude precision. Phaedra divides encoder outputs into an 8-dimensional morphology latent and a 1-dimensional amplitude latent, supporting structural representation and dense numerical scales separately. The analogy of morphology as basis functions and amplitude as coefficients is explanatory: there is no predefined physical basis and no additional disentanglement label.
This distinguishes Phaedra from VQ-VAE-2, which factorizes information into coarse and fine spatial scales rather than semantic channel roles on a shared latent grid. It also differs from a sequential approach that first reconstructs a low-frequency component and then quantizes the residual; parallel encoding lets the morphology branch use the original field rather than depend entirely on what the first stage leaves behind. The branches' roles arise primarily from their different quantization configurations, not an explicit loss guarantee, and therefore constitute soft disentanglement.
2. Split-FSQ: use multidimensional combinations for morphology and dense scalar levels for amplitude
The morphology branch uses Finite Scalar Quantization (FSQ), with \([5,4,4,3,3,3,2,2]\) levels across its 8 channels; their Cartesian product yields 8640 possible morphology codes. Each dimension has few levels, but their combinations distinguish diverse local geometric patterns; this configuration is identical to the single-stream FSQ baseline. The amplitude branch instead uses 1024-level one-dimensional FSQ, allowing intensity to vary finely along an ordered scalar axis without learning 1024 unrelated high-dimensional patterns.
These vocabularies do not form a single 9664-class joint classifier: each location stores two indices, one selected from 8640 classes and one from 1024 classes. This assigns fine intensity subdivision and combinatorially rich morphology to suitable quantization geometries. The equal-cardinality ablation replaces the one-dimensional amplitude vocabulary with five-dimensional FSQ at \([4,4,4,4,4]\) levels; it retains 1024 codes but worsens reconstruction, showing that code arrangement is not equivalent to code count.
FSQ applies a bounded nonlinearity before rounding, which can place much of the latent distribution in the saturated region of tanh if used directly. The paper multiplies by a scale factor before quantization and divides by the same factor afterward to avoid mismatched numerical domains before and after quantization. The morphology and amplitude scale factors are 10 and 0.1, respectively; the smaller amplitude factor keeps more of the bulk distribution in tanh's approximately linear region. This remains a finite-range representation, not a genuinely unbounded scalar encoding: extreme values are softly bounded, and the model does not gain unlimited dynamic range.
3. Learned recombination: let the decoder learn nonlinear amplitude–morphology interactions
After retrieving both groups of quantized vectors, the model concatenates them and passes them through a learnable channel-mixing convolution to the decoder rather than imposing morphology-times-amplitude multiplication. Fixed multiplication is interpretable but may not suit a complex latent space: different local structures may require different amplitude modulation, which convolutional recombination can learn from reconstruction supervision. Both encoder and decoder use residual convolutional backbones; a single attention layer is included when the latent grid is no larger than \(32\times32\), and nearest-neighbor upsampling restores resolution in the decoder.
This flexibility means amplitude and morphology cannot be interpreted as fully independent physical quantities. After amplitude-token swapping, reconstructed energy correlates with the amplitude-source field at 0.99; swapping only morphology tokens leaves energy correlated with the original field at 0.97. These interventions support the claim that the amplitude branch primarily controls macroscopic energy, but they are not a strict factorization theorem. The appendix also shows that zeroing amplitude embeddings produces approximately 300% high-frequency energy retention: nonphysical artifacts rather than stronger detail recovery.
A Worked Example¶
For a \(128\times128\) density snapshot, encoding produces a \(32\times32\) grid with an 8-dimensional morphology representation and one amplitude scalar at each location. The morphology vector maps to one of 8640 combination codes and the amplitude scalar to one of 1024 levels; the stored representation is two \(32\times32\) index maps rather than 9-channel floating-point features. During reconstruction, each index map retrieves quantized embeddings, and learned recombination plus the decoder restores the density field, with amplitude supporting large-scale intensity and morphology adding local boundaries and vortex structure.
Considering only tightly packed theoretical payloads, each latent location requires 10 bits for amplitude and 14 bits for morphology, or 24 bits in total. Relative to a raw \(128\times128\) 32-bit single-channel array, this saves approximately 95.3%; the actual dataset shrinks from 65.0 GB to 3.44 GB, corresponding to 94.7% measured savings. This measures data storage, not an equivalent reduction in model parameters, and it does not imply a shorter Transformer input sequence than FSQ.
Loss & Training¶
The tokenizer uses neither perceptual losses nor an explicit PDE residual loss; its main objective is pointwise \(L_1\) reconstruction, with constraints keeping encoder latents close to their quantized versions.
Here \(\operatorname{sg}\) denotes stop-gradient; dense amplitude quantization already gives a small commitment term, so it is not additionally multiplied by 0.25. The commitment loss does not label either branch as morphology or energy; it prevents encoder outputs from moving far outside the quantizer's effective range. The standard model has approximately 97M parameters and uses AdEMAMix with a base learning rate of \(10^{-4}\), weight decay of 0.01, EMA decay of 0.999, and 1 training epoch. Pretraining uses six Compressible Euler and Incompressible Navier-Stokes datasets with approximately 4.8M training samples; the main reconstruction comparison shares the backbone and training recipe.
Downstream experiments freeze the tokenizer and train an approximately 38M-parameter sequence-to-sequence Transformer to predict future physical states. For operator learning, two output heads predict categorical distributions over morphology and amplitude codes using cross-entropy; all2all training organizes 7 time points into 28 input–output combinations, with time-difference embeddings and two-dimensional RoPE. Masked autoencoding instead uses a single decoder for both token types, randomly masking 75% of spatial locations with the same spatial mask across physical variables. Predicted downstream tokens are converted to physical fields by the frozen decoder; this is distinct from evaluating tokenizer reconstruction alone.
Key Experimental Results¶
Main Results¶
The following excerpts from Tables 2 and 7 average across variables and retain the source reporting scale; nMAE is not relabeled as relative \(L_1\) error. ID denotes in-distribution data; OD1 changes initial conditions or boundary geometry; OD2 includes unseen PDEs such as Poisson, Darcy, Allen-Cahn, and Acoustic Wave.
| Model | Distribution | nMAE ↓ | nRMSE ↓ | Local variance error ↓ | Minimum spectral coherence ↑ |
|---|---|---|---|---|---|
| Continuous AE | ID | 0.672 | 1.122 | 2.98 | 98.4% |
| FSQ | ID | 2.603 | 4.292 | 11.29 | 85.3% |
| Phaedra | ID | 1.522 | 2.489 | 5.96 | 93.6% |
| Continuous AE | OD1 | 1.154 | 2.079 | 3.40 | 99.2% |
| FSQ | OD1 | 1.878 | 4.331 | 20.55 | 93.9% |
| Phaedra | OD1 | 1.224 | 2.435 | 6.47 | 98.1% |
| Continuous AE | OD2 | 1.967 | 2.821 | 3.02 | 97.0% |
| FSQ | OD2 | 4.865 | 6.675 | 11.66 | 64.3% |
| Phaedra | OD2 | 3.147 | 4.237 | 5.83 | 77.6% |
nMAE divides mean absolute residuals by the variable's global standard deviation over the training data; nRMSE divides root mean squared residuals by the same standard deviation. Local variance error first computes variance maps with a \(7\times7\) sliding window, then measures their maximum discrepancy:
Minimum spectral coherence is neither energy-spectrum error nor topological accuracy; it measures the lowest coherence between ground-truth and reconstructed fields across frequency bands. With \(G\) denoting auto- or cross-spectra, the paper defines:
Relative \(L_1\) instead divides summed absolute residuals by summed absolute ground-truth values; downstream tables report percentages, which must not be conflated with nMAE.
Ablation Study¶
The following CEU RC density results come from Tables 12, 15, and 16, all at \(4\times4\) downsampling; these are single-variable results, not the cross-variable means above.
| Config | nMAE ↓ | Local variance error ↓ | Minimum spectral coherence ↑ | Note |
|---|---|---|---|---|
| FSQ | 8.4078 | 20.8387 | 65.41% | Morphology-only stream, 1024 tokens |
| Phaedra | 5.5397 | 9.5770 | 84.27% | Parallel amplitude–morphology, 2048 tokens |
| Codebook Ablation | 7.8021 | 19.2576 | 72.37% | Equal-cardinality five-dimensional amplitude FSQ |
| Residual Ablation | 7.2957 | 20.0013 | 74.11% | Amplitude first, then morphology residuals |
FSQ and Phaedra share the spatial grid but not token count; more targeted evidence comes from the two ablations that retain 2048 tokens and a 1024-class amplitude vocabulary. They support contributions from both one-dimensional amplitude quantization and parallel factorization, although the residual variant performs better on some smooth-data variables, so this single-variable result does not establish universal superiority.
The downstream excerpt below follows Table 3. Operator learning compares final-time predictions using each model's best of three stepping strategies, whereas MAE averages across all time points.
| Task / Model | KH relative L1 ↓ | RC relative L1 ↓ | RKH relative L1 ↓ |
|---|---|---|---|
| Operator learning / CNO | 10.15% | 29.92% | 11.82% |
| Operator learning / Continuous Transformer | 9.28% | 119.0% | 11.56% |
| Operator learning / FSQ | 11.28% | 38.05% | 17.99% |
| Operator learning / Phaedra | 9.50% | 27.21% | 10.23% |
| MAE / FSQ | 5.84% | 29.16% | 9.45% |
| MAE / Phaedra | 4.75% | 22.17% | 7.27% |
Key Findings¶
- ID nMAE is approximately 41.5% lower than FSQ, but Continuous AE still reconstructs more accurately; reducing quantization loss does not mean surpassing continuous representations.
- The continuous model degrades more in relative terms from ID to OD1, yet its absolute OD1 nMAE of 1.154 remains below Phaedra's 1.224; this does not establish comprehensive discrete-model superiority.
- At \(16\times16\) downsampling on high-resolution PDE data, Phaedra reaches nMAE 2.33 versus Cosmos's 14.53; Cosmos uses per-sample normalization, and the comparison matches neither parameter count nor pretraining data.
- An amplitude vocabulary with only 32 levels already improves over FSQ, with diminishing gains beyond approximately 512 levels; increasing latent-grid token count also helps, but returns diminish around \(32\times32\).
- Several source passages refer reconstruction and downstream results to Table 4, although their numerical tables are Tables 2 and 3; Table 4 is titled token scaling. The prose describes ID spectral coherence as near 100%, while this note retains the actual mean of 93.6%.
- Continuous Transformer RC is 119.0% in Table 3 and 119.05% in supplementary Table 27, a reporting-precision difference; the table here follows Table 3 rather than silently standardizing precision.
Highlights & Insights¶
- Quantization geometry matters more than simply enlarging the vocabulary. A set of 1024 ordered scalar levels and 1024 five-dimensional combination codes have equal cardinality but different amplitude-representation capabilities.
- Interpretability comes from intervention rather than embedding plots alone. Token swapping, branch zeroing, and Fourier low-pass/high-pass comparisons make the proposed division of roles testable while revealing artifacts caused by nonlinear recombination.
- Data compression and downstream accuracy require separate validation. Reporting measured storage savings alongside dynamics learning avoids relying solely on attractive reconstructions to argue for scientific-modeling utility.
Limitations & Future Work¶
- Independent encoding of physical variables ignores density–velocity–pressure coupling; no joint-channel variant is trained, so the associated information loss remains unquantified.
- There are no hard conservation-law, symmetry, or PDE-residual constraints; low local variance error and high spectral coherence do not guarantee physical conservation.
- Main evaluations concern two-dimensional gridded data; the appendix includes only one preliminary three-dimensional CEU RC experiment, not comprehensive coverage of three-dimensional PDE families or irregular meshes.
- Downstream evaluation remains a proof of concept, operator comparisons use each model's best stepping strategy, and Phaedra takes approximately 0.0398 seconds per sample versus FNO's 0.0126 seconds; sequence modeling is not a cost-free improvement.
- Cross-domain performance is not uniformly superior: ImageNet local variance error remains relatively high in Table 5, and Phaedra is not best on NDVI or some RGB metrics in Table 23; domain transfer should not be recast as universal image-compression superiority.
- Future directions include joint-variable encoding, conservation constraints, and geometry-aware backbones; a shared vocabulary spanning grids, dimensions, and symbolic equations remains a motivation rather than a demonstrated system.
Related Work & Insights¶
- vs FSQ: Phaedra shares morphology quantization and adds the amplitude stream plus learned recombination; this is the most direct structural baseline, but its doubled token count warrants interpreting gains alongside equal-budget ablations.
- vs VQ-VAE-2 / VAR: They factorize by scale or spatial residual, whereas Phaedra separates amplitude and morphology; the paper's VAR uses an FSQ modification, so results do not directly characterize the original VAR system.
- vs Cosmos: The pretrained image tokenizer targets natural-image statistics, while Phaedra trains specifically on physical fields with pointwise accuracy; matching compression does not match training, capacity, and token budgets simultaneously.
- vs continuous neural operators / continuous latents: Continuous AE reconstructs more accurately, while discrete representations offer compression and categorical-distribution modeling; RC/RKH downstream gains suggest representation affects learning stability, not that every continuous generative method is inferior.
Rating¶
- Novelty: 4/5 — A concise, targeted dual-stream scientific-field quantizer built from shape-gain principles.
- Experimental Thoroughness: 4/5 — Covers PDE reconstruction, transfer, interventions, and downstream tasks, but lacks joint-variable and large-scale foundation-model validation.
- Writing Quality: 3/5 — The method is clear, but some table references and generalization statements do not fully match the boundaries of the numerical evidence.
- Value: 4/5 — Provides high-fidelity representations and measured compression evidence for connecting physical fields to discrete Transformers.