Skip to content

TimeTok: Granularity-Controllable Time-Series Generation via Hierarchical Tokenization

Conference: NeurIPS2026
arXiv: 2605.01418
Area: Time Series
Keywords: granularity-controllable generation, hierarchical tokenization, finite scalar quantization, conditional flow matching, granularity-block autoregression

TL;DR

TimeTok encodes time series into discrete token prefixes with coarse-to-fine semantics, adds detail through granularity-block autoregression, and controls output granularity through conditional flow matching, achieving strong conditional refinement and distributional metrics without leading on every downstream predictive metric.

Background & Motivation

Time series contain slowly varying trends alongside local fluctuations and transient events. Generators such as TimeGAN, DiffusionTS, and TimeVQVAE mainly learn the distribution at the granularity of their training data. They typically cannot directly answer which plausible details could accompany a given coarse outline, or expose different levels of output detail through one native interface. Adding noise to a coarse signal does not solve this problem: it creates variation without necessarily matching the conditional distribution of real sequences.

The central difficulty is to give the representation an actionable granularity ordering. Conventional time-series tokens often correspond to positions or local patches; removing the latter tokens usually discards the latter part of the time interval rather than producing a coarse version of the entire sequence. Even adding a granularity label leaves the model to learn many coarse–fine relationships if its latent representation lacks hierarchical organization. Direct generation in raw sequence space instead repeatedly processes the entire preceding trajectory, increasing sequence length and learning difficulty.

The paper defines granularity-controllable time-series generation (GC-TSG) as a unified task: generate at a specified granularity from an empty condition, or refine a coarser input toward a finer target. In the main experiments, granularity is defined by information-reducing operators, with every level retaining the same sequence length; it should not be equated directly with a physical sensor sampling rate. Core idea: make token-prefix length specify the temporal detail retained across an entire sequence, then complete token blocks along the granularity axis so that coarse-outline conditioning and unconditional generation share one model.

Method

Overall Architecture

TimeTok first learns “Hierarchical Tokenization and Flow Decoding,” compressing a continuous sequence into ordered discrete tokens and enabling prefixes of different lengths to decode into different granularities. It then freezes the tokenizer and decoder and trains “Granularity-Block Autoregression” to predict the tokens added at the next level. At inference, “Prefix Selection and Target Control” retains the coarse information in an input, generates up to the requested level, and uses the learned flow decoder to transform noise into a continuous sequence.

Training uses real sequences and their smoothed versions as supervision. Inference has no true fine-scale detail as supervision: it has only a coarse input, generated tokens, and random noise. The added detail is a plausible conditional sample, not a unique recovery of the original observation. Dashed arrows below denote training supervision or reuse of the learned decoder; solid arrows denote token modeling and inference data flow.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    X["Real sequence / coarse input"] --> A["Hierarchical Tokenization<br/>and Flow Decoding"]
    S["Training smoothing targets"] -.->|Multigranular flow-matching supervision| A
    A -->|Training tokens after freezing| B["Granularity-Block<br/>Autoregression"]
    B -->|Block sampling at inference| C["Prefix Selection<br/>and Target Control"]
    A -->|Candidate prefixes of coarse input| C
    C -->|Target-level token prefix| O["Continuous sequence output"]
    N["Gaussian noise"] -->|Conditional flow decoding| O
    A -.->|Reuse learned flow decoder| O

Key Designs

1. Hierarchical Tokenization and Flow Decoding: turn token order into granularity order through prefix supervision

The encoder splits the input into non-overlapping patches, linearly projects them, and appends 128 learnable register tokens. A Transformer lets the registers aggregate information across the complete sequence. The encoded registers are discretized; the representation is not a time-ordered subset of raw measurements. Consequently, a short token prefix can describe a trend across the entire time interval rather than only the beginning of the sequence.

Finite scalar quantization (FSQ) maps each register to a six-dimensional discrete code with 4 values per dimension, yielding a combined vocabulary of \(4^6=4096\) codes. The 128 token positions must be distinguished from the 4096 possible codes at each position; the six code dimensions are not six independent temporal segments. FSQ provides the discrete generation interface, while multigranular decoding supervision gives the positions their ordered semantics.

The model uses 8 granularity levels with prefix budgets of 1, 2, 4, 8, 16, 32, 64, and 128 tokens. Small budgets describe trends; larger budgets progressively carry detail. Training does not discard arbitrary tokens while still requiring reconstruction of the original fine signal. Instead, each budget is paired with a target smoothed to the corresponding degree. Otherwise, a small prefix would also be forced to retain all detail, undermining the granularity interpretation of its length.

The core Gaussian-smoothing target is defined below. The finest level, level 8, uses the original sequence directly rather than treating a zero bandwidth as an ordinary Gaussian kernel.

\[ \mathbf{x}^{(\ell)}=g_{\sigma(\ell)}*\mathbf{x},\qquad \sigma(\ell)=\frac{T}{12}\frac{8-\ell}{7},\qquad \mathbf{x}^{(8)}=\mathbf{x}. \]

The kernel is normalized and truncated at three standard deviations. Smoothing removes high-frequency detail without changing sequence length, so the control concerns information granularity under this training definition rather than directly changing the number of sampled points. The coarse-to-fine transition is not deterministically invertible: multiple fine signals can correspond to the same smoothed outline.

The decoder uses conditional flow matching (CFM). Given the current noisy state, flow time, and selected token prefix, it learns to transport Gaussian noise toward the distribution of sequences at the corresponding smoothing level. Shorter prefixes receive smoother targets, and longer prefixes receive finer targets. The discrete representation and continuous generator therefore jointly establish the prefix-length–granularity relationship instead of applying a common filter after generation.

2. Granularity-Block Autoregression: predict newly added information at the next level, not the next time point

After tokenizer training, the representation is fixed and a VAR Transformer models it. The level-0 budget is 0, and subsequent budgets follow the exponential schedule above. At each step, the model conditions on the existing coarse prefix and predicts the new token block needed to reach the next budget. The first two blocks each contain 1 token, followed by blocks of 2, 4, 8, 16, 32, and 64 tokens, totaling 128 positions.

This is neither conventional next-token prediction nor prediction of the next observation on the time axis. It is next-block prediction along the granularity axis: the new tokens at one level are predicted as a group and become the context for the next level. Generation can stop at an intermediate granularity without first producing all fine detail. Traversing the entire hierarchy takes 8 granularity stages, but continuous decoding still incurs flow-sampling costs; this alone does not establish any particular end-to-end speedup.

The raw-space comparison concatenates continuous sequences from all 8 levels into a trajectory of approximately \(8T\) positions and models it with an autoregressive Transformer and mixture-density head. In conditional experiments, it additionally receives the ground-truth trajectory of every preceding coarse level. TimeTok's advantage is therefore not simply the presence of a coarse-to-fine objective: it uses a compact, semantically ordered prefix to represent preceding granularity information. The experiments support this comparison, but do not isolate all performance differences to one component.

3. Prefix Selection and Target Control: estimate the input level and add only the missing detail

For a coarse input of unknown granularity, the model first encodes a full candidate token sequence. It reconstructs the input using all 8 prefix budgets and computes the L2 distance at each level. Rather than selecting the most complex level with the smallest error, it chooses a level based on relative reconstruction-error improvement between adjacent levels and retains that prefix. This is a reconstruction-curve heuristic, not a measurement of the true sampling rate, and it is not proven to identify a “true granularity” in every case. The paper does not give sufficiently precise closed-form selection details, so no exact algorithmic formula is supplied here.

Given a selected source level \(i\) and finer target \(j\), VAR keeps the original prefix and generates the blocks from \(i+1\) through \(j\). The corresponding token budget and random noise are then decoded into a continuous sequence. Fixing the discrete prefix constrains coarse structure, but does not guarantee that re-smoothing the output reproduces the input pointwise. The objective does not impose that hard constraint, so consistency must be checked empirically with C-Cons. With no input, generation starts from [BOS] and stops at any requested level; reaching level 8 gives standard generation.

Class conditioning can be incorporated through AdaLN. A further extension trains a foundational tokenizer on UTSD: multigranular representations are constructed, and detrended fluctuation analysis (DFA) exponents are partitioned into 8 uniform bins to assign levels. DFA-based assignment in heterogeneous pretraining is distinct from reconstruction-error detection in single-dataset inference. After tokenizer transfer, VAR is still trained on downstream data; this is not zero-shot transfer of the complete generator.

A Worked Example

Consider an ordinary, fixed-length coarse electricity-demand waveform that rises in the morning and falls in the afternoon. The following walkthrough illustrates the mechanism; it is not an additional measured example from the paper.

The input is encoded into 128 candidate tokens and compared through reconstructions at different budgets. Suppose the reconstruction heuristic selects level 3. The first 4 tokens are retained: they encode coarse structure across the full waveform, not its first 4 measurements.

If the target is level 6, VAR generates 4 additional tokens for level 4, 8 for level 5, and 16 for level 6, producing a 32-token prefix. Conditioned on this prefix and Gaussian noise, the flow decoder outputs an equally long waveform with more local variation.

A different random seed can produce another refined sample. Its detail should be interpreted as a draw from the learned conditional distribution, not as the true missing measurements. A level-8 target requires two further blocks of 32 and 64 tokens. Starting directly from [BOS] and stopping at level 6 instead produces an unconditional intermediate-granularity sample, not a refinement of this waveform.

Loss & Training

The first stage jointly trains the encoder and conditional flow decoder while sampling granularity levels uniformly. With Gaussian noise \(\boldsymbol{\epsilon}\) and flow time \(\tau\), the linear interpolation path and target velocity define the multigranular flow-matching loss:

\[ \mathbf{x}^{(\ell)}_\tau=(1-\tau)\boldsymbol{\epsilon}+\tau\mathbf{x}^{(\ell)},\qquad \mathcal{L}_{\mathrm{CTF}}= \mathbb{E}_{\mathbf{x},\boldsymbol{\epsilon},\tau,\ell} \left\|\mathbf{v}_\theta(\mathbf{x}^{(\ell)}_\tau\mid\mathbf{Q}_{1:n_\ell},\tau) -(\mathbf{x}^{(\ell)}-\boldsymbol{\epsilon})\right\|_2^2. \]

The second stage freezes the encoder and decoder and optimizes blockwise negative log-likelihood over the training sequences' discrete tokens, with \(n_0=0\):

\[ \mathcal{L}_{\mathrm{VAR}}=-\mathbb{E}_{\mathbf{Q}} \sum_{\ell=1}^{8}\log p_\phi \left(\mathbf{Q}_{n_{\ell-1}+1:n_\ell}\mid\mathbf{Q}_{1:n_{\ell-1}}\right). \]

The UTSD tokenizer is trained on 2,954,239 samples of length 256. Single-dataset experiments are univariate, with only the first ETTh1 channel used. Classification evaluation uses InceptionTime and forecasting uses DLinear, both under Train-Synthetic-Test-Real (TSTR). UCR data are actually split 60%/20%/20%, rather than simply using the original train/test counts listed in Table 7. The paper reports standard deviations over three random seeds, but this cache does not fully specify reproducibility hyperparameters such as the optimizer and learning rate; these are not invented here.

Key Experimental Results

Main Results

Standard generation targets the finest level. Lower FID and forecasting MSE are better; higher ACC/F1 are better. The table selects representative generators from Table 2. Oracle denotes a downstream model trained on real data, not another generator.

Dataset and metric TimeTok ImagenTime DiffusionTS Oracle
ECG5000 FID 0.004 ± 0.001 0.020 ± 0.010 0.041 ± 0.009 —
ECG5000 ACC 0.948 ± 0.006 0.940 ± 0.007 0.896 ± 0.015 0.959 ± 0.001
ECG5000 F1 0.564 ± 0.047 0.477 ± 0.104 0.552 ± 0.028 0.620 ± 0.010
ItalyPowerDemand FID 0.007 ± 0.001 0.043 ± 0.006 0.023 ± 0.005 —
ItalyPowerDemand ACC 0.982 ± 0.000 0.978 ± 0.002 0.979 ± 0.003 0.982 ± 0.008
Nasdaq FID 0.091 ± 0.017 0.142 ± 0.032 0.208 ± 0.015 —
Nasdaq MSE 0.045 ± 0.001 0.046 ± 0.000 0.056 ± 0.001 0.033 ± 0.000
ETTh1 FID 0.070 ± 0.006 0.074 ± 0.007 0.083 ± 0.011 —
ETTh1 MSE 0.103 ± 0.000 0.099 ± 0.000 0.101 ± 0.001 0.098 ± 0.000

TimeTok has the best ETTh1 FID in Table 2, but worse MSE than ImagenTime, DiffusionTS, and Oracle. Its Nasdaq MSE beats the generation baselines but remains worse than Oracle's 0.033. Better FID does not imply improvement on every downstream task.

Conditional generation samples \(K=5\) sequences per coarse input. CRPS assesses pointwise predictive-distribution accuracy and calibration and is summed over time. It contains both error relative to the true target and a sample-to-sample distance term; a lower score alone does not establish greater sample diversity.

\[ \mathrm{CRPS}_{\mathrm{sum}}=\sum_{t=1}^{T} \left[\frac{1}{K}\sum_{k=1}^{K}|\hat{x}^{(j,k)}_t-x^{(j)}_t| -\frac{1}{2K^2}\sum_{k=1}^{K}\sum_{k'=1}^{K}|\hat{x}^{(j,k)}_t-\hat{x}^{(j,k')}_t|\right]. \]

C-Cons is the RMSE between a re-coarsened generated sequence and the coarse input. Under Gaussian levels, re-coarsening must use the variance difference; applying the full source-level filter defined for the original finest sequence to an intermediate output would be incorrect:

\[ \mathrm{C\text{-}Cons}=\mathrm{RMSE} \left(\mathcal{R}_{j\rightarrow i}(\hat{\mathbf{x}}^{(j)}),\mathbf{x}^{(i)}\right),\qquad \sigma_{j\rightarrow i}^{2}=\sigma_i^{2}-\sigma_j^{2}. \]

The following values average over all valid coarse–fine level pairs, and lower is better throughout. The first three rows come from Table 3 and the raw-space AR row from Table 15.

Conditional generator ECG5000 CRPSsum ECG5000 C-Cons Nasdaq CRPSsum Nasdaq C-Cons
TimeTok 1.587 ± 0.867 0.010 ± 0.006 15.87 ± 12.48 0.096 ± 0.090
TimeGAN 1.737 ± 0.737 0.011 ± 0.005 39.11 ± 22.14 0.196 ± 0.105
TOTEM 2.639 ± 1.286 0.035 ± 0.019 35.97 ± 22.54 0.294 ± 0.247
Raw-space AR Transformer 5.804 ± 1.658 0.040 ± 0.018 22.98 ± 18.95 0.119 ± 0.107

ImagenTime and TSGDiff are not included in conditional GC-TSG comparisons because the authors regard adapting their frequency-domain or graph representations as beyond a reasonable modification. Other baselines are adapted for granularity conditioning. The conclusions therefore apply to these implementations, not to an impossibility of granularity control in other generative paradigms.

Ablation Study

This table combines level-detection analysis from Table 8 and the foundational-tokenizer extension from Table 6. These are different experimental settings, not a single controlled one-factor ablation across every row.

Config ECG5000 results Nasdaq results Question addressed
TimeTok with reconstruction-curve level detection CRPSsum 1.587 ± 0.867 — Source-level estimation at inference
TimeTok with ground-truth constructed level CRPSsum 1.509 ± 0.883 — Known-level reference; gap 0.078
UTSD TimeTok without DFA FID 0.042; ACC 0.944; F1 0.583 FID 0.111; MSE 0.044 Levels defined using the single-dataset approach
UTSD TimeTok with DFA FID 0.010; ACC 0.951; F1 0.668 FID 0.046; MSE 0.042 DFA-based assignment for heterogeneous data

The exponential-versus-uniform budget comparison fixes the total budget at 128. Uniform allocation uses cumulative budgets of 16, 32, …, 128 tokens. Figures 6–7 show similar training and validation reconstruction curves, supporting smaller budgets for coarse levels. Exact final values are not tabulated, so none are guessed from the figures.

Key Findings

  • Table 4 reports cross-level average unconditional FID of 0.002 ± 0.001 on ECG5000 and 0.103 ± 0.014 on Nasdaq. These are not the single finest-level standard-generation scores in Table 2.
  • Under alternative coarsening operators, TimeTok's ECG5000 CRPSsum is 2.354 for wavelets, 2.080 for moving averages, 2.126 for downsampling, and 1.587 for Gaussian smoothing (Table 17). This supports some cross-operator generalization, but C-Cons is not reported for non-Gaussian cases; it does not establish exact preservation of every coarse structure.
  • Multiple evaluators in Tables 18–21 reinforce task dependence: among synthetic-data methods, TimeTok has mean ECG5000 F1 rank 1.4, but mean ETTh1 MSE rank 3.2, versus ImagenTime's 1.2.
  • The source contains unexplained numerical differences: finest-level ECG5000 FID is 0.004 ± 0.001 in Table 2 and 0.006 ± 0.001 in Table 10. The all-pairs ItalyPowerDemand/ETTh1 averages in Table 13 do not directly correspond to the individual-pair rows in Table 12. This note retains each table's boundaries rather than silently reconciling the numbers.

Highlights & Insights

  • Granularity is a representation interface, not merely an output-filter parameter: short prefixes describe full-sequence trends, while added tokens introduce detail. Intermediate granularities consequently become independently generatable and evaluable targets rather than scaffolding for a final sample.
  • Stochastic continuous decoding complements discrete hierarchy: VAR learns which codes to add between levels, while CFM learns how conditioned codes become continuous waveforms. Conditional sampling is useful, but the respective contributions of token and flow randomness merit separate investigation.
  • Transfer requires checking whether granularity labels are comparable: the UTSD extension uses DFA to recalibrate levels instead of treating every observation as finest-grained. The transferable insight is calibration of cross-domain representation supervision, not treating the DFA exponent as a universally valid physical ground truth for granularity.

Limitations & Future Work

  • Training granularity depends on a constructed definition: the main experiments largely use Gaussian smoothing to form 8 levels of fixed-length univariate sequences. Aliasing, irregular sampling, missing observations, and multivariate coupling in real devices are not automatically covered.
  • Consistency is empirical rather than a hard constraint: a fixed coarse prefix does not ensure pointwise equality after re-coarsening. A useful extension would add verifiable consistency constraints and study their interaction with conditional-distribution coverage.
  • Unknown-level detection remains heuristic: the ground-truth-level comparison reports an average gap on one dataset, not a reliability guarantee for unfamiliar domains, noisy inputs, or hand-drawn outlines.
  • Foundational-tokenizer evidence is preliminary: multidomain pretraining and transfer to two downstream generation datasets do not establish general time-series foundation-model capability. Downstream VAR retraining costs also belong in the assessment.
  • Evaluation and reproducibility have boundaries: CRPSsum depends on sequence length and scale, so its magnitude should not be compared directly across datasets. Fuller hyperparameters, runtime costs, and explanations of numerical discrepancies are needed. Generated detail is not true recovered information, and these benchmarks do not justify medical diagnosis or financial decision advice.
  • vs TimeVQVAE / TOTEM: all use discrete time-series representations, but TimeTok binds prefix budgets to smoothed targets so that token order explicitly corresponds to granularity. Control comes from this binding, not simply from replacing the quantizer.
  • vs DiffusionTS / TimeGAN: these methods principally optimize standard generation, whereas TimeTok adds a unified coarse–fine interface and compares with adapted implementations. Its gains concern conditional refinement and distribution matching, not universally superior prediction.
  • vs VAR / FlowAR / FlexTok: TimeTok draws on next-scale block generation and variable-length ordered prefixes, then makes the hierarchy a user-facing temporal input–output control variable. Its novelty lies primarily in aligning these elements with the task, rather than introducing every underlying component.
  • vs temporal super-resolution: a fixed low-to-high mapping is one conditional-refinement instance; GC-TSG additionally supports intermediate targets and unconditional generation. Testing the unified interface under genuine downsampling rather than predominantly constructed smoothing remains important.

Rating

  • Novelty: 4/5 — Integrates ordered prefixes, block generation, and temporal-granularity supervision into a unified interface, with precedents for the underlying components.
  • Experimental Thoroughness: 4/5 — Covers four datasets, conditional pairs, multiple evaluators, and transfer, but lacks broader real granularity-mismatch and multivariate validation.
  • Writing Quality: 3/5 — The methodological thread is clear, while some table references, numerical relationships, and reproducibility details need clarification.
  • Value: 4/5 — Provides a reusable granularity-control representation design; the authenticity of generated detail and deployment reliability require independent validation.