Kairos: Toward Adaptive and Parameter-Efficient Time Series Foundation Models¶
Conference: NeurIPS2026 (task-list placement; this note follows arXiv v4)
arXiv: 2509.25826
Paper: Project page
Area: Time Series
Keywords: adaptive patching, mixture-of-size encoding, dynamic rotary positional encoding, multi-patch prediction, zero-shot forecasting
TL;DR¶
Kairos handles local time-series complexity and instance-level spectral differences through mixture-of-size encoding and dynamic rotary positional encoding, then predicts multiple future patches in parallel, reaching a normalized MASE of 0.738 on GIFT-Eval with 53M parameters, without uniformly leading in probabilistic or task-level performance.
Background & Motivation¶
Time series foundation models aim to pretrain on extensive cross-domain data without retraining for every new dataset. Yet hourly electricity consumption, daily sales, and signals with local abrupt changes do not share a single temporal scale. Point-wise representations in Chronos, fixed-length patching in models such as TimesFM, and multi-scale selection applied uniformly across a whole sequence can assign the same representation budget to stable and rapidly changing regions. Large patches can lose detail, whereas small patches create redundant tokens in low-complexity regions; enlarging the Transformer only compensates indirectly for this mismatch.
Positional encoding has a related problem. Fixed RoPE assigns every sequence the same rotation frequencies and treats token indices as temporal distances. Once an encoder permits different local patch lengths, adjacent tokens no longer represent equal observation spans; even with equal patch lengths, different sequences have different periodic structures. Dynamic patching without a corresponding positional change can therefore introduce a new distortion of the time axis. Kairos does not simply increase sparse expert capacity: lightweight input-side modules absorb this heterogeneity before it burdens the backbone.
Here, “information density” is an analytical concept describing spectral complexity, not a supervision label directly computed by the router. Appendix M applies a Hann window to windows with both length and stride set to 128, then computes Shannon entropy from the normalized power spectrum; concentrated spectra indicate lower complexity, while dispersed spectra indicate higher complexity. Routing is still learned end-to-end through the forecasting objective, and high entropy should not be equated with greater predictability. Core idea: select the granularities needed by each segment and calibrate attention's temporal coordinates using instance spectra and actual observation spans, reserving limited parameters for transferable temporal structure rather than compensation for a uniform representation mismatch.
Method¶
Overall Architecture¶
The input is a univariate observation history, with multivariate data processed independently by channel; outputs are future quantiles, with the median used for point forecasting. “Mixture-of-Size Encoding” first selects patching experts independently for each coarse segment and fuses their outputs into a variable-length token sequence. “Dynamic Temporal Calibration” then uses DRoPE inside the Transformer to adjust rotation frequencies and positions. Finally, “Multi-Patch Parallel Decoding” retrieves historical representations through cross-attention using learnable forecast tokens.
Parallelism occurs within a decoding round, not in a single forward pass for every possible horizon. The standard configuration outputs 2 future patches of length 64 per round; beyond that coverage, median predictions are appended to the history for another round. Future ground truth provides quantile-loss supervision during training and is not an input to zero-shot inference.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Historical observations"] --> B["Mixture-of-Size Encoding"]
B --> C["Dynamic Temporal Calibration"]
C --> D["Multi-Patch Parallel Decoding"]
D --> E["Future quantiles"]
E -->|Long horizon: append medians to history| A
Y["Training future ground truth"] -.->|Training supervision only| L["Weighted quantile loss"]
E -.->|Training only| L
Key Designs¶
1. Mixture-of-Size Encoding: let local complexity determine representation resolution
The history is first divided into non-overlapping coarse segments of fixed length. What changes dynamically is the patching granularity within each segment, not arbitrary segmentation boundaries. A lightweight linear router produces expert scores from the segment observations, applies softmax with load-balancing biases, and selects Top-K candidates from real granularity experts and null experts together. Each real expert corresponds to a fixed patch length and uses a two-layer MLP for encoding. Null experts produce no representation and perform no encoding computation; they occupy selection slots, allowing the number of active real granularities to be smaller than K.
Base has 4 real experts with patch lengths 32, 64, 128, and 256, plus 2 null experts, with Top-K set to 4. Since there are fewer null experts than selection slots, real experts always remain active, preventing a segment from having no representation. Experts differ in observation granularity rather than being equally sized Transformer FFNs used to expand capacity. Local variation can use fine granularity, while stable regions can use coarser representations.
Different granularities produce different token counts, so their embeddings cannot be added directly or simply concatenated indiscriminately. Kairos takes the finest active patch length in the current segment as its target resolution. Because patch lengths are nested integer multiples, coarse embeddings can be repeated into the corresponding fine positions before fusion. Repeating a coarse embedding shares context; it does not reconstruct details that were never encoded. Fine detail still comes from the fine-grained expert. The final token count is determined by the finest active granularity, not by the sum of all expert token counts.
The valid set contains only selected real experts, and fusion weights are renormalized over that set. Null experts are not averaged in as zero vectors. Fused coarse segments are concatenated in temporal order, allowing neighboring segments to contribute different numbers of tokens while converting the multi-granularity outputs into a unified sequence for the backbone.
Training balances loads without adding an auxiliary loss: routing biases are adjusted using each expert's total softmax weight across a batch and a target load distribution. The target is not necessarily uniform and includes null-expert shares. This helps prevent undertrained experts, but also means computation allocation is influenced by manually specified target proportions rather than unconstrained discovery of an optimal granularity.
2. Dynamic Temporal Calibration: correct spectral differences and the time axis of variable-length tokens
DRoPE retains RoPE's paired-dimension rotation of queries and keys, changing the rotation frequencies and positional coordinates. On the frequency side, the observation mask is applied to the history before a real FFT. The first 128 low-frequency magnitudes pass through LayerNorm and a lightweight MLP to generate instance-specific, layer-specific scales and offsets. Modulation occurs in log-frequency space to avoid disproportionate adjustments across the broad range of base frequencies, followed by exponentiation back to positive frequencies.
On the position side, equally spaced token indices are no longer appropriate. Each fused token represents the finest active patch length in its segment. Accumulating these patch lengths and dividing by the globally smallest candidate patch length yields calibrated coordinates. A token covering a larger span advances temporal distance further instead of taking the same unit step as a fine patch. Here, “physical time” means the time axis measured in original observation spans; it does not imply automatic access to external units such as hours or days, or support for arbitrary irregular timestamps.
The scales and offsets come from the current historical spectrum, while patch lengths come from the actual encoded token resolution; layer indices are omitted for brevity. The two parts address different problems: frequency modulation adapts to periodic and trend structure, while positional calibration prevents mixed patch lengths from distorting relative distances. The derivation in Appendix H establishes periodicity of each two-dimensional subspace's RoPE attention contribution for fixed content, not that dynamic modulation must improve overall generalization. Its benefits still require ablations and intervention experiments.
DRoPE does not estimate a single “true period” and directly substitute it into positional encoding. It generates a frequency profile that varies by instance and layer to alter attention's preference over relative lags. Multiple rotation frequencies can jointly represent multi-periodic, non-stationary signals. The spectrum conditions the model at the whole-history level, complementing the local adaptation of segment routing.
3. Multi-Patch Parallel Decoding: reduce long-horizon rounds rather than eliminate autoregression
The decoder assigns independent learnable forecast tokens to future patches. Through cross-attention, these tokens retrieve Transformer history representations, and a shared residual feed-forward prediction head produces fixed-length patches. Multiple future patches are generated in parallel within a round, so a later patch does not depend on values already generated for the preceding patch. Compared with single-patch rolling forecasts, this reduces how often errors are written back into the history.
The default is 2 forecast tokens, each covering 64 future points, for 128 points per round. Shorter targets can use the required predicted patches and intervals; longer targets still require rolling rounds. Appendix D compares a single token producing 128 points against 2 tokens producing 64 points each. The latter performs better despite equal direct forecast lengths, indicating that the benefit is not solely a longer forecast per round: separate queries can retrieve information needed for different future patches.
More forecast tokens are not always better. Training samples must supply both sufficient history and supervision for every future patch, and longer direct forecasts are harder. Appendix D counts samples whose total length is at least twice the per-round forecast length: coverage is 82.69% for 2 tokens and 71.53% for 12. A small number of parallel patches therefore represents a concrete trade-off among forecasting difficulty, training coverage, and rolling rounds.
A Worked Example¶
Use the encoding example from Appendix H: a coarse segment has length 128, candidate patch lengths are 32, 64, and 128, there are 2 null experts, and Top-K is 3. Suppose the router selects patch length 32, patch length 128, and a null expert. The real experts produce 4 fine embeddings and 1 coarse embedding, respectively.
Before fusion, the coarse embedding is repeated 4 times to align with the fine embeddings. The aligned outputs are summed using weights renormalized over the real experts, producing 4 tokens rather than 5. If another segment retains only a coarser valid granularity, it contributes fewer tokens, and DRoPE measures temporal distance using the associated observation spans rather than the indices after concatenation.
Under standard decoding, a horizon of 192 requires 2 rounds: the first produces 128 points, median predictions are appended to the history, and the next supplies predictions for the remaining 64 points. This inference example explains round counts and is not an additional experiment from the paper; the subsequent round cannot access future ground truth.
Loss & Training¶
The model directly outputs 9 quantiles with levels from 0.1 to 0.9. The median 0.5 is used for point forecasting and multi-round feedback. Training uses weighted pinball loss, placing more weight on times near the forecast origin. This is not simply mean-squared-error training, and probabilistic forecasts do not require stochastic generation sampling.
The overall objective averages over batches and quantiles and sums weighted losses over future timesteps. The paper maps discrete timesteps over the range from \(1+10^{-5}\) to \(H-10^{-3}\) into an evenly spaced continuous sequence \(t'\) to avoid a zero final weight; the original \(t'\) should not simply be replaced by \(t\).
PreSTS contains over 300B real time points plus 15B synthetic time points. The abstract and main text summarize its scale as “over 300B,” while Appendix I gives the finer composition. Real datasets are grouped into 5 predictability tiers, preferentially sampling clearer patterns. The loader samples real and synthetic data at 80% and 20%, respectively, which is not their storage-size ratio. Synthetic data include combinations of seasonality, trend, and noise alongside idealized industrial periodic signals, supplementing regularity and period distributions missing from real corpora.
The authors state that the training, validation, and test splits of GIFT-Eval and TSLib are excluded, supporting the reported zero-shot setting. This is the paper's exclusion rule; this note does not independently audit the corpus. Base uses 6 layers, 8 attention heads, and a hidden dimension of 512, trained for 300,000 steps with batch size 512, AdamW, and linear learning-rate decay. DRoPE-related parameters use learning rate \(10^{-5}\) and the remaining parameters use \(10^{-3}\). Reported training uses 4 A100 GPUs, TF32, and approximately 15 hours; this is not an independently reproduced runtime.
Key Experimental Results¶
Main Results¶
GIFT-Eval contains 97 tasks: 55 short-term, 21 medium-term, and 21 long-term, originating from 28 source datasets. Each task's MASE and CRPS are normalized by Seasonal Naïve, then aggregated geometrically; lower is better. The table selects several foundation models from the paper's Table 1. The central advantage concerns point forecasting and parameter count, not leadership on every probabilistic metric.
| Model | Params | Normalized MASE | Normalized CRPS |
|---|---|---|---|
| Chronos | 709M | 0.870 | 0.574 |
| ChronosBolt | 205M | 0.808 | 0.574 |
| TimesFM | 500M | 0.758 | 0.550 |
| Toto | 151M | 0.750 | 0.517 |
| Sundial | 128M | 0.750 | 0.559 |
| Kairos-Small | 23M | 0.748 | 0.554 |
| Kairos-Base | 53M | 0.738 | 0.548 |
Base reduces overall normalized MASE by 1.6% relative to Sundial. However, the 8.1% geometric-average improvement applies only to the 64 tasks it wins, while the 33 lost tasks show a 12.3% geometric-average worsening. A win-only statistic must not be presented as the overall gain. Grouping by source dataset gives 20 wins and 8 losses, with a 9.36% improvement among wins and an 8.02% worsening among losses.
TSLib uses ETTh1, ETTh2, ETTm1, ETTm2, Weather, and daily and weekly Saugeen, comprising 27 dataset–horizon pairs; weekly Saugeen excludes horizon 720. Kairos uses context length 2048, while some foundation models use 512 or 1536. Supervised baselines select the better original-paper or long-context candidate. Table 23 reports overall MSE / MAE of 0.471 / 0.365 for Base, 0.477 / 0.367 for Mini, and 0.456 / 0.397 for full-shot PatchTST. Kairos therefore should not be described as beating full-shot models on every overall metric.
Ablation Study¶
The following normalized MASE values are selected from Table 2. Each row replaces or removes an architectural component rather than changing the pretraining corpus. The full model is best overall, but single-patch autoregression is better for short horizons, demonstrating a horizon-dependent benefit.
| Config | Short | Medium | Long | Overall |
|---|---|---|---|---|
| Fixed patch length (32) | 0.724 | 0.802 | 0.820 | 0.761 |
| Mixture-of-size without null experts | 0.720 | 0.770 | 0.800 | 0.748 |
| Standard RoPE | 0.729 | 0.807 | 0.835 | 0.767 |
| Instance frequency modulation only | 0.719 | 0.767 | 0.797 | 0.746 |
| Granularity position calibration only | 0.727 | 0.796 | 0.807 | 0.758 |
| Single-patch autoregression | 0.705 | 0.794 | 0.848 | 0.753 |
| Full model | 0.709 | 0.761 | 0.794 | 0.738 |
Table 13 crosses models and training corpora, showing that Kairos's advantage does not exist only with PreSTS. However, parameter counts, objectives, and training implementations can still differ; matching the corpus is not causal proof isolating architecture or training objective.
| Model | Params | Training corpus | Normalized MASE |
|---|---|---|---|
| Kairos | 53M | PreSTS | 0.738 |
| Kairos | 53M | Chronos corpus | 0.761 |
| ChronosBolt | 205M | PreSTS | 0.781 |
| ChronosBolt | 205M | Chronos corpus | 0.808 |
Key Findings¶
- Local adaptation has additional support: sequence-level Pathformer-style routing scores 0.759, uniform granularity weights at inference score 0.831, and shuffled routing scores 1.205, versus 0.738 for the full model. These results support useful learned selections, but interventions also alter the representation distribution the model expects and do not establish a uniquely optimal routing policy.
- DRoPE instance conditions cannot be exchanged arbitrarily: within-dataset shuffling scores 0.751, cross-dataset shuffling 0.947, and standard RoPE 0.767. This ordering against the full model's 0.738 supports instance matching. However, fixed RoPE is a retrained structural baseline, whereas shuffling intervenes at inference; they are different kinds of manipulation.
- Statistical tests use 28 source datasets rather than treating correlated task variants as independent observations. After Holm correction, the comparison against Sundial has a non-significant sign-test value of 0.0714 and a significant Wilcoxon value of 0.0402; against fixed patching, the values are 0.0872 and 0.0358. Comparisons against removing null experts and standard RoPE are significant under both tests, so “all comparisons are significant” is inaccurate.
- Parameter matching addresses capacity more directly: with the same PreSTS, training steps, and optimization settings, a conventional 53M Transformer scores 0.797 versus Kairos's 0.738. Mini trained from scratch on Weather averages MSE / MAE of 0.228 / 0.267, showing that benefits do not depend entirely on pretraining, but this finding remains limited to that supervised dataset and the listed baselines.
- Latency is averaged per batch on one TITAN RTX with input length 2048 and output length 96: Base 0.061 s, Mini 0.030 s, TTM 0.009 s, ChronosBolt 0.055 s, and Toto 13.717 s. TTM instead uses input length 1536. These are not full-benchmark runtimes or hardware-independent throughput claims, and Kairos is not the fastest model. The accompanying MASE is the overall GIFT-Eval metric, not the timed sample's error.
Highlights & Insights¶
- Selecting representation budgets before the backbone tackles heterogeneity more directly than increasing parameter count alone. Null experts allow the router to choose one fewer granularity instead of invoking every candidate needed to fill a fixed Top-K budget.
- Multi-scale fusion requires more than offering several patch lengths: repetition-based alignment and renormalization over valid experts are essential. Simply concatenating all scales would not reproduce the same token time axis or permit the same positional calibration.
- Positional encoding must change alongside the encoder's temporal granularity. A transferable lesson for nonuniform sampling or adaptive compression is to check whether representation indices and actual spans have diverged instead of assuming adjacent representations are equally spaced.
- Multi-patch decoding reduces feedback between rounds rather than guaranteeing no error accumulation. In the ETTh1 analysis of Appendix D, rounds 2–5 still incur a 9.79%–14.85% MAE penalty relative to a ground-truth-feedback oracle, supporting mitigation rather than elimination.
Limitations & Future Work¶
- Channel-independent processing does not explicitly model inter-variable dependencies. A channel mixer is an author-proposed extension, but its effectiveness and cost are not evaluated here.
- Probabilistic forecasting still trails Toto: CRPS is 0.548 versus 0.517. YingLong also scores 0.548 in Table 1, so the main text's “second-best” claim should be read as a tie at displayed precision, not an exclusive ranking.
- Generalization evidence primarily concerns forecasting. Appendix C's frozen encoder achieves mean accuracy 0.818 across 91 UCR datasets versus MOMENT's 0.794, but it still trains a lightweight classifier. This is not zero-shot classification and does not replace anomaly-detection or imputation evaluation.
- Repeated training reports only 3 Small runs, with normalized MASE \(0.747\pm0.001\); it does not establish seed uncertainty for every Base ablation. Deterministic quantile inference also does not make training non-random.
- Data and explanations have scope details: the main text summarizes over 300B, while the appendix separates over 300B real points plus 15B synthetic points. Appendix I's interpretation of near-zero ARCH-LM statistics is ambiguous relative to the usual statistical reading, so it is not used to strengthen claims about synthetic-data distributions. Publishing tiering rules, sampling details, and statistical implementations would improve auditability.
Related Work & Insights¶
- vs Pathformer: it selects multiple scales using sequence-level features, whereas Kairos selects granularities per coarse segment. The replacement study in Appendix E directly tests the extra value of local adaptation, not whether all multi-scale methods are unsuitable for time series.
- vs Chronos / ChronosBolt: Kairos uses dynamic granularity and DRoPE rather than leaving heterogeneity entirely to the backbone. Matched-corpus results support the combination, but do not independently prove that a continuous quantile objective is superior to alternative objectives.
- vs ElasTST: tunable RoPE adapts through dataset-specific training, whereas Kairos generates layer-specific modulation dynamically from each instance's historical spectrum. Updating spectral conditions after abrupt changes could be a research direction for non-stationary forecasting, not a result already validated here.
- vs single-patch autoregressive foundation models: multiple queries distribute retrieval across future patches and reduce iterations while retaining variable-horizon forecasting. Comparisons should jointly control direct forecast length per round, training coverage, and context budgets; Appendix D's fixed-length control is a useful starting point.
Rating¶
- Novelty: 4/5; coordinated local resolution and instance-spectral positional encoding differ from generic sparse capacity expansion.
- Experimental Thoroughness: 4/5; structural ablations, routing interventions, matched corpora, and source-dataset statistics are provided, but seeds and cross-task coverage remain limited.
- Writing Quality: 4/5; appendices clarify fusion and long-horizon mechanisms, while rankings and corpus-size scopes require careful reading.
- Value: 4/5; reusable architectural ideas for compact zero-shot forecasters, not a parameter-efficient fine-tuning method.