Progressive Memory Transformer: Memory-Aware Attention for Time-Series¶
Conference: NeurIPS2026
arXiv: 2609.31351
Code: https://github.com/dsb-ifi/pmt
Area: Time Series
Keywords: writable memory, multi-scale representations, contrastive learning, low-label classification, time-series forecasting
TL;DR¶
PMT exposes writable memory updated along sliding windows as an explicit mid-range representation and separately supervises tokens, memory, and sequence summaries, achieving 84.4% average accuracy across seven low-label classification datasets while a cue-retention probe verifies that memory carries information across windows.
Background & Motivation¶
Useful information in time series exists beyond two endpoints: fine-grained changes support forecasting, whole-sequence summaries support classification, and intermediate recurring patterns such as gait or vibration motifs also require stable representations. TS2Vec builds multi-resolution supervision through repeated max-pooling, TS-TCC combines cross-view future prediction with contextual discrimination, and SoftCLT assigns soft positives using temporal distance and instance similarity. These approaches exploit temporal structure, but intermediate states often remain compressed or indirectly shaped internal variables rather than an interface that can be read and supervised directly at the mid-range scale.
Adding memory alone does not automatically address this problem. Transformer-XL treats historical activations as a read-only cache, while recurrent networks propagate context through hidden states; both can retain history without organizing it into motif representations aligned with the current temporal window. Global latent tokens or memory banks compress context, but make it difficult to specify what a particular slot in a particular window should preserve across augmentations. The paper therefore changes the representation interfaces exposed by the backbone to training objectives and downstream tasks, rather than simply enlarging the global classification vector.
Architecture and objectives must consequently be designed together: each window first produces sample-specific writable slots, which receive supervision matching their granularity, while local tokens and the global summary remain available. Core Idea: make window-aligned progressive memory an independent mid-range representation, then shape the three levels through local continuity, mid-range cross-view consistency, and global instance agreement.
Method¶
Overall Architecture¶
A multichannel time series is first converted into patch tokens by a one-dimensional convolution and unfolded into overlapping windows of length \(W\) and stride \(S\). Progressive Memory Attention (PMA) advances along these windows, updating current tokens and a small set of memory slots. Stacking PMA blocks also supplies each window with tokens and memory from the previous block, creating horizontal propagation through time and vertical propagation through depth.
Overlap Aggregation and Hierarchical Readout then merge duplicated positions back into a fixed-length token stream and use neighborhood-masked encoders plus a global [CLS] token to expose local tokens, mid-range memory, and a global summary. Three-Scale Supervision acts on these interfaces only during training; inference requires neither paired augmentations nor contrastive losses, and downstream tasks select the interface matching their granularity.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
X["Time series<br/>convolution and windowing"] --> A["Progressive Memory Attention"]
H["Previous-window memory"] -->|horizontal propagation| A
V["Previous-block tokens and memory"] -->|vertical propagation| A
A --> B["Overlap Aggregation<br/>and Hierarchical Readout"]
A -->|updated memory| H
B --> R["Local tokens / mid-range memory<br/>global summary"]
R --> O["Downstream probes<br/>inference data flow"]
R -.-> C["Three-Scale Supervision<br/>training only: HGCL / PCL / ICL"]
Key Designs¶
1. Progressive Memory Attention: supply each window with historical summaries and previous-block representations
When processing a new window, each PMA block receives memory produced by the previous window in the same block instead of rereading all historical tokens. It also receives the current window's memory and tokens from the previous block, concatenates these context sources with current evidence, and updates them jointly through one multi-head attention operation. Slots participate as queries as well as keys and values, allowing historical reading and evidence writing in one attention computation. They are not shared prompt tokens discarded after processing, but states that evolve window by window for the current sample.
Here \(w\) indexes windows and \(b\) indexes blocks; the horizontal term comes from the preceding window, while the two barred terms come from the preceding block after gating. A memory gate mixes inherited states with a learnable reset state to reduce persistence of stale motifs after a regime change. A token gate mixes processed tokens with original patch representations so that historical context does not erase fine-grained variation. The first window of each block starts from the reset state, rather than from memory accumulated across independent samples.
The mask is asymmetric: memory queries can read the concatenated input; window queries can read previous-window memory and the strictly causal part of their own window, but cannot read the previous block's current-window memory. This blocks summaries potentially containing later-window evidence from leaking backward. These are local update constraints, not a guarantee that every implementation of the entire backbone is strictly causal. Forecasting uses past-only input slices to guarantee feature causality separately.
The main text describes the memory gate as resetting horizontally carried memory, whereas Appendix A.1 and Algorithm 1 explicitly apply gating to the previous block's current-window memory. This implementation-description discrepancy is retained here: horizontal and vertical sources follow Eq. (1), without guessing a unified placement for the reset gate.
2. Overlap Aggregation and Hierarchical Readout: preserve local detail while exposing mid-range states independently
Overlapping windows produce several contextualized outputs for each global token position. Concatenating them would expand sequence length across blocks, while averaging would ignore differences in contextual quality between windows. The paper uses single-query cross-attention at each position to select and fuse its overlapping copies, with a learned attention-head-by-overlap-position bias. Scaling, normalization, and a learnable gate then combine the result with a masked-mean skip path. Each PMA block therefore preserves token length instead of passing all sliding-window copies to the next layer.
After the PMA stack, a short encoder stack further processes tokens: ordinary tokens read only a local historical neighborhood, while [CLS] reads the full sequence to form a global summary. The final block's per-window memory remains separately accessible and is not replaced by this global readout. Forecasting thus uses local representations, cue retention uses final-window memory, and low-label classification uses [CLS]. The interfaces can be tested independently instead of forcing every task into one pooled vector.
This also explains why progressive processing does not require a deep network merely to access the whole history. Horizontal memory can carry preceding-window information in the first block, while deeper blocks recompose summaries and local evidence. Appendix A.3 provides nominal receptive-field upper bounds, not a guarantee that each token actually uses every input within those bounds.
3. Three-Scale Supervision: constrain local continuity, memory motifs, and global instances separately
The mid-range PMA Contrastive Loss (PCL) directly uses normalized memory slots from the final PMA block without an additional projection head. A slot's positive is the corresponding state at the same window and slot index in the second view of the same sample. Negatives come from other sequences in the batch; every slot from the same source sequence is excluded. This targets window-level motifs rather than raw-token caching, although describing each slot as a motif detector remains a mechanistic intuition rather than a guarantee of manually identifiable slot semantics.
Similarity is cosine similarity, and losses are averaged over both view directions. The appendix samples 512 negative slots per anchor and computes anchors in chunks to control activation memory. Directly supervising the slot vector is what turns memory from an internal state into a usable representation, rather than merely extending context propagation.
The local Hierarchical Gaussian Contrastive Loss (HGCL) computes both a token-within-window term and a window-across-sequence term on projected token representations. Both use the aligned second-view position as the positive, give neighboring positions or windows Gaussian weights based on temporal distance, and draw negatives only from other sequences. Window representations are normalized token means. Gaussian widths depend on window length and stride rather than precomputed DTW or data-space distances, making window geometry the source of the local structural prior.
The authors' soft-positive interpretation must be distinguished from the actual equation: the numerator in Appendix A.11 contains only the aligned positive. Weighted similarities of other neighbors are first combined into a bucket, which enters the denominator as an additional exponential term; they are not summed into the numerator as soft positives. HGCL therefore cannot be reduced to conventional multi-positive InfoNCE, nor does the soft-positive terminology alone imply direct attraction of every neighbor.
Global Instance Contrastive Loss (ICL) projects [CLS] through a two-layer GELU MLP, bringing two views of the same sequence together and separating other sequences. Each anchor has \(2(\mathcal B-1)\) other-instance views as negatives. HGCL and ICL use scale-specific projections, whereas PCL does not; this distinction directly subjects the memory itself to cross-view agreement.
A Worked Example¶
Appendix A.3 illustrates a token sequence with \(K=500\), \(W=100\), \(S=25\), and \(B=4\), yielding 17 valid non-padded windows. This is a receptive-field example, not the raw-waveform tokenizer configuration used for FordA experiments.
The first window processes 100 tokens and writes memory; the second receives this summary in addition to overlapping local evidence. By the seventeenth window, nominal historical coverage in the first block can already reach 500 tokens. The untruncated four-block upper bound is 725, but actual coverage remains capped at \(K=500\); additional depth recomposes covered information rather than creating new inputs.
Overlap aggregation fuses different-window outputs for the same position back into 500 positions. The final interfaces are local tokens, memory for 17 windows, and one global summary. During training, the two augmentations establish correspondences at all three levels; at inference, a probe testing whether earlier evidence survives reads only final-window memory rather than searching every window for the cue.
Loss & Training¶
The weak view uses channel-wise z-scaling and low-variance Gaussian noise. The strong view additionally uses time-warping and magnitude-warping, with both views sharing the backbone. HGCL combines its token and window terms internally, and the outer objective combines the three scales:
Training uses AdamW, cosine learning-rate scheduling, 5% warmup, a peak learning rate of \(10^{-4}\), a minimum of \(10^{-6}\), and batch size 256. Main-experiment loss weights are tuned per dataset. Setting every weight to 1 is the Table 5 ablation reference and a suggested starting point, not a shared configuration for Table 1. Appendix E.6 makes this distinction explicit, so headline results should not be described as obtained without per-dataset tuning.
Classification pretrains without labels, freezes the backbone, and trains a linear SVM using 1% or 5% labels. Forecasting freezes the encoder and trains a separate ridge regressor for each horizon; its tokenizer uses patch size and stride 1, and features are generated by sliding-encode over left-padded past windows with \(P=200\), unlike the classification patch configuration.
Key Experimental Results¶
Main Results¶
The following low-label linear-probe results come from Table 1. Each cell is accuracy / macro-F1 (%); the comparison includes two strong time-series SSL baselines without labeling either as SOTA in every setting.
| Dataset / label fraction | PMT | TS2Vec+SoftCLT | TS-TCC+SoftCLT |
|---|---|---|---|
| HAR / 1% | 92.9 / 93.2 | 91.0 / 91.0 | 82.9 / 82.8 |
| HAR / 5% | 95.5 / 95.9 | 92.1 / 92.1 | 92.6 / 92.6 |
| Wafer / 1% | 98.7 / 96.5 | 95.3 / 88.1 | 96.5 / 96.5 |
| Wafer / 5% | 99.0 / 97.2 | 98.8 / 96.8 | 98.2 / 98.2 |
| FordA / 5% | 89.5 / 89.5 | 92.5 / 92.5 | 93.2 / 93.2 |
| FordB / 5% | 76.7 / 76.7 | 78.8 / 78.6 | 88.0 / 88.0 |
| ElectricDevices / 5% | 65.9 / 60.8 | 62.4 / 54.4 | 65.1 / 63.8 |
Across seven datasets and two label fractions, average accuracy is 84.4% for PMT, 83.1% for TS-TCC+SoftCLT, and 82.5% for TS2Vec+SoftCLT. Advantages do not extend to every metric: ElectricDevices at 5% has higher accuracy but lower macro-F1 than TS-TCC+SoftCLT, and FordA/B at 5% show clear deficits.
The main text claims accuracy wins in 11 of 14 settings, but Table 1 gives the same 96.7% for PMT and TS2Vec+SoftCLT on Epilepsy / 5%, alongside three settings below other baselines. Strictly counting wins therefore gives 10 wins, 1 tie, and 3 losses. This note retains the difference between the prose claim and the table's counting convention without changing the original results.
Forecasting Table 3 uses SoftCLT retrained under the same windowed protocol rather than importing results from its original paper. Averaged over five horizons, multivariate Electricity MSE is 0.404 for PMT versus 0.505 for SoftCLT, and multivariate ETTh1 is 0.728 versus 0.742. However, univariate ETTh1 at horizon 24 is 0.103 versus 0.045, showing a substantial short-horizon deficit; its average MAE is also worse, at 0.322 versus 0.309.
Ablation Study¶
Table 2 tests FordA cue retention using a Hann-tapered burst absent from pretraining, with the probe reading only the final-window state. AUC is ROC-AUC; Top-1 and macro-F1 use probability threshold 0.5, and all values lie on a 0–1 scale.
| Config | EOS AUC | EOS Top-1 | EOS macro-F1 |
|---|---|---|---|
| TS2Vec+SoftCLT | 0.649 | 0.638 | 0.604 |
| Transformer-XL | 0.933 | 0.795 | 0.790 |
| xLSTM | 0.956 | 0.820 | 0.817 |
| PMT, PCL disabled | 0.983 | 0.831 | 0.824 |
| PMT, full model | 0.994 | 0.921 | 0.920 |
This probe uses a non-overlapping-window variant. Burst amplitude is 0.5 times the channel standard deviation and width is 20% of sequence length; the standard-deviation calculation excludes the cue span. A frozen backbone demonstrates that cue evidence can be recovered from the final state, not that the model learns arbitrary long-range causal relationships in natural data.
Table 5(a) averages classification over FordA, HAR, POC, and Wafer and both label fractions. The full reference sets all three outer loss weights to 1.
| Config | Top-1 (%) | macro-F1 (%) |
|---|---|---|
| Three losses, all weights 1 | 87.49 | 86.06 |
| ICL disabled | 83.33 | 81.23 |
| HGCL disabled | 84.69 | 82.37 |
| PCL disabled | 87.33 | 86.01 |
| ICL only | 81.37 | 79.59 |
| HGCL only | 82.51 | 79.17 |
| PCL only | 81.58 | 79.25 |
Key Findings¶
- PCL provides little direct gain in [CLS] classification: removing it reduces Top-1 by only 0.16 percentage points. On the final-memory probe, however, the full model gains 0.090 Top-1 and 0.096 macro-F1 over the no-PCL variant. Matching probe granularity to representation granularity reveals the module's value.
- The paper interprets the small AUC increase and larger threshold-metric gains as improved calibration. No dedicated probability-calibration metric is reported, so the more conservative conclusion is improved discrimination at the fixed threshold.
- Tokenizer stride matters more than patch length: Table 6(b) increases stride from 10% to 100% of patch length and lowers FordA accuracy from 87.7% to 62.5%. Loss-weight tuning cannot replace attention to input resolution.
- Short-run HAR window ablation in Table E.4 reports 92.7%–92.8% accuracy at window fractions 0.10–0.30 and 92.3% with a full-sequence window. This supports smaller windows rather than assuming greater depth or window size is always preferable.
Highlights & Insights¶
- The design separates state propagation from whether that state merits direct supervision. Window alignment and explicit readout give memory a testable role beyond supplying the final classifier with more context.
- Cue retention compares the full model with the same architecture trained without PCL. The former tests supervision gains and the latter tests the interface itself, locating contributions more precisely than classification accuracy alone.
- Forecasting, the memory probe, and classification read three distinct interfaces. This evaluation strategy can transfer to other multi-scale models, but requires tasks specifically matched to intermediate states rather than substituting the final readout for every test.
Limitations & Future Work¶
- The authors acknowledge that fixed-stride convolution requires per-dataset selection; coarse strides can damage brief phase and frequency events. Signal-adaptive tokenization or predictive objectives are possible extensions, but are not evaluated here.
- The three InfoNCE objectives require sufficiently large in-batch negative sets, excluding extremely small datasets and making batch size 256 demanding on constrained hardware. Momentum queues are a possible direction, not a demonstrated result.
- Architectural replacements xLSTM and Transformer-XL lack an equivalent interface for PCL and train only with HGCL and ICL. Table 4 therefore compares interfaces together with the supervision they can host, rather than isolating architecture under identical three-loss objectives.
- Streaming PMA cost advantages should not be extrapolated to full PMT: in Table C.2, full PMT has peak memory 12,337 MB and latency 5,381.27 ms, versus 8,336 MB and 3,476.55 ms for the FlashAttention baseline. This non-streaming configuration materializes overlapping windows and includes additional encoders; parameter counts are 20.1M and 7.4M, respectively.
- Configuration descriptions differ within the paper: Appendix B gives convolution kernel and stride both as \(k\), Table 6 varies patch length and stride independently, and Appendix C uses a 16-sample patchifier with 8-fold compression. These should be distinguished by experiment rather than asserting that every experiment uses non-overlapping patches.
Related Work & Insights¶
- vs TS2Vec / SoftCLT: these methods shape multi-scale representations through hierarchical pooling or soft similarity. PMT preserves the token stream and assigns separate PCL supervision to window memory. HGCL derives neighbor weights from window geometry rather than SoftCLT's data-space similarities and DTW.
- vs Transformer-XL / xLSTM: these models propagate history through caches or recurrent states; PMT adds a writable, window-aligned, directly supervisable interface. Its advantage does not hold in every low-label setting, and FordB at 5% still indicates the value of predictive pretraining.
- vs HiMTM / PatchTST: the former learns hierarchies through multi-scale reconstruction and self-distillation, while the latter extends context through patches and attention. PMT focuses on three independent contrastive representation interfaces. Its frozen ridge-regression forecasting protocol does not establish superiority over dedicated end-to-end forecasting models.
Rating¶
- Novelty: 4/5. Window-aligned writable memory and scale-matched supervision constitute a clear interface-level contribution.
- Experimental Thoroughness: 4/5. All three granularities have probes and ablations, but dataset scale, tuning coverage, and pure architectural isolation remain limited.
- Writing Quality: 3/5. The main argument is clear, but gate placement, tokenizer configurations, and some benefit claims require careful reading alongside the appendices.
- Value: 4/5. Useful for time-series representation learning requiring mid-range states, rather than an unconditional forecasting or efficiency breakthrough.