Revitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language Perspective¶
Conference: NeurIPS2026
arXiv: 2609.34652
Code: https://github.com/Levi-Ackman/ViRe
Area: Time Series
Keywords: medical time series, waveform morphology, Vision Query, dual-path retrieval, frozen CLIP
TL;DR¶
ViRe renders the same numerical EEG/ECG segment as a waveform and uses frozen CLIP visual features as a query to retrieve evidence from temporal and channel numerical tokens; the mean six-metric score across six benchmarks rises from Medformer's 77.90 to 82.90, but ViRe does not lead on every dataset, and this is not patient-level diagnostic accuracy.
Background & Motivation¶
EEG/ECG classification requires recognizing both local waveform events and relationships among channels. CNNs, recurrent networks, and Transformers such as Medformer already process numerical signals directly, but must learn from task labels which variations deserve preservation. Clinical waveform inspection instead exposes spikes, waveform complexes, and synchronous cross-lead changes in a visible layout. That morphological organization can provide a prior for numerical feature aggregation without learning everything again from limited patient labels.
Rendering a waveform as an image does not collect new information: the image and numerical matrix originate from the same preprocessed segment. The potential complementarity lies in representation and pretraining knowledge. Numerical compression may attenuate some shape structures, whereas an image-text-pretrained visual encoder may retain a different morphological summary. General-purpose CLIP is not equivalent to a medical expert, however, and image-text pretraining does not establish that it understands pathological concepts; ablations, concept probes, and clinical association analyses must test that hypothesis.
The method neither asks the visual model to predict disease classes directly nor supplies patient reports or invokes the CLIP text encoder. Core Idea: let a frozen visual representation determine which numerical evidence to attend to, use one global Vision Query to aggregate temporal and channel evidence separately, and classify from the resulting numerical evidence summaries.
Method¶
Overall Architecture¶
The input is a multichannel medical time-series segment, and the output is its class logits. Morphology Query Construction encodes a deterministic waveform image into a global query; Dual-Axis Numerical Encoding produces temporal and channel tokens; Dual-Path Query Retrieval uses the visual query to read both numerical feature sets, after which the two summaries are added and linearly classified.
“Retrieval” here means cross-attention within the current sample, not searching an external case database. The paths read different numerical organizations of the same segment. CLIP supplies only the Query, while Keys and Values come from numerical encoders; there is no visual-feature shortcut that bypasses retrieval to perform classification.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
X["Same-source numerical segment"] --> Q["Morphology Query Construction<br/>Deterministic rendering and frozen CLIP"]
X --> N["Dual-Axis Numerical Encoding<br/>Temporal and channel tokens"]
Q -->|Shared Query| R["Dual-Path Query Retrieval"]
N -->|Numerical Keys and Values| R
R --> O["Add two summaries<br/>Linear classification"]
O -.->|Training only| L["Cross-entropy supervision"]
Y["Segment label"] -.-> L
Solid arrows indicate data flow shared by training and inference, and dashed arrows indicate training-only supervision. CLIP weights remain frozen, but the dimension-alignment projection, numerical encoders, two retrieval blocks, and classifier are optimized.
Key Designs¶
1. Morphology Query Construction: turn waveform layout into a condition for evidence selection
Each channel is plotted separately, and the panels are vertically stacked in channel order into an RGB image. The Appendix D.3 operator uses fixed axes and removes ticks, spines, legends, and grids, encouraging the image encoder to represent morphology and cross-channel temporal correspondence rather than plot decorations. The frozen CLIP vision encoder produces a global feature for the image, and a trainable linear projection maps it to the common numerical feature dimension as a single Vision Query.
This does not produce a visual token for every timestamp or a separate diagnosis for every channel. The segment shares one morphological summary, making the visual branch a condition for deciding which numerical patterns to aggregate; precise numerical information remains in the downstream Keys and Values. Since the image is a deterministic transformation of the original signal, its contribution is a pretrained representational inductive bias, not an independent sensor observation.
Rendering still introduces an information bottleneck: many channels share a limited image area, and line width and resolution affect visible morphology. Appendix H shows that this preprocessing step matters. However, the cached DPI text contains duplicated digits and cannot establish exact default settings.
2. Dual-Axis Numerical Encoding: let the query retrieve both time intervals and channels
The temporal path divides the signal into non-overlapping patches of length \(L\). Each patch contains all channels and is flattened, linearly projected, and given positional encoding. A temporal token therefore represents a local cross-channel state, and a separate Transformer encoder models relationships among intervals. The number of temporal tokens is \(P=\lceil T/L\rceil\); when the sequence length is not divisible by the patch length, the cached text does not specify final-patch completion, so this note does not invent an implementation rule.
The channel path instead linearly maps each channel's complete trajectory into one token, adds channel positional encoding, and uses another Transformer to learn inter-channel relationships. It is not a second collection of temporal patches: it preserves an organization describing each channel's behavior within the segment. The temporal path supports local event selection, and the channel path provides cross-lead structure; both use a common feature dimension but separate encoders.
The two organizations also define the explanation granularity. Temporal retrieval weights refer to temporal tokens, and channel retrieval weights refer to channel tokens; neither should be indiscriminately described as sample-by-sample lesion localization. TDBrain uses channel encoder depth 0 as an implementation exception. This does not remove channel tokens or the channel retrieval path.
3. Dual-Path Query Retrieval: vision determines weights, while numerics supply the aggregated content
The same aligned global query enters two cross-attention encoders. Temporal features provide one Key–Value set, and channel features provide the other, yielding two evidence summaries of length 1. Using the notation of the original Equation (6), the central operation is:
The asymmetry means that vision occupies only the Query position and is not directly inserted into the classification representation as a Value. Sharing a query does not imply that all parameters are shared: the text describes two cross-attention encoders. After retrieval, the summaries are added and a linear layer produces class logits; visual features are not additionally concatenated as an independent diagnostic branch.
Why is this more than adding another branch? A single attention head performs weighted pooling of numerical Values based on matching the query against numerical Keys. Changing the waveform summary can therefore change how the current numerical tokens are read. Appendix A shows that identical attention scores recover mean pooling. ViRe thus includes a morphology-independent numerical aggregation special case, while the usefulness of the visual prior still depends on the learned weights.
This inclusion concerns empirical risk under ideal optimization, not whether actual training finds a better solution or reduces risk on unseen patients. The conditional-information argument likewise assumes that the visual representation retains label information relative to already compressed numerical features. It does not imply that images contain more information than the original signal.
A Worked Example¶
Consider a PTB heartbeat segment from the primary benchmark: it has 15 channels and 300 timestamps, with temporal granularity \(L=1\). The temporal path therefore produces 300 tokens, and the channel path produces 15 tokens. The same segment is rendered as 15 vertically stacked panels; frozen CLIP compresses the image into one query, projected to 128 dimensions.
Temporal retrieval aggregates one summary from the 300 temporal tokens, and channel retrieval aggregates another from the 15 channel tokens. Their sum produces binary classification logits. If several leads change synchronously, the query may increase weights on the corresponding interval or channels. This is an illustration of the mechanism, not a claim that every sample attends to the same location or a description of final patient diagnosis.
During training, the segment label supplies cross-entropy supervision only at the classification end. Testing requires neither labels nor clinical reports. Online prediction for a new input still requires rendering and visual encoding unless that segment's visual features have already been cached.
Loss & Training¶
The model uses standard classification cross-entropy without an additional image-text contrastive loss or morphology annotations. ViRe uses batch size 128, feature dimension 128, and temporal encoder depth 6. Channel encoder depth is normally 6 and is 0 for TDBrain. The values of \(L\) for APAVA, TDBrain, ADFTD, PTB, PTB-XL, and MIMIC are 1, 3, 8, 1, 6, and 6, respectively.
Training uses Adam at learning rate \(10^{-4}\) for up to 100 epochs, with early stopping on validation macro-F1 and patience 10. Main experiments use fixed subject-disjoint splits and five random seeds, and evaluate the checkpoint with the best validation result. The baseline appendix specifies batch size 32 for APAVA/TDBrain, whereas ViRe uses 128; a common training protocol therefore does not mean every hyperparameter is identical.
Appendix C uniformly selects one of six augmentations per forward pass: temporal flipping, channel shuffling, temporal masking, frequency masking, jittering, or dropout. It requires augmenting the numerical segment first and regenerating its image from that same transformed segment, preventing view mismatch; validation and testing disable augmentation. Appendix D.2 nevertheless states that frozen CLIP features are cached once per sample. The relationship between this caching and dynamic re-rendering is unresolved, and this note does not combine them into a supposedly verified execution procedure.
Key Experimental Results¶
Main Results¶
The six primary benchmarks split by subject, but model inputs and the main metric units are segmented windows or heartbeats. EEG and PTB-XL use 1-second windows, while PTB/MIMIC use R-peak-aligned heartbeats. Test segments originate from unseen patients; this is not equivalent to reporting diagnoses aggregated for individual patients.
Avg is the arithmetic mean of Accuracy, macro-Precision, macro-Recall, macro-F1, macro-AUROC, and macro-AUPRC, all expressed on a percentage scale. The final row then averages Avg across six datasets and must not be called accuracy. Values below come from Table 2.
| Dataset | ViRe Avg | Medformer Avg | Strongest non-ViRe baseline Avg | Difference from strongest baseline (points) |
|---|---|---|---|---|
| APAVA | 92.79 | 79.74 | Medformer 79.74 | +13.05 |
| ADFTD | 59.83 | 54.63 | Medformer 54.63 | +5.20 |
| TDBrain | 95.54 | 91.91 | Medformer 91.91 | +3.63 |
| PTB | 89.08 | 84.69 | iTransformer 84.95 | +4.13 |
| PTB-XL | 69.47 | 69.28 | PatchTST 69.90 | -0.43 |
| MIMIC | 90.68 | 87.14 | Reformer 87.90 | +2.78 |
| Six-dataset mean | 82.90 | 77.90 | Medformer 77.90 | +5.00 |
The abstract's 6.42% is the relative improvement of the overall mean over Medformer: the increase of 5.00 divided by 77.90. It is neither a 6.42-percentage-point gain nor the average of per-dataset relative improvements. PTB-XL particularly requires metric-wise reading: Table 13 reports macro-F1 of 61.59 for ViRe, 62.02 for Medformer, and 62.61 for PatchTST. A slightly higher Avg than Medformer does not mean every metric improves.
Ablation Study¶
The table selects macro-F1 means on four datasets from Tables 3–5. Avg. Gain retains the authors' aggregate values: average relative improvement over the no-retrieval variant across eight entries comprising Accuracy and F1 on four datasets, not the six-metric Avg used above.
| Config | ADFTD F1 | APAVA F1 | PTB F1 | MIMIC F1 | Avg. Gain |
|---|---|---|---|---|---|
| No retrieval; add numerical paths | 51.73 | 82.45 | 74.58 | 84.81 | — |
| Zero Query retrieval | 51.23 | 83.73 | 80.21 | 86.06 | +1.92% |
| Gaussian Query retrieval | 51.54 | 81.65 | 77.34 | 86.44 | +0.76% |
| Direct visual-feature addition | 53.53 | 85.45 | 81.26 | 87.49 | +4.05% |
| Visual-feature concatenation | 53.55 | 83.02 | 79.13 | 87.81 | +3.02% |
| Randomly initialized ViT Query | 27.34 | 77.18 | 73.50 | 83.58 | -11.81% |
| ImageNet ViT Query | 50.55 | 84.86 | 75.72 | 86.42 | +0.34% |
| CLIP Vision Query retrieval | 54.02 | 91.19 | 85.81 | 88.54 | +7.72% |
The results separately support two observations: informative visual queries outperform content-free queries, and cross-attention retrieval outperforms simply adding visual features. A random ViT substantially degrades performance, so adding an image encoder alone is insufficient. However, the CLIP–ImageNet difference cannot be attributed exclusively to language supervision from this table; pretraining data and implementation settings may also contribute.
Key Findings¶
- No retrieval is not the worst variant on every individual metric: its ADFTD F1 of 51.73 exceeds the Zero Query's 51.23 and Gaussian Query's 51.54. This note does not reproduce the text's absolute claim that removing retrieval yields the worst results.
- Appendix F measures APAVA at batch size 128. Cached training takes 211.9/313.1 ms and 698.2/838.4 MB peak memory for ViRe/Medformer; online inference takes 267.7/152.6 ms and 740.8/408.5 MB. Cached training measurements cannot establish cheaper online inference for new inputs.
- Appendix I reports 64.93 versus 61.57 on the supplementary PTB-XL 13-diagnosis protocol, a +3.36 difference with patient-cluster 95% CI [+1.46,+5.10], and 7/7 wins on repeated splits. This is a different task/split protocol, not a replacement for the primary PTB-XL result.
- In the patient-matched visual interventions of Table 8, zero masking reduces macro-F1 by 6.19 points and cross-patient shuffling by 15.61 points. This supports matching query content to the current segment, but interventions also introduce distribution shifts and do not independently establish a clinical causal explanation.
The paper additionally tests whether retrieval favors morphological regions using attention-density enrichment. Let \(A_t\) be normalized temporal attention and \(\rho_A(S)\) the mean attention density per timestamp in a region. \(S_{\mathrm{curv}}\) contains the top 20% timestamps by channel-averaged absolute second differences, \(S_{\mathrm{qrs}}\) is the QRS region, and bars denote temporal complements:
| ECG dataset | ViRe MAR | Gaussian Query MAR | ViRe CAR | Gaussian Query CAR |
|---|---|---|---|---|
| PTB | 1.81 | 1.46 | 1.65 | 1.33 |
| PTB-XL | 1.67 | 1.38 | 1.51 | 1.27 |
| MIMIC | 1.72 | 1.41 | 1.63 | 1.35 |
Values above 1 in Table 7 indicate denser attention inside the target region than in its complement, and ViRe exceeds the Gaussian Query on both metrics. High-curvature and QRS enrichment demonstrate selective attention, not correct disease localization or causal interpretability. Frozen-representation probes on MEETI add evidence of decodable concepts: amplitude-range Spearman correlation is 0.903 versus 0.501 for random features, and prolonged-QT AUROC is 0.904 versus 0.656. These are not accuracy results for ViRe's main classification task.
Highlights & Insights¶
- Visual priors control reading rather than replace numerical evidence: the Query-only structure helps localize the contribution to the aggregation rule. It is transferable to tasks where numerical precision matters but morphology can guide selection, rather than indiscriminately concatenating images and numerics.
- Same-source representations can complement through different compression: learning with limited labels can improve without adding observations. A more informative comparison is morphology-conditioned pooling versus numerical adaptive pooling, not merely a larger number of input modalities.
- Mechanistic evidence covers several levels: query substitution, fusion, pretraining, attention enrichment, and concept probes examine different stages. They are complementary evidence, not a combined demonstration that CLIP already possesses clinical diagnostic knowledge.
Limitations & Future Work¶
- Clinical extrapolation remains limited: the primary study is retrospective segment classification. Subject-disjoint splits prevent same-patient leakage, but many heartbeats do not imply many independent patients. Patient-level aggregation, calibration, cross-device/center validation, and prospective studies remain necessary; the authors position the system as decision support.
- Rendering and augmentation assumptions need testing: general-purpose CLIP may miss subtle pathology, and the stated rationale does not establish that temporal flipping and channel shuffling always preserve label semantics. These should be tested by disease and lead definition, alongside rendering scale, resolution, and extreme-noise robustness.
- Implementation conflicts are preserved: the relationship between re-rendering augmented signals and once-per-sample caching is unspecified. Table 1 lists 72 TDBrain subjects, while Appendix B states 25 PD + 25 HC and lists test subject IDs reaching 53. This note does not guess which cohort size or partition is correct.
- Numerical conflicts are preserved: APAVA ViRe F1 is \(91.19\pm1.03\) in the ablation tables but \(91.19\pm1.01\) in Table 13; ADFTD is respectively \(54.02\pm2.22\) and \(54.02\pm3.19\). The Table 5 discussion's claimed 7.11% CLIP improvement over ImageNet does not directly match aggregate gains of +7.72%/+0.34%, and its calculation is unexplained.
- Cache extraction and reproducibility details are incomplete: Appendix H contains duplicated low/moderate-resolution digits, “55” and “5050–200200 DPI,” which are not silently corrected into exact parameters. The text does not clearly specify the CLIP checkpoint or final-patch handling. These require checking the public implementation, which this offline note has not inspected.
- Extensions have separate evaluation boundaries: Appendix I's irregular forecasting/classification supports transfer, but does not use the same evaluation as the six primary EEG/ECG benchmarks. Joint clinical-report modeling, specialized waveform encoders, and adaptive rendering remain future directions, not implemented ViRe components.
Related Work & Insights¶
- vs Medformer: Medformer models numerical signals using multi-granularity temporal patches, whereas ViRe adds a global visual query to control dual-axis numerical aggregation. Its primary overall mean is higher, but the online visual front end is heavier, and PTB-XL macro-F1 does not exceed Medformer.
- vs iTransformer / PatchTST: channel tokens and temporal patches are established modeling choices; the contribution is using a shared morphological query to read both organizations. iTransformer is the strongest non-ViRe Avg baseline on PTB, and PatchTST on PTB-XL, so Medformer must not be labeled the best baseline in every column.
- vs ViTime / Time-VLM / TS-CLIP: the paper characterizes these as image-space forecasting, visual/textual augmentation fusion, and time-series–text alignment, respectively. ViRe's main model has no text input; vision is a numerical-retrieval condition rather than the primary prediction space.
- Research direction: compare visual queries, learned numerical queries, and waveform-specific queries with identical numerical backbones and parameter budgets, then test patient-level generalization. This could better distinguish generic adaptive pooling, general-purpose pretraining, and medical morphological knowledge; it is a direction proposed by this note, not a reported experiment.
Rating¶
- Novelty: 4/5. Restricting frozen waveform priors to a shared Query distinguishes the method from direct fusion, although dual-axis encoding and attention components have precedents.
- Experimental Thoroughness: 4/5. Six primary benchmarks, ablations, interventions, and transfer analyses provide broad evidence, but patient-level clinical validation and some reproduction details remain missing.
- Writing Quality: 3/5. The mechanism is clear, but cohort-size, caching/augmentation, and statistical inconsistencies reduce reproducibility.
- Value: 4/5. Morphology-conditioned aggregation is reusable for medical time series, while deployment cost and cross-center validity still require evaluation.