DeepArrhythmia: Segment-Contextualized ECG Arrhythmia Classification via Selective Evidence Acquisition¶
Conference: NeurIPS2026 (task-list grouping; based on arXiv v2, 2026-09-27)
arXiv: 2605.16441
Code: https://github.com/jimmylihui/DeepArrhythmia_repo
Area: Medical Imaging
Keywords: ECG arrhythmia, segment context, beat-level classification, physiological evidence, confidence routing
TL;DR¶
DeepArrhythmia produces R-peak-aligned beat labels within 10-second ECG segments and uses initial segment confidence to decide whether to acquire numerical and morphology evidence, demonstrating the value of context and physiological grounding across four datasets without outperforming always-rich inference on every metric.
Background & Motivation¶
Beat-level arrhythmia classification assigns a class to each heartbeat, but its evidence need not be confined to that heartbeat. Ventricular and supraventricular ectopic beats can both occur prematurely and be followed by a longer interval; discrimination also requires comparing QRS morphology, P-wave behavior, and rhythm changes against neighboring normal beats. Isolated-beat crops can lose these relationships, while end-to-end models processing longer waveforms do not necessarily expose physiological measurements that make their decisions verifiable.
Connecting specialized peak detection and feature extraction tools to a vision-language model (VLM) separates precise measurement from heterogeneous evidence integration. More tools, however, are not necessarily better: easy segments may not need additional measurements, and features derived from poor-quality signals can mislead classification. On VitalDB, the rich-evidence model has Macro-F1 0.858, below the minimal-evidence model's 0.877, illustrating that the problem is not simply adding modalities but identifying when auxiliary evidence is worth acquiring.
The paper reduces this problem to choosing between two evidence budgets: make an initial prediction from the signal, waveform image, and peak locations, then use confidence to decide whether to obtain numerical and textual evidence. Core Idea: preserve multi-beat context and explicit R-peak alignment while using segment-level confidence to control physiological evidence acquisition, rather than invoking every tool for every segment.
Method¶
Overall Architecture¶
The inputs are the raw multichannel signal and rendered waveform image from the same 10-second ECG interval; the output is an N/S/V/F label for each R peak. The three key designs are Peak-Aligned Context, Confidence-Gated Routing, and Rich-Evidence Integration: localize beats and predict with minimal evidence first, then acquire rhythm–morphology measurements and morphology text only when confidence is insufficient, before classifying again.
A frozen ECG-CoCa encoder processes the signal, while the image is encoded separately; both representations are projected into the language-model embedding space of Qwen3.5-4B. The central model calls tools and integrates evidence rather than performing every precise measurement itself. The peak detector, handcrafted feature computation, and confidence calculation should not all be interpreted as autonomous agents negotiating with one another.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["10-second signal and waveform image"] --> A["Peak-Aligned Context<br/>R peaks and minimal-evidence prediction"]
A --> B["Confidence-Gated Routing"]
B -->|Confidence reaches threshold| O["Beat-level N/S/V/F"]
B -->|Confidence below threshold| C["Rich-Evidence Integration<br/>Numerical features and morphology text"]
C --> O
D["Development labels<br/>Threshold selection and routed SFT"] -.Training supervision.-> B
T["Training-subject labels<br/>Teacher-rationale distillation"] -.Training supervision.-> C
Solid edges represent test-time data flow; dashed edges represent training supervision only. At test time, development labels are unavailable and the morphology analyzer receives no ground-truth beat class.
Key Designs¶
1. Peak-Aligned Context: classify individual beats using evidence from the whole segment
Ten seconds denotes a fixed physical duration, not an identical sample count across datasets. Signal length varies with sampling rate, and the number of beats varies with heart rate; the input is neither a fixed number of isolated beats nor a segment assigned one overall class. The coarse AAMI mapping retains N (normal), S (supraventricular ectopic), V (ventricular ectopic), and F (fusion), excluding unknown or unclassifiable class Q.
The peak detector uses 1D U-Net++ to produce point-wise probabilities, followed by thresholding and peak selection to obtain R peaks. Peak locations jointly specify where final labels are emitted, where numerical features are calculated, and which beat the morphology analyzer examines. The central model serializes answers as “sample position:class,” making temporal anchors the shared index that binds evidence to an instance rather than incidental metadata.
Preserving neighboring beats makes prematurity, subsequent pauses, and morphological consistency comparable relationships. Numerical features quantify how early a target beat occurs, whereas the waveform image helps assess whether its shape remains close to neighboring normal beats. The image is rendered from the raw waveform and adds no independent measurement information: its contribution is an alternative representation, not an additional independent sensor.
This design also exposes upstream dependencies. A missed peak removes a classification instance, a shifted peak contaminates local features, and a spurious peak creates unrealistic interval relationships. The no-peak-detector ablation cannot perform reliable alignment and is marked unavailable, not assigned zero performance. Beats at fixed-window boundaries may also lack complete preceding or following intervals, so 10-second context does not imply equally complete evidence for every beat.
2. Confidence-Gated Routing: average beat posteriors to decide whether the whole segment needs richer evidence
Minimal evidence consists of the raw signal, waveform image, and R peaks. The central model first generates beat-level predictions and applies a softmax over canonical N/S/V/F class-token logits at each label-generation position. Each beat's confidence is its largest class probability, and segment confidence is the mean across all beats. For a segment with \(K_x\) beats, the routing statistic is:
This mean is neither the probability that the segment contains an abnormal beat nor a direct estimate of the benefit of additional tools. It is a practical proxy for evidence sufficiency. At or above the dataset threshold, the initial prediction is returned; below it, the Feature Extractor and Morphology Analyzer are invoked and a rich-evidence answer is generated. Routing occurs once per segment because both optional tools exploit segment rhythm and cross-beat relationships; independent per-beat routing was not evaluated.
Averaging prevents one uncertain beat from automatically escalating the entire segment, but many confident normal beats can also raise the average. Appendix E compares mean confidence with minimum beat confidence on MIT-BIH: both achieve Macro-F1 0.588, while Micro-F1 is 0.951 for the mean rule and 0.948 for the minimum rule. This is empirical support within the reported protocol, not proof that the mean is always preferable for rare abnormalities.
The threshold is neither chosen purely by intuition nor selected using test labels. Development subjects held out from Stage 1 optimization are used to compare minimal- and rich-evidence predictions, sweep thresholds, and maximize beat-level Micro-F1. The selected threshold is fixed before testing: 0.990529 for MIT-BIH, 0.980933 for MIT-BIH-SUP, 0.993532 for INCART, and 0.98519 for VitalDB.
These near-one thresholds reflect normal-class dominance and concentrated high initial confidence; they should not be interpreted like a generic 0.5 binary-classification cutoff. Because threshold selection optimizes Micro-F1, deployments prioritizing rare-beat misses or expert-review workload would require a different objective. Confident errors after a device shift may also fail to trigger escalation.
3. Rich-Evidence Integration: numerical measurements and distilled morphology text play different roles
The Feature Extractor calculates RR intervals, normalized RR intervals, R-peak amplitude, higher-order statistics, and numerical morphology descriptors around the same peak anchors. Higher-order statistics split a beat into five intervals and calculate skewness and kurtosis, yielding ten dimensions; morphology descriptors use distances from the R peak to representative local points. These measurements turn prematurity, amplitude changes, and shape differences into comparable numbers rather than requiring the central language model to infer precise timing and amplitude from an image.
The Morphology Analyzer is a separate Qwen3.5-4B-based visual-language student. It reads the ECG image, target-beat location, and auxiliary numerical context to generate morphology text. During training, the Gemini 3.1 teacher sees ground-truth classes and produces label-supported explanations exclusively for training subjects. At test time, the student receives no true labels; its text becomes supportive evidence for the central model and is retained for inspection.
This distinction matters: the text is neither an independent clinician's judgment nor an observation entirely independent of class supervision. Label-conditioned teacher distillation can transmit bias or rationalization tendencies. “Auditable” therefore means that intermediate outputs can be inspected, not that explanations have been proven fully faithful. The paper's evidence-deletion experiment supports local evidence dependence but does not eliminate all explanation risks.
The central model serializes explicit tool-call and tool-output spans into its context, combines signal, image, peaks, numbers, and text under rich evidence, and generates peak-aligned labels. Three autonomous agents do not vote on the answer; specialized modules produce evidence and the central model decides. Additional evidence changes the conditioning inputs without changing the requirement that each label identify a specific R peak.
A Worked Example¶
Appendix L presents a segment with 13 beats. The minimal-evidence model initially labels all beats N, but segment confidence is 0.989392, below the MIT-BIH threshold of 0.990529, so both optional tools are called.
The target beat at position 2045 occurs at approximately 5.68 seconds. Its preceding RR interval is 236 samples and its following interval is 358 samples. The morphology text describes premature timing, a narrow QRS resembling neighboring beats, and a subsequent longer interval. The central model changes that position from N to supraventricular ectopic while retaining N for the other positions.
The example illustrates complementarity: prematurity and a pause alone do not distinguish S from V; QRS comparison is also needed. The source example writes the final label as SVEB, whereas the formal task uses canonical S. This is a semantic mapping, not an additional class, and the single example should not be treated as a clinical diagnostic rule.
Loss & Training¶
The central model undergoes two-stage masked supervised fine-tuning (SFT). Generated tool-call and final-answer tokens contribute to the loss; externally returned tool outputs and other context tokens do not. Training therefore targets calling behavior and classification, rather than making the model imitate precise tool measurements.
Training and test subjects are first separated. Training subjects are then split 9:1 into Stage 1 optimization set \(D_1\) and Stage 2 route-induction set \(D_2\). Stage 1 trains separate minimal- and rich-evidence specialists to learn output syntax, peak alignment, and prediction under different evidence budgets. Stage 2 selects the threshold on \(D_2\) and constructs threshold-induced tool-use trajectories.
The routed model is initialized from the minimal specialist and fine-tuned on these trajectories. Routing supervision denotes “stop or acquire evidence,” not “normal or abnormal.” Development labels measure classification outcomes for threshold selection, but test labels are excluded from threshold selection, tool-rationale distillation, and central-model fine-tuning. Reuse of training data must be understood within this subject-level separation.
Class imbalance is mitigated through shifted window re-anchoring: additional 10-second windows are generated around non-N beats, placing their R peaks at different relative positions and assigning more offsets to rarer classes. Windows are clipped to valid record boundaries and deduplicated. Only the training partition is augmented; the test distribution is not altered to improve scores.
The ECG-CoCa encoder remains frozen and the central backbone is Qwen3.5-4B. Training uses learning rate \(2\times10^{-5}\), maximum sequence length 8000, bfloat16, FlashAttention-2, and DeepSpeed ZeRO-3. Three A6000 GPUs, per-device batch size 1, and gradient accumulation 4 yield effective batch size 12; this differs from the three independent single-GPU workers used for inference-latency measurement.
Key Experimental Results¶
Main Results¶
All four datasets use subject-disjoint DS1/DS2-style evaluation and report beat-level Macro-F1 and Micro-F1. Machine-learning and deep-learning baselines average three random seeds, while VLM decoding uses temperature 0. Macro-F1 excludes classes absent from the test ground truth.
The following selection from Table 1 highlights evidence budgets and baseline differences without reproducing every model. Each cell is Macro-F1 / Micro-F1.
| Configuration | MIT-BIH | MIT-BIH-SUP | INCART | VitalDB |
|---|---|---|---|---|
| Minimal-evidence specialist | 0.534 / 0.917 | 0.562 / 0.938 | 0.620 / 0.963 | 0.877 / 0.923 |
| Rich-evidence specialist | 0.593 / 0.950 | 0.614 / 0.952 | 0.710 / 0.981 | 0.858 / 0.911 |
| Threshold-induced routed model | 0.588 / 0.951 | 0.615 / 0.953 | 0.689 / 0.979 | 0.885 / 0.922 |
| Full-evidence non-agentic fusion | 0.567 / 0.937 | 0.536 / 0.926 | 0.641 / 0.967 | 0.821 / 0.900 |
| SVM | 0.523 / 0.934 | 0.575 / 0.938 | 0.681 / 0.985 | 0.600 / 0.747 |
| Gemini 3.1 | 0.410 / 0.878 | 0.522 / 0.849 | 0.653 / 0.924 | 0.389 / 0.693 |
Rich evidence improves over the minimal specialist on the first three datasets but hurts on VitalDB. Routing has the highest VitalDB Macro-F1, although its Micro-F1 0.922 remains slightly below the minimal specialist's 0.923. On INCART, the rich specialist beats routing on both metrics, and SVM Micro-F1 0.985 exceeds the rich specialist's 0.981. The source's broad best-performance wording should therefore not be expanded into “state of the art on every dataset and metric.”
Language-model baselines receive peaks and numerical features but not outputs from the dataset-supervised Morphology Analyzer, so their supervision budget is not identical to full DeepArrhythmia. The non-agentic fusion baseline uses comparable signal, image, and textual representations with a supervised classification head, providing a more direct integration comparison; nevertheless, the entire difference cannot be attributed solely to tool calling.
Ablation Study¶
The full MIT-BIH configuration in Table 2 corresponds to the rich-evidence performance level, not the routed model with Macro-F1 0.588 in Table 1. Original precision is retained below; unavailable results are not replaced with invented scores.
| MIT-BIH configuration | Macro-F1 | Micro-F1 |
|---|---|---|
| Full DeepArrhythmia | 0.5929 | 0.9501 |
| Without shift augmentation | 0.4259 | 0.9120 |
| Without ECG image | 0.5394 | 0.9426 |
| Without Feature Extractor | 0.3527 | 0.9066 |
| Without Morphology Analyzer | 0.5694 | 0.9433 |
| Without Peak Detector | — | — |
Removing numerical features lowers Macro-F1 by 0.2402, compared with 0.0235 for removing the Morphology Analyzer; removing shift augmentation lowers it by 0.1670. These ablations establish contributions within the full configuration, not an isolated estimate of routing's benefit.
Table 3 latency is measured in seconds per segment and includes the corresponding evidence-acquisition pipeline, not just language-model decoding. Evaluation uses three independent single-GPU A6000 workers, batch size 1, bfloat16, and greedy decoding with up to 8192 new tokens.
| Inference configuration | MIT-BIH | MIT-BIH-SUP | INCART | VitalDB |
|---|---|---|---|---|
| Minimal-evidence specialist | 2.83 | 2.73 | 2.88 | 3.01 |
| Rich-evidence specialist | 3.89 | 3.85 | 3.94 | 4.02 |
| Rich-evidence non-agentic fusion | 3.40 | 3.39 | 3.50 | 3.59 |
| Simple routing | 3.24 | 3.28 | 3.53 | 3.48 |
| Threshold-induced routing | 3.58 | 3.40 | 3.45 | 3.27 |
Threshold routing is faster than always-rich inference but not always faster than simple routing. Removing the autoregressive central backbone also leaves substantial evidence-processing cost. The source latency paragraph gives a simple-routing range of 3.24–3.48 seconds, whereas Table 3 lists 3.53 seconds for INCART; the table value is retained and the conflict explicitly flagged rather than silently corrected.
Key Findings¶
- Evidence utility differs by dataset: Appendix D reports feature-only Macro-F1 0.5209 on MIT-BIH and 0.3851 on VitalDB, supporting weaker auxiliary numerical evidence on the latter without independently establishing whether signal noise or label noise causes it.
- Minority classes remain difficult: Table 16 gives MIT-BIH S-class F1 0.3922 and F-class F1 only 0.0659; INCART F-class F1 is 0.0992. High aggregate Micro-F1 cannot substitute for rare-abnormality recognition.
- Explanation evaluation is not a utility trial: Three experts blindly rate 200 outputs per system; the morphology student scores 4.92/5 for physiological correctness and 4.19/5 for overall usefulness, below the label-conditioned teacher's overall 4.97/5. This does not validate clinical outcomes or workflow benefit.
Highlights & Insights¶
- Acquisition matters more than tool count: The same tools can support classification on one dataset and introduce noise on another. Making evidence need part of the policy addresses budget control more directly than mandatory full-tool invocation.
- Temporal anchors organize multimodal alignment: Peak positions bind waveforms, numbers, morphology text, and final labels to the same heartbeat. The transferable idea is the event-instance alignment interface, not applying ECG labels unchanged to every physiological signal.
- Redundant representations can improve usability: The waveform image adds no independent information but makes morphology discarded by scalar features more accessible to the backbone. Its value should be measured through representation ablation, not explained merely by a larger modality count.
Limitations & Future Work¶
- Cross-dataset transfer remains unresolved: Appendix J evaluates a fixed-seed 10% sample of each target test set, retaining source models and source routing without target-data tuning. Mean Micro-F1 over twelve off-diagonal transfers is 0.698 versus 0.457 for 1D ResNet, but VitalDB-to-MIT-BIH-SUP reaches only 0.0935. This is not certification of clinical generalization.
- Mean confidence can hide an individual abnormality: Both mean routing and Micro-F1-based threshold selection are influenced by majority classes. Risk-aware quantiles or per-beat routing merit study, but their costs, calibration, and miss rates require validation; more aggressive tool use cannot be assumed beneficial.
- Explanation and noise evidence have boundaries: The morphology student depends on label-conditioned distillation. Appendix Q keeps clean peaks and RR intervals fixed to isolate waveform-quality effects, so it does not represent deployment where noise simultaneously disrupts detection and classification.
- The intended role is offline assistance: The study is retrospective and targets Holter review and second reading of flagged segments, not hard real-time device alarming or autonomous diagnosis. Multicenter prospective validation, privacy governance, out-of-distribution detection, and human review remain necessary.
Related Work & Insights¶
- vs SVM and handcrafted ECG features: Both use rhythm and morphology measurements; this framework additionally integrates signal, image, and text through a central model. SVM's higher INCART Micro-F1 demonstrates that conventional models remain competitive.
- vs GEM, PULSE, and ECG-R1: These methods emphasize multimodal ECG understanding or interpretation, whereas this paper explicitly outputs R-peak-aligned beat labels and consumes external measurements as intermediate decision evidence. Table 1 under this task and supervision protocol does not establish superiority over all their ECG capabilities.
- vs ECG-Agent and general tool use: The emphasis is not open-ended multi-turn planning but supervised confidence gating between two evidence states. Directly predicting the marginal benefit of additional tools, rather than using classification confidence as its proxy, is a direction worth testing.
Rating¶
- Novelty: 3.5/5. Individual components are established; the contribution is their combination for contextual beat prediction, explicit physiological interfaces, and selective acquisition.
- Experimental Thoroughness: 4/5. Four datasets, subject separation, ablations, and transfer analysis provide broad evidence, but prospective clinical and real-device validation are absent.
- Writing Quality: 3.5/5. The method is clear, but broad best-performance statements do not match every table entry, and the latency range conflicts with the table.
- Value: 4/5. Concrete interfaces support inspectable offline ECG assistance, without yet justifying autonomous clinical deployment.