PhenoLeaf-TS: A Time-Series Benchmark for Leaf Instance Segmentation, Tracking, and Growth Stage Classification¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://pisyntor.github.io/PhenoLeaf-TS
Area: Segmentation
Keywords: Plant Phenotyping, Leaf Instance Segmentation, Multi-Object Tracking, Growth Stage Classification, Time-Series Benchmark
TL;DR¶
PhenoLeaf-TS presents the first large-scale time-series plant phenotyping benchmark spanning 21 Arabidopsis genotypes, 318 plant replicates, and 17,082 top-down RGB images with temporally consistent 31-colour-coded leaf instance masks, systematically evaluating 21 baseline models across segmentation, tracking, and growth classification while demonstrating up to +51 mAP cross-dataset transfer gains.
Background & Motivation¶
Non-invasive quantification of leaf-level morphological traitsโincluding projected surface area, relative expansion rate, length, width, and rosette geometryโis foundational to high-throughput plant phenotyping. These phenotypic measurements fundamentally necessitate instance-level scene understanding: each individual leaf must not only be segmented against complex background soil and severe mutual leaf occlusions, but also tracked continuously across multi-week developmental trajectories and mapped to its broader whole-plant developmental stage. However, computer vision applications in plant phenotyping have long been hindered by a critical lack of temporal depth and persistent annotation consistency, precluding unified benchmarking across segmentation, tracking, and developmental classification.
Existing public plant phenotyping benchmarks exhibit severe structural trade-offs. The seminal CVPPP Leaf Segmentation Challenge datasets (Ara2012, Ara2013, Ara2014) established rigorous per-leaf instance segmentation baselines, yet their collections comprise isolated, un-ordered snapshots without temporal alignment. Field-scale datasets such as PhenoBench offer crop-weed panoptic masks under agricultural conditions, but omit per-leaf temporal identity tracking. Conversely, tracking-focused datasets like Komatsuna provide temporal sequences with tracking annotations, but remain strictly confined in scale (only 15 plant sequences across 900 images) and limited to a single crop species; multi-modal benchmarks such as MSU-PID offer time series but lack full leaf instance-level annotations. Consequently, the research community lacked a public benchmark that simultaneously delivers high-frequency temporal sequencing of the same plants, temporally consistent per-leaf instance identities, and sufficient genetic scale to benchmark segmentation, tracking, and growth stage classification jointly.
To resolve this limitation, the authors developed PhenoLeaf-TS using automated gantry-mounted RGB imaging in environmentally controlled growth chambers across the full rosette life cycle. Core idea: construct the first benchmark of 17,082 temporally ordered RGB images across 21 Arabidopsis genotypes, employing a deterministic 31-colour palette where emergence-ordered colour assignments persist across the entire developmental trajectory, enabling direct evaluation of leaf instance segmentation, tracking, and growth stage classification without post-hoc bipartite matching.
Method¶
Overall Architecture¶
The PhenoLeaf-TS benchmark establishes an integrated pipeline spanning automated multi-genotype imaging, deterministic colour-coded temporal annotation, and standardized multi-task evaluation protocols. During data acquisition, 318 individual plants from 21 Arabidopsis thaliana genotypes were grown on height-adjustable tables inside climate-controlled growth chambers. An automated gantry-mounted RGB camera (4.92 MP) captured top-down imagery at regular intervals (2 or 4 captures per day) from approximately 14 days after sowing through full rosette maturity, yielding 17,082 plant-centered sequence frames (\(\sim 530 \times 525\) resolution). For annotation, a fixed palette of 31 perceptually distinct colours was deployed to assign sequential identity codes based on initial leaf emergence, guaranteeing that every physical leaf maintains an identical colour mask across its lifespan. The benchmark then defines three standardized evaluation tasks evaluated across 21 representative computer vision models.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Gantry Time-Series Imaging<br/>21 Genotypes / 318 Replicates / 17,082 RGB Frames"] --> B["Deterministic Palette Temporal Annotation<br/>Fixed 31-Colour Palette / Emergence Ordering / Persistent IDs"]
B --> C["Three Standardized Benchmark Tasks<br/>70/15/15 Within-Replicate Split / Protocol Definition"]
C --> D["Single-Frame Leaf Instance Segmentation<br/>Mask R-CNN / YOLO / Mask2Former / SAM2/3"]
C --> E["Time-Series Leaf Instance Tracking<br/>ByteTrack / BoTSORT / DeepSORT via Tracking-by-Detection"]
C --> F["Plant Growth Stage Classification<br/>Early(4-6) / Interm.(7-10) / Mature(11+)"]
Key Designs¶
1. Deterministic Palette Temporal Annotation: Direct Evaluation without Bipartite Matching
Standard multi-object tracking and video instance segmentation benchmarks conventionally depend on post-hoc Hungarian matching to align predicted bounding boxes or masks with ground truth across frames, introducing algorithmic overhead and ambiguity when leaf contours undergo subtle non-rigid deformations. PhenoLeaf-TS resolves this bottleneck by leveraging the biological phenomenon of centripetal leaf emergence through a deterministic sequential 31-colour palette. The very first cotyledon/leaf to emerge is deterministically assigned Colour 1 (red), the second Colour 2 (blue), and subsequent newly emerging leaves receive the next sequential colour in the predefined index. Once assigned, this colour code remains fixed across all frames throughout the plant's life cycle regardless of leaf expansion, petiole elongation, or temporary partial occlusion. Twelve trained annotators used CVAT with boundary-conservative contours, verified by four senior reviewers via frame-to-frame bootstrapping to eliminate identity swaps, thereby enabling direct colour-to-identity pixel-level comparison against model predictions without post-hoc matching.
2. Hierarchical Multi-Task Protocol: From Local Instances to Macro Phenology
To systematically examine computer vision architectures across multiple biological observation scales, the benchmark establishes three complementary evaluation tasks. Task 1 (Leaf Instance Segmentation) evaluates single-frame mask delineation using COCO-standard mAP (mAP50, mAP75) along with Symmetric Best Dice (SBD) and Dice coefficients. Task 2 (Leaf Tracking) formulates temporal sequence tracking under the tracking-by-detection paradigm, adopting CLEAR MOT (MOTA, MOTP), HOTA, IDF1, and identity switches (IDSW) to quantify trajectory continuity and identity preservation. Task 3 (Growth Stage Classification) maps global rosette development into three biological stages determined by instance leaf counts: Early (4โ6 leaves, 22.6%), Intermediate (7โ10 leaves, 39.8%), and Mature (11+ leaves, 37.6%), evaluated via Accuracy, Macro-F1, Precision, and Matthews Correlation Coefficient (MCC). All tasks utilize a strict chronological 70/15/15 within-replicate split to reflect authentic developmental progression.
3. Growth-Specific Tracking Adaptation: Two-Threshold Association for Emergent Leaves
Mainstream multi-object trackers are historically engineered for urban surveillance or autonomous driving, where rigid objects frequently enter and leave the camera view. In contrast, plant phenotyping tracking operates under an inverse kinematic regime: leaves rarely disappear, while the core challenge resides in continuous non-rigid morphological expansion, gradual petiole displacement, and the emergence of microscopic young leaves from the central apical meristem. Standard trackers relying on high-threshold detection filtering (such as DeepSORT) prematurely discard weak initial detector activations from emerging leaves, capping tracking recall at 70.9%. By contrast, ByteTrack's two-threshold hierarchical matching retains low-confidence detection proposals in a secondary association pool, successfully preserving newly emerged leaf tracks and achieving 84.1% MOTA with 91.0% IDF1.
Key Experimental Results¶
Main Results¶
The benchmark evaluates 9 instance segmentation architectures, 6 multi-object trackers (all utilizing IS-1 YOLOv11-seg detections as inputs), and 6 growth stage classifiers under standardized protocols. Table 1 summarizes the instance segmentation and multi-object tracking benchmarks on the PhenoLeaf-TS test set.
| Task Family | Model / ID | Backbone | mAP / MOTA (โ) | mAP50 / MOTP (โ) | mAP75 / IDF1 (โ) | IDSW (โ) / Precision (โ) | Recall (โ) |
|---|---|---|---|---|---|---|---|
| Seg. (Single-stage) | IS-1: YOLOv11-seg | CSPDarknet-X | 68.2 | 84.7 | โ | โ | โ |
| Seg. (Single-stage) | IS-2: YOLO26-seg | CSPDarknet-L | 67.3 | 84.3 | โ | โ | โ |
| Seg. (Two-stage) | IS-3: Mask R-CNN | R50-FPN | 73.2 | 94.3 | 82.9 | โ | โ |
| Seg. (Two-stage) | IS-4: Mask R-CNN | R101-FPN | 72.4 | 93.4 | 81.6 | โ | โ |
| Seg. (Two-stage) | IS-5: Mask R-CNN | X101-FPN | 71.3 | 91.4 | 80.4 | โ | โ |
| Seg. (Two-stage) | IS-6: Cascade Mask R-CNN | R50-FPN | 71.9 | 91.5 | 81.8 | โ | โ |
| Seg. (Transformer) | IS-8: Mask2Former | Swin-L | 71.9 | 94.6 | 82.7 | โ | โ |
| Seg. (Foundation) | IS-7: SAM2 (AMG) | Hiera-L | 55.0 | 66.7 | 63.2 | โ | โ |
| Seg. (Foundation) | IS-9: SAM3 (Prompted) | Hiera-L | 71.1 | 87.9 | 81.2 | โ | โ |
| Tracking (MOT) | TR-5: ByteTrack | YOLOv11 Detections | 84.1% | 29.8 | 91.0% | 590 | 99.5% |
| Tracking (MOT) | TR-1: BoTSORT | YOLOv11 Detections | 83.1% | 27.1 | 90.3% | 692 | 99.5% |
| Tracking (MOT) | TR-2: DeepSORT | YOLOv11 Detections | 68.4% | 33.7 | 81.7% | 43 | 96.7% |
| Tracking (MOT) | TR-3: Deep-OC-SORT | YOLOv11 Detections | 66.1% | 23.4 | 79.5% | 417 | 99.6% |
| Tracking (MOT) | TR-6: StrongSORT | YOLOv11 Detections | 64.2% | 24.2 | 78.0% | 345 | 99.8% |
| Tracking (MOT) | TR-4: Norfair | YOLOv11 Detections | 57.8% | โ | 73.4% | 71 | 99.3% |
In Growth Stage Classification, Swin-T (CL-6) leads with 91.7% accuracy, 91.7% Macro-F1, and 0.873 MCC, followed closely by EfficientNetV2-S (CL-3) at 91.4% accuracy and 0.869 MCC. Notably, DINOv2-B (CL-5) with a frozen backbone and linear probe achieves 87.1% accuracy and 0.803 MCC without task-specific feature tuning.
Ablation Study¶
To evaluate PhenoLeaf-TS as an effective pre-training source for general plant phenotyping, the authors conducted zero-shot transfer and fine-tuning experiments across intra-species benchmarks (Ara2012, Ara2013) and an external cross-species benchmark (Komatsuna).
| Target Dataset | Task Type | Model | Zero-shot | Fine-Tuned (FT) | Gain (\(\Delta\)) | Note |
|---|---|---|---|---|---|---|
| Ara2012 (Arabidopsis) | Seg. (mAP) | Mask R-CNN (IS-3) | 55.3 | 73.9 | +18.6 | Intra-species transfer; fine-tuning converges to strong precision |
| Ara2012 (Arabidopsis) | Seg. (mAP) | Mask2Former (IS-8) | 59.9 | 71.1 | +11.2 | Attention decoders yield stronger un-adapted boundary transfer |
| Ara2012 (Arabidopsis) | Cls. (Acc%) | EfficientNetV2 (CL-3) | 100.0 | 100.0 | 0.0 | All 120 images are mature rosettes; collapses to single-class check |
| Ara2013 (Arabidopsis) | Seg. (mAP) | Mask R-CNN (IS-3) | 53.5 | 75.2 | +21.7 | Substantial adaptation gains on small benchmark splits |
| Ara2013 (Arabidopsis) | Seg. (mAP) | Mask2Former (IS-8) | 51.3 | 66.3 | +15.0 | Consistently demonstrates pre-training value across architectures |
| Ara2013 (Arabidopsis) | Cls. (Acc%) | EfficientNetV2 (CL-3) | 21.0 | 67.0 | +46.0 | Captures developmental spread across multi-stage distribution |
| Ara2013 (Arabidopsis) | Cls. (Acc%) | Swin-T (CL-6) | 21.0 | 78.0 | +57.0 | Hierarchical vision transformer exhibits strongest phenotypic transfer |
| Komatsuna (Komatsuna) | Seg. (mAP) | Mask R-CNN (IS-3) | 34.7 | 86.1 | +51.4 | Massive cross-species performance breakthrough via fine-tuning |
| Komatsuna (Komatsuna) | Seg. (mAP) | Mask2Former (IS-8) | 47.1 | 76.7 | +29.6 | Robust zero-shot generalization against radical morphology shifts |
Key Findings¶
- Backbone Inverted Scaling: Across two-stage instance segmentation, scaling model parameters inversely impacts performance: ResNet-50 (73.2 mAP) > ResNet-101 (72.4 mAP) > ResNeXt-101 (71.3 mAP). In controlled indoor phenotyping with 17,082 images, phenotypic diversity rather than network capacity forms the primary bottleneck; excessive capacity induces slight overfitting to uniform background soil textures.
- The Domain Gap of Foundation Vision Models: Zero-shot SAM2 with unprompted automatic mask generation achieves only 55.0 mAP due to severe over-segmentation on background soil granules and pot edges. Guiding SAM3 with Grounding DINO text prompts dramatically boosts performance to 71.1 mAP (+16.1 points), demonstrating that open-world foundation models require explicit domain prompt constraints for agricultural tasks.
- Appearance vs. Association Trade-off in Growth Tracking: DeepSORT produces only 43 identity switches (IDSW) owing to deep appearance embeddings that stably identify leaves under expansion; however, its strict matching thresholds cause heavy false negatives on emerging leaves (70.9% recall). In contrast, ByteTrack prioritizes low-threshold detection recovery over appearance re-identification, attaining a superior 85.9% recall and 84.1% MOTA.
- Substantial Cross-Species Transferability: Fine-tuning PhenoLeaf-TS pre-trained weights on Komatsuna yields an exceptional +51.4 mAP gain for Mask R-CNN (from 34.7 to 86.1 mAP), confirming that the learned representations capture universal leaf boundary topology and geometric primitives that generalize across botanical families.
Highlights & Insights¶
- Persistent 31-Colour Palette Eliminates Bipartite Matching: Linking physical leaf emergence directly to an immutable palette converts temporal instance consistency evaluation into straightforward pixel-level colour verification, eliminating algorithmic ambiguity in downstream benchmarks.
- Unveiling Inverse Kinematics in Biological Tracking: Demonstrates that tracking biological growth fundamentally contradicts traditional pedestrian tracking assumptions; retaining low-confidence candidate detections in secondary association pools is far more impactful than complex non-rigid appearance modeling.
- Efficiency and Deployment Implications for Edge Phenotyping: Demonstrates that compact backbones (ResNet-50) and lightweight classifiers (EfficientNetV2-S, Swin-T) match or exceed heavyweight foundation models in controlled agricultural screening, providing clear architectural guidance for embedded robotic phenotyping platforms.
Limitations & Future Work¶
- Domain Confinement to Controlled Indoor Setups: All imagery originates from indoor climate-controlled growth chambers with static top-down perspectives and soil backgrounds, leaving outdoor agricultural challengesโsuch as variable sunlight, weed entanglements, and weather dynamicsโunaddressed.
- Boundary Ambiguity in Discrete Phenological Stages: Discretizing developmental progress strictly by leaf counts (thresholds at 6 and 10 leaves) introduces ambiguity near transitional boundaries, suggesting that future iterations should incorporate continuous growth age regression and parametric biomass modeling.
- Mismatch Between Standard MOT Metrics and Biological Growth: Classical CLEAR MOT metrics penalize natural temporal delays in detecting microscopic newly emerged leaves and fail to reward monotonic leaf area growth curves; developing specialized botanical tracking metrics remains an open priority.
Related Work & Insights¶
- vs. CVPPP / LCC Benchmarks: CVPPP provides high-quality static instance masks but lacks temporal continuity. PhenoLeaf-TS expands the data scale by two orders of magnitude (17k+ vs. hundreds of images) and introduces tracking and growth stage classification dimensions.
- vs. Komatsuna Dataset: While Komatsuna introduced temporal tracking annotations, it is limited to 15 sequences and 900 images of one species. PhenoLeaf-TS spans 21 genotypes and 318 sequences, facilitating deep pre-training and comprehensive multi-task benchmarking.
- vs. Foundation Models (SAM2 / SAM3): Directly deploying zero-shot vision foundation models generates extensive soil over-segmentation. PhenoLeaf-TS highlights the necessity of task-specific grounding prompts or direct pre-training on botanical benchmarks.
Rating¶
- Novelty: โญโญโญโญโ First large-scale, multi-genotype, temporally continuous plant phenotyping benchmark featuring persistent colour-coded instance identity.
- Experimental Thoroughness: โญโญโญโญโญ Rigorously evaluates 21 baseline architectures across three distinct vision tasks alongside extensive cross-dataset transfer experiments.
- Writing Quality: โญโญโญโญโญ Exceptionally clear task formulation, comprehensive figures, and insightful empirical analysis of botanical tracking dynamics.
- Value: โญโญโญโญโญ Resolves a longstanding public data gap for continuous plant phenotyping, offering immediate practical utility for both academic vision research and precision agricultural automation.