On-Orbit Real-Time Wildfire Detection Under On-Board Constraints¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Remote Sensing
Keywords: wildfire detection, Edge AI, DenseMAE, uncalibrated infrared, lightweight segmentation
TL;DR¶
Addressing the extreme computational bottlenecks of LEO thermal microsatellites and uncalibrated single-band mid-wave infrared data, this paper introduces DenseMAEโa lightweight fully convolutional masked autoencoder paired with a high-resolution refinement headโachieving a 0.699 pixel-level full-distribution AP and 0.744 event-level Fire-F1 with a sub-megabyte footprint and 65 ms TensorRT inference.
Background & Motivation¶
Wildfires are intensifying globally in frequency, severity, and duration, posing severe threats to ecosystems, infrastructure, and human life while contributing substantially to carbon emissions. Early and rapid detection of incipient ignitions is paramount for operational containment and mitigation. However, conventional spaceborne surveillance architectures suffer from an inherent spatio-temporal trade-off: geostationary meteorological satellites offer high temporal revisit but coarse spatial resolution (2โ3 km) that misses small early fires, whereas low Earth orbit (LEO) satellites provide high-resolution imaging but typically only a few daily passes, frequently missing the critical afternoon peak in fire activity. To bridge this divide, commercial thermal microsatellite constellations (such as OroraTech's OTC-P1 mission and precursor satellite FOREST-2) deploy uncooled microbolometers providing 200 m ground sampling distance, enabling high-revisit monitoring during the peak fire window and targeting an end-to-end alert pipeline under 10 minutes from overpass to communication. This strict latency ceiling completely rules out traditional downlink-and-process architectures, mandating real-time autonomous inference directly on board the satellite.
Yet, on-orbit wildfire detection operates under severe compounded constraints. On the hardware side, edge compute nodes are restricted to low-power modes (NVIDIA Jetson Xavier NX running in 10 W mode, with strictly limited compute throughput and memory bandwidth), imposing a hard budget of sub-megabyte engine sizes and sub-150 ms per-batch inference latencies. On the sensor side, to eliminate the latency and compute overhead of radiometric calibration, the pipeline must ingest uncalibrated single-channel mid-wave infrared (MWIR, 3.8 ยตm) digital numbers (DN), rendering it vulnerable to sensor striping noise, non-uniformity, and sunglint specular reflections. Crucially, early-stage fires manifest as isolated sub-pixel or single-pixel thermal anomalies amidst vast noisy backgrounds, resulting in extreme class imbalance where fire pixels comprise a mere 0.0182% of valid pixels and labeled training scenes are severely limited.
Faced with noisy single-band inputs, extreme imbalance, and sub-megabyte compute limits, the paper investigates lightweight dense self-supervised representation learning tailored to the on-orbit thermal regime. Core idea: develop DenseMAE, a lightweight staged convolutional masked autoencoder that learns robust spatial thermal representations directly on uncalibrated MWIR imagery while masking out sensor artifacts, coupled with a high-resolution skip refinement head to achieve sub-megabyte, 65 ms on-orbit wildfire segmentation without requiring pruning or compression.
Method¶
Overall Architecture¶
The system processes uncalibrated single-band MWIR tiles of size 224ร224 (robustly scaled across valid pixels while masking invalid sensor pixels). The workflow is divided into two distinct phases: in the self-supervised pretraining phase, a fully convolutional DenseMAE encoder and a lightweight reconstruction decoder are trained on unlabeled MWIR imagery using block-masked L1 reconstruction to learn context-aware spatial representations; in the downstream deployment phase, the reconstruction decoder is discarded, and the DenseMAE encoder is coupled with a lightweight TensorRT-optimized head and an high-resolution refinement head that re-injects high-resolution features from the shallow stem to predict the final fire segmentation mask.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Uncalibrated MWIR Tile<br/>1ร224ร224 Raw DN"] --> B["Stem Shallow High-Res Conv<br/>HรW / 32 ch (Outputs fโ)"]
B --> C["DenseMAE Staged Lightweight Encoder<br/>Single Stride-2 Conv + Dilated Conv"]
B -.->|"High-Res Skip fโ"| E["High-Resolution Refinement Head<br/>HR Refinement Module"]
C --> D["Dense Embedding Map z<br/>DรH/2รW/2"]
D --> E
E --> F["Predicted Wildfire Mask<br/>1ร224ร224"]
Key Designs¶
1. DenseMAE Staged Convolutional Masked Autoencoder: Lightweight SSL Tailored to Edge MWIR
Standard Vision Transformer-based MAE models divide images into non-overlapping patch tokens and execute compute-heavy self-attention, incurring excessive memory and latency overheads that hinder deployment on embedded edge GPUs. DenseMAE adopts a fully convolutional staged design: a 3ร3 convolutional stem first extracts shallow spatial features \(f_0\) at resolution \(H \times W\) with 32 channels; Stage 1 then applies a single stride-2 3ร3 convolution to downsample spatial features to \(H/2 \times W/2\); subsequent Stage 2 and Stage 3 operate strictly at this \(H/2 \times W/2\) resolution using dilated convolutions with dilation rates \(d \in \{1, 2, 4\}\) to expand the receptive field without further resolution loss, concluding with a 1ร1 projection layer to produce dense embeddings \(z\) of dimension \(D\) (\(D \in \{32, 64\}\)). GroupNorm (8 groups) and GELU activations are used throughout for numerical stability under small batch sizes. During pretraining, random spatial block masking with ratio \(r\) and block size \(b\) is applied, and sensor artifacts/blind pixels are strictly excluded using a valid-pixel mask \(V\). The L1 reconstruction loss is evaluated exclusively over masked, valid pixels against the area-downsampled input \(x_{\downarrow 2}\):
2. High-Resolution Skip Refinement Head: Sub-Pixel Fire Localization Under Coarse Grid Embeddings
Operating predominantly at the downsampled \(H/2 \times W/2\) grid drastically cuts computational complexity, but risks losing sharp structural boundaries and edge details critical for detecting isolated single-pixel or sub-pixel thermal anomalies. To resolve this tension, the downstream architecture employs an asymmetric two-tier prediction strategy: coarse logits \(L_c\) are first generated from the dense embedding \(z\) by a lightweight head (such as TRT-Head or DW-Res Head); subsequently, a shallow refinement module \(\phi\) concatenates the coarse logits with the full-resolution stem feature map \(f_0\) to predict refined final logits:
By re-injecting uncompressed high-frequency features from the stem at negligible compute cost (<0.1 MB parameter overhead), this skip refinement preserves sharp spatial localization for weak point-source fires.
3. Operator-Fused Lightweight Downstream Heads: Pushing Sub-Megabyte TensorRT Inference Boundaries
To maximize kernel fusion and memory efficiency on the 10 W Jetson Xavier NX, downstream segmentation heads avoid complex branching and dynamic tensor operations. Two standardized heads are established: TRT-Head is an ultra-lightweight module composed of a 1ร1 projection, a stack of three 3ร3 Conv-BatchNorm-SiLU blocks, and a 1ร1 output layer (hidden width 32), compiling to a tiny 0.52 MB TensorRT engine with 65.34 ms batch inference latency; DW-Res Head introduces depthwise-separable residual blocks with dilated convolutions to expand spatial modeling capacity, achieving superior accuracy while retaining an engine footprint below 1.0 MB.
Loss & Training¶
To combat extreme class imbalance (fire pixels accounting for only 0.0182% of valid pixels), training batches are biased via positive patch oversampling to ensure representative fire patterns in each step. Loss evaluation uses masked binary cross-entropy (Masked BCE) that excludes invalid and no-data pixels via \(V\). Optimization uses AdamW with mixed-precision training and cosine annealing learning rate schedules. SSL checkpoints are initially screened via linear probe AP on a static, class-balanced embedding pool, while all headline results are evaluated on the complete test split under true operational prevalence alongside connected-component event-level Fire-F1.
Key Experimental Results¶
Main Results¶
Main experiments are evaluated across 839 annotated scenes collected over two years across 9 satellites (FOREST-2 and FOREST-4 through 11). The test set comprises 9,592 tiles evaluated at operational prevalence using pixel-level AP and validation-tuned event-level Fire-F1. Latency and engine footprint are measured on an NVIDIA Jetson Xavier NX using TensorRT FP16 (\(224 \times 224\), Batch \(B=8\)).
| Model Config | Pretraining / Transfer Paradigm | Test AP โ | Event Fire-F1 โ | TRT FP16 Latency (ms) โ | Engine Size (MB) โ |
|---|---|---|---|---|---|
| HR-U-Net++ (depth=2, h=32) | Supervised end-to-end | 0.610 | 0.680 | 94.25 | 0.77 |
| HR-U-Net++ (depth=3, h=32) | Supervised end-to-end | 0.635 | 0.690 | 97.53 | 1.62 |
| HR-U-Net++ (depth=4, h=32) | Supervised end-to-end | 0.650 | 0.730 | 104.34 | 2.10 |
| DenseMAE (emb32) + TRT-Head | SSL Pretrained + Fine-tuned | 0.640 | 0.690 | 65.34 | 0.52 |
| DenseMAE (emb32) + DW-Res + HR | SSL Pretrained + Fine-tuned | 0.677 | 0.732 | 98.01 | 1.10 |
| DenseMAE (emb64) + TRT-Head | SSL Pretrained + Fine-tuned | 0.677 | 0.689 | 110.46 | 0.55 |
| DenseMAE (emb64) + DW-Res + HR | SSL Pretrained + Fine-tuned | 0.699 | 0.744 | 141.00 | 0.91 |
Ablation Study¶
Ablations investigate transfer modes, embedding dimensionality, and self-supervised training variants.
| Config | Strategy & Transfer Setting | Full Test AP โ | Note |
|---|---|---|---|
| DenseMAE (emb64) + HR Head | DenseMAE Pretraining + Full Fine-tuning | 0.689 ยฑ 0.004 | Best fine-tuned performance across folds |
| DenseMAE (emb32) + HR Head | DenseMAE Pretraining + Full Fine-tuning | 0.671 ยฑ 0.005 | Compact representation, lower latency |
| DenseMAE (emb64) + HR Head | DenseMAE Pretraining + Frozen Encoder (Probe) | 0.612 ยฑ 0.006 | Head-only adaptation confirms strong representation |
| DenseMAE (emb64) + HR Head | Random Init + Full Fine-tuning (From Scratch) | 0.551 ยฑ 0.008 | Severe 13.8% AP drop confirms necessity of SSL |
| DenseMAE + Hybrid EMA Distill (\(\lambda=0.02\)) | Masked Reconstruction + EMA Teacher Consistency | 0.2796 (Probe) | Marginal gains; masked distillation hurts full-stream AP |
Key Findings¶
- SSL pretraining provides essential regularization under extreme imbalance: Under identical model architectures, DenseMAE pretraining provides a massive +13.8 percentage point AP gain over training from scratch with random initialization (0.689 vs. 0.551), confirming that self-supervised masked reconstruction builds strong resilience against sensor noise and background thermal variations in low-data regimes.
- High-resolution skip refinement is decisive for small fire detection: While pixel-level AP can be dominated by large burning perimeters, adding high-resolution skip refinement yields a pronounced jump in event-level Fire-F1 (rising from 0.689 to 0.744 in emb64 variants), demonstrating that re-injecting shallow full-resolution features crucially recovers sub-pixel and isolated single-pixel fire signals.
- Pure masked reconstruction outperforms hybrid EMA distillation in noisy MWIR: Diagnostic ablations show that adding an EMA momentum distillation loss fails to yield consistent gains on full-stream evaluation, and restricting distillation to masked areas degrades performance. For single-band noisy infrared imagery, unadulterated L1 masked reconstruction over valid context provides the most robust learning signal.
- Cross-sensor validation against operational VIIRS: In a coincident overpass evaluation (\(\pm 10\) min window, 800 m buffer clustering) against the operational VIIRS 375 m active fire product, the model achieves 103 joint detection clusters and successfully detects small early fires below the VIIRS spatial detection threshold (200 m vs. 375 m GSD).
Highlights & Insights¶
- Adapting masked autoencoding to a fully convolutional edge regime: Unlike standard transformer MAE architectures, DenseMAE tailors single early downsampling and dilated convolutions to remote sensing edge constraints, maintaining spatial grid correspondence at \(H/2 \times W/2\) and maximizing GPU cache locality on embedded devices.
- Compression-free sub-megabyte deployment Pareto frontier: Without resorting to post-training quantization (INT8) or structured pruning, careful operator selection and channel budgeting allow DenseMAE engines compiled directly to TensorRT FP16 to achieve 0.52โ0.91 MB sizes and sub-100 ms latencies, defining a new Pareto frontier for spaceborne edge AI.
- Physics-grounded false positive attribution via specular geometry: By analyzing the dot product \(N \cdot H\) (surface normal and sun-satellite halfway vector), the authors show that 67% of event false positives are concentrated in near-ideal specular reflection geometries (\(N \cdot H \approx 1.0\)), pinpointing sunglint as the primary single-band failure mode and charting clear paths for geometric filtering.
Limitations & Future Work¶
- Physical ill-posedness of single-band inputs under specular glint: Lacking complementary long-wave infrared (LWIR) or visible channels, single-band MWIR cannot fundamentally disentangle extreme specular sunglint (e.g. from water bodies or solar arrays) from true high-temperature combustion solely via spatial patterns. Future work requires multi-band ratio tests or integrating orbital illumination geometry.
- Absence of multi-temporal observation context: Each satellite overpass is currently processed independently in isolation, precluding the use of multi-temporal heating curves to filter static hot spots. Strict memory and inter-process synchronization constraints on the Jetson Xavier NX currently prevent maintaining persistent frame buffers across orbits.
Related Work & Insights¶
- vs MODIS (MOD14) / VIIRS (375m) Operational Contextual Pipelines: Conventional operational fire pipelines require rigorous absolute radiometric calibration and rely heavily on MWIR-TIR dual-band brightness temperature contrast with fixed decision rules. In contrast, DenseMAE operates directly on raw uncalibrated single-band DNs, learning end-to-end contextual anomaly representations that match or exceed operational sensitivity at 200 m resolution.
- vs Standard Deep Learning Fire Segmentation (e.g., Fire-Net, Smoke-Unet): Existing deep learning fire detection models assume ground-based server deployments, ingest multi-spectral calibrated inputs, and utilize models with tens of megabytes of parameters. DenseMAE demonstrates that a sub-megabyte fully convolutional SSL model running under 10 W power constraints can match and outperform heavy supervised baselines on orbit.
Rating¶
- Novelty: โญโญโญโญ [Well-crafted fully convolutional DenseMAE tailored to uncalibrated single-band MWIR and edge satellite constraints]
- Experimental Thoroughness: โญโญโญโญโญ [Rigorous two-year multi-satellite dataset across 9 spacecraft, full-distribution pixel AP under extreme imbalance, event-level Fire-F1, on-device TensorRT profiling, and coincident VIIRS cross-sensor validation]
- Writing Quality: โญโญโญโญโญ [Clear articulation of edge constraints, transparent ablation analysis, and honest treatment of single-band limitations]
- Value: โญโญโญโญโญ [Directly deployed aboard commercial LEO constellation, demonstrating real-world viability of sub-10-minute on-orbit wildfire alerting]