Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://humansensinglab.github.io/AVODDiag/
Area: Object Detection
Keywords: Aerial Object Detection, Foundation Generative Models, Model Diagnosis, Attribute-Controlled Generation, Targeted Data Supplementation
TL;DR¶
This paper introduces a synthetic diagnostic framework powered by state-of-the-art foundation image generative models (Imagen 3) that couples an explicit attribute taxonomy, text-guided synthesis, complementary image editing, and human-in-the-loop verification to pinpoint detector vulnerabilities across scenes, enabling targeted supplementation with minimal real data to achieve up to 13.4% AP50 gains.
Background & Motivation¶
Aerial-view object detection (AVOD) serves as a cornerstone of modern Earth observation (EO) pipelines, underpinning critical tasks ranging from disaster response and infrastructure inspection to defense intelligence and urban planning. In these high-stakes deployment scenarios, visual detection systems must operate with high reliability, as systematic performance degradation under unfamiliar or complex environmental conditions can lead to catastrophic downstream decisions. However, widely used aerial benchmarks such as DOTAv2, LINZ, and UGRC exhibit severe intrinsic distribution imbalances: standard daytime scenes with clear skies and suburban layouts dominate the distributions, while complex industrial facilities, dense urban cores, and extreme seasonal weather remain severely underrepresented. Consequently, standard aggregate evaluation metrics (such as dataset-wide mAP) gloss over localized, systematic failure modes.
Traditional diagnostic methodologies for object detection suffer from an inability to conduct controlled causal intervention. Post-hoc diagnostic tools like TIDE can classify detection errors into localization mistakes, duplicate detections, or background confusions, and attribution methods like Grad-CAM can highlight spatial attention regions, but neither can disentangle tightly coupled environmental variables such as background texture, illumination, shadow casting, and object scale in real imagery. Concurrently, collecting geographically and environmentally balanced real-world aerial imagery with fine-grained control is cost-prohibitive and practically infeasible. This leaves researchers without a controlled stress-testing benchmark to measure model sensitivity along isolated attribute axes.
To break this bottleneck, this paper reframes the role of generative AI: foundation image generative models should not merely serve as generic data augmentation tools, but rather function as high-fidelity diagnostic probes for trained vision systems. Core idea: construct an attribute-controlled synthetic diagnostic testbed via foundation generative models to disentangle environmental and object factors, identify scene-level detector vulnerabilities, and guide targeted, highly data-efficient real-world data supplementation.
Method¶
Overall Architecture¶
The proposed diagnostic and targeted repair pipeline consists of four cohesive stages: first, defining a structured taxonomy tree that decouples macroscopic environmental variables from microscopic object attributes; second, generating nadir aerial images via text-to-image foundation models conditioned on uniform attribute prompts, followed by MLLM-based semantic validation; third, executing image-and-text-to-image editing on underrepresented categories to mitigate generative distribution biases; and finally, curating zero-shot VLM proposals with rapid human-in-the-loop binary verification to establish a synthetic benchmark for scene-conditioned model diagnosis and targeted real-world supplementation.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Aerial diagnostic requirements and evaluation goals"] --> B["Hierarchical attribute taxonomy & prompt-driven synthesis<br/>Decouple environmental & object factors; generate nadir scenes"]
B --> C["Multimodal feedback loop & inverse-distribution image editing<br/>MLLM attribute validation; targeted text editing for long-tail balance"]
C --> D["Multi-model zero-shot ensemble & rapid human-in-the-loop annotation<br/>Ensemble lightweight VLM proposals; binary human verification"]
D --> E["Attribute-conditioned diagnosis & targeted real-data supplementation<br/>Slice-based performance probing; supplement real data to fix weak categories"]
E --> F["Output: Robust aerial detector & comprehensive diagnostic report"]
Key Designs¶
1. Hierarchical attribute taxonomy & prompt-driven synthesis: decoupling environmental and object factors for controlled probing In real aerial imagery, environmental variables are heavily entangled, preventing isolated sensitivity analysis of individual factors. The framework establishes a two-level semantic taxonomy tree \(\mathcal{T}\) divided into Primary Environmental Attributes and Secondary Object Attributes. Primary attributes govern the scene globally across scene type (dense, desert, forest, industrial, plain, residential, rural, scrubland, and urban), season (spring, summer, fall, winter), and weather (sunny, cloudy, rainy, foggy, snowy) to simulate background clutter, illumination, and seasonal variations. Secondary attributes specify object-level details including vehicle count, vehicle color, and vehicle type to test foreground contrast and spatial crowding. In the initial synthesis stage, attribute tuples \(a_i = \{a_i^{(c)}\}_{c \in \mathcal{T}}\) are uniformly sampled from \(\mathcal{T}\). A large language model \(f_{\text{LLM}}^{(G)}\) (GPT-5) translates each tuple into an aerial nadir-view prompt \(p_i^{(G)}\), which directs a text-guided diffusion model \(f_G\) (Imagen 3) to synthesize a realistic aerial image \(x_i\) at aligned resolution.
2. Multimodal feedback loop & inverse-distribution image editing: correcting generation defects and mitigating attribute bias Text-guided diffusion models often deviate from input prompts due to generation defects and implicit training priors, resulting in an empirical distribution that diverges from the intended uniform sampling. To restore distributional fidelity, the framework incorporates an automated multimodal validation and editing loop. A multimodal large language model \(f_{\text{MLLM}}\) (Gemini 2.5 Flash) inspects the generated image \(x_i\) to extract the observed attribute value \(\tilde{a}_i^{(c)}\). A text encoder \(f_{\text{TE}}\) (Sentence-BERT) computes cosine similarity \(S_i^{(c)}\) against text embeddings of valid candidate categories \(C(c)\) in \(\mathcal{T}\): $\(k_{\max} = \arg\max_k S_i^{(c)}[k], \quad \hat{a}_i^{(c)} = \begin{cases} C(c)[k_{\max}], & \text{if } S_i^{(c)}[k_{\max}] \ge \tau \\ \tilde{a}_i^{(c)}, & \text{otherwise} \end{cases}\)$ If the highest similarity falls below threshold \(\tau\), the observed attribute is dynamically appended to \(\mathcal{T}\) to preserve semantic taxonomy completeness. To correct remaining distributional imbalances, an image-and-text-to-image editing stage samples underrepresented categories from a complementary distribution \(P^\Theta(C(c))\). An editing LLM \(f_{\text{LLM}}^{(E)}\) (Gemini 2.5 Flash Lite) generates an instruction that strictly modifies only the targeted attribute while maintaining invariant scene semantics. This produces an enriched set \(I_E\), yielding a balanced diagnostic benchmark of 5,453 synthetic images spanning 304 unique environmental attribute combinations (\(I = I_G \cup I_E\)).
3. Multi-model zero-shot ensemble & rapid human-in-the-loop annotation: ensuring unbiased high-throughput ground truth Generative models do not natively output bounding boxes, and deploying detectors pretrained on real aerial data for pseudo-labeling would inevitably re-inject real-world dataset biases into the diagnostic benchmark. The authors circumvent this by employing foundation vision-language models for zero-shot vehicle proposal generation. Predictions from Gemini 2.5 Flash Lite and Moondream 2 are ensembled and overlapping bounding boxes are averaged to suppress duplicates while retaining high recall. To eliminate false-positive hallucinations inherent to zero-shot detection on tiny remote sensing targets, human annotators perform rapid binary verification ("keep or discard"). Compared to manual bounding box drawing (~26 seconds per object), binary decision verification requires only ~1 second per box, achieving an approximately 13-fold annotation speedup while securing high ground-truth quality. On the synthetic dataset, 31,519 vehicle bounding boxes were verified and approved (an approval rate of 58.5%).
4. Attribute-conditioned diagnosis & targeted real-data supplementation: pinpointing failure modes to guide efficient data collection Trained aerial detectors are deployed on the synthetic benchmark and evaluated independently across scene-type slices, yielding scene-conditioned average precision \(AP_{50}^{\text{scene}}\). Scene types exhibiting an AP drop exceeding 1.0% relative to the overall dataset average \(\overline{AP}_{50}\) are flagged as systematic model vulnerabilities. The diagnostic probing reveals that urban scenes are problematic across all 9 detector-dataset configurations, while industrial scenes fail in 7 of 9 configurations. Guided by these diagnostic insights, the authors curate compact, real single-environment datasets (Miami urban with 2,284 images, Los Angeles industrial with 2,000 images, and Phoenix desert with 2,000 images)โsizes that are an order of magnitude smaller than original training corpora (e.g., UGRC's 147k images). Supplementing models with these targeted real samples achieves substantial performance recovery while preventing the negative transfer and domain interference caused by unguided random data scaling.
Key Experimental Results¶
Main Results¶
The experiments benchmark three representative detector paradigmsโtwo-stage (Faster R-CNN R50), single-stage anchor-free (YOLOv8-M), and vision transformer-based (ViTDet-B)โpretrained on three standard aerial datasets (DOTAv2, LINZ, UGRC). Models are diagnosed on the synthetic testbed and retrained with targeted real data subsets (Miami urban, LA industrial, Phoenix desert). The table below summarizes dataset-level AP50 and the resulting performance gains.
| Initial Training Set | Detector Model | Supplemented Categories | Baseline AP50 (%) | With Suppl. AP50 (%) | Absolute Gain (%) |
|---|---|---|---|---|---|
| DOTAv2 (28k imgs) | Faster R-CNN (R50) | Urban + Industrial | 93.8 | 94.7 | +0.9 |
| DOTAv2 (28k imgs) | YOLOv8-M | Urban + Industrial | 87.2 | 92.3 | +5.1 |
| DOTAv2 (28k imgs) | ViTDet-B | Urban | 86.3 | 95.9 | +9.6 |
| LINZ (87k imgs) | Faster R-CNN (R50) | Urban + Industrial | 87.7 | 95.1 | +7.4 |
| LINZ (87k imgs) | YOLOv8-M | Urban + Industrial + Desert | 92.7 | 95.8 | +3.1 |
| LINZ (87k imgs) | ViTDet-B | Urban | 91.5 | 96.1 | +4.6 |
| UGRC (147k imgs) | Faster R-CNN (R50) | Urban + Industrial + Desert | 80.9 | 93.5 | +12.6 |
| UGRC (147k imgs) | YOLOv8-M | Urban + Industrial | 94.2 | 95.9 | +1.7 |
| UGRC (147k imgs) | ViTDet-B | Urban + Industrial | 81.6 | 95.0 | +13.4 |
Ablation Study: Targeted Supplementation vs. Controlled Budget Random Supplementation (DOTAv2)¶
To rigorously evaluate the diagnostic utility, targeted supplementation is benchmarked against random aerial sampling (random LINZ images) and random natural imagery (COCO) under an identical image-budget constraint.
| Detector Model | Evaluated Scene Slice | Targeted Real Suppl. Gain (%) | Random Aerial Suppl. (LINZ) Gain (%) | Random Natural Suppl. (COCO) Gain (%) | Diagnostic Advantage |
|---|---|---|---|---|---|
| Faster R-CNN | Urban | -0.0 | -1.1 | -4.2 | Prevents performance regression |
| Faster R-CNN | Industrial | +2.5 | +1.0 | -3.3 | Outperforms random baseline |
| Faster R-CNN | Forest | +0.3 | -3.7 | -6.8 | Suppresses domain interference |
| YOLOv8-M | Urban | +11.7 | -0.2 | -3.3 | Resolves primary weakness |
| YOLOv8-M | Industrial | +10.0 | -0.4 | +4.1 | Maximizes targeted gain |
| YOLOv8-M | Plain | +1.5 | -0.0 | -9.3 | Avoids severe negative transfer |
| ViTDet-B | Urban | +10.8 | +8.7 | +8.8 | Delivers superior peak accuracy |
| ViTDet-B | Industrial | +6.6 | +3.4 | +4.3 | Dominates across weak scenes |
| ViTDet-B | Desert | +5.6 | +3.0 | +3.5 | Consistently ahead of random |
Key Findings¶
- Urban and industrial scenes represent pervasive Achilles' heels across detectors: Under scene-conditioned synthetic evaluation, urban scenes fell below dataset-wide performance in all 9 configurations, and industrial environments degraded in 7 of 9 configurations (e.g., UGRC-trained Faster R-CNN achieved only 65.4% AP50 on industrial scenes and 73.7% on urban scenes, compared to its 83.8% aggregate mean). Dense vehicle crowding, complex artificial structural shadows, and low foreground-background contrast are key factors driving these failures.
- Synthetic diagnostic predictions transfer faithfully to real domains: Injecting a compact targeted set of ~2,000 real industrial or urban images led to drastic performance jumps on weak categoriesโsuch as a +24.1% AP50 surge in industrial scenes and +15.0% in urban scenes for UGRC-trained Faster R-CNN. This demonstrates that synthetic probing reliably isolates authentic real-world representation deficits.
- Blind data scaling incurs severe risks of negative transfer: As demonstrated in Figure 9, injecting randomly sampled images under identical data budgets (even from high-grade aerial imagery like LINZ) causes notable accuracy drops on conventional CNN architectures (Faster R-CNN and YOLOv8 drop by 2% to 9% AP across various scenes). Non-targeted additions introduce feature interference and redundant representations. While Vision Transformer detectors (ViTDet) exhibit better resilience to random shifts, targeted supplementation still achieves markedly higher peak performance gains.
Highlights & Insights¶
- Repositioning generative AI from data generator to diagnostic probe: While previous efforts predominantly utilize diffusion models for indiscriminate training data synthesis, this work pioneers their deployment as controlled, counterfactual diagnostic instruments capable of stress-testing trained visual models along isolated semantic axes.
- Counteracting generative long-tail bias with inverse-distribution editing: Generative models naturally harbor mode bias toward common, visually straightforward scenes. The proposed LLM/MLLM-directed editing pipeline systematically rebalances attribute distributions, ensuring fair and rigorous benchmark statistics.
- The "diagnose-then-target" closed loop transforms operational data curation: Collecting exhaustive, multi-environment Earth observation data at scale is economically prohibitive. This study formalizes an efficient paradigm: synthesize controlled probes to uncover blind spots \(\rightarrow\) acquire a modest batch (~2k) of targeted real samples \(\rightarrow\) eliminate double-digit performance deficits.
Limitations & Future Work¶
- Reliance on proprietary foundation model APIs: The framework currently relies on commercial APIs (Imagen 3 and Gemini 2.5) for high-fidelity nadir perspective control. Open-weight models (e.g., SDXL, SD 3.5) exhibit scale inconsistencies and off-nadir tilt without domain-specific fine-tuning. Future work should transition toward fully open-weight pipelines as open remote sensing foundation models mature.
- Single-modality and nadir viewpoint constraints: The empirical validation focuses strictly on visible-spectrum (RGB) nadir-view vehicle detection, leaving oblique perspectives and non-RGB modalities (thermal infrared, multispectral, or SAR) unaddressed.
- Higher-order attribute interactions remain unexplored: While the taxonomy accommodates multiple environmental and object attributes, the current evaluation predominantly slices data by scene type. Investigating multi-factor combinatorial stresses (e.g., severe snowstorms \(\times\) micro-scale dark vehicles \(\times\) ultra-dense occlusion) warrants further dedicated modeling.
Related Work & Insights¶
- vs TIDE (Bolya et al., ECCV 2020): TIDE is an error-decomposition toolbox applied post-hoc to real test sets, constrained by the natural confounding of real images. In contrast, this framework proactively generates counterfactual test cases with isolated causal factors.
- vs DatasetDM (Wu et al., NeurIPS 2023): DatasetDM synthesizes training data with perception annotations for generic data augmentation; this work leverages generative models to diagnose hidden representation gaps, advocating for targeted real-data supplementation rather than synthetic replacement.
- vs RarePlanes (Shermeyer et al., WACV 2021): RarePlanes relies on traditional Unreal Engine simulators to generate synthetic aircraft, incurring notable texture artifacts and domain gaps. This paper leverages foundation diffusion models with superior context realism and flexible prompt controllability.
Rating¶
- Novelty: โญโญโญโญโญ Pioneering framework transforming foundation image generative models into interpretable, attribute-controlled diagnostic probes for aerial perception.
- Experimental Thoroughness: โญโญโญโญโญ Rigorous validation across 3 detector architectures, 3 major aerial benchmarks, 5,453 synthetic diagnostic images, multi-city targeted real datasets, and controlled random augmentation baselines.
- Writing Quality: โญโญโญโญโญ Cohesive narrative, precise technical formulation, and highly persuasive empirical demonstration of the diagnose-then-target paradigm.
- Value: โญโญโญโญโญ Delivers an actionable, cost-effective blueprint for robust model verification and data collection in high-consequence Earth observation applications.