Skip to content

Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/be-chen/REMSA
Area: Remote Sensing / LLM Agent / Segmentation
Keywords: Remote sensing foundation models, model selection, intelligent agents, constraint-aware reasoning, foundation model database

TL;DR

To tackle the challenge of selecting suitable remote sensing foundation models under heterogeneous documentation and strict deployment constraints, this paper introduces RS-FMD (a structured database of 160+ RSFMs) and Remsa, a constraint-aware LLM agent that automates transparent model selection via structured metadata grounding, adaptive orchestration, and in-context ranking.

Background & Motivation

With the growing availability of remote sensing (RS) satellite missions and airborne platforms (e.g., Sentinel-2 multispectral imagery, Sentinel-1 synthetic aperture radar SAR, EnMAP hyperspectral imaging, and LiDAR sensors), multi-sensor Earth observation data are increasingly integrated into critical downstream pipelines such as land cover mapping, flood extraction, wildfire assessment, and urban expansion monitoring. To alleviate the reliance on massive labeled downstream datasets, remote sensing foundation models (RSFMs) have proliferated rapidly. These models span unimodal vision encoders (e.g., MMEarth, MA3E), cross-sensor fusion architectures (e.g., OmniSat), and multimodal vision-language models (e.g., GeoText, LHRS-Bot, SkySense). Each foundation model demonstrates distinct advantages across spatial, spectral, and temporal resolutions, as well as downstream tasks like semantic segmentation, change detection, classification, and visual question answering (VQA).

However, selecting an optimal RSFM for a practical operational workflow remains an arduous and error-prone hurdle. Unlike generic computer vision settings, real-world remote sensing deployments are heavily constraint-driven: practitioners must balance diverse sensor modalities, variable spectral band sets, heterogeneous geospatial extents, strictly bounded GPU memory budgets, and scarce downstream annotation. Currently, information on hundreds of available RSFMs is scattered across disparate publications, model cards, and open-source repositories without any standardized machine-readable schema. Public benchmarks such as GEO-Bench-2 primarily focus on standardized downstream adaptation accuracy on static tasks rather than addressing upstream selection under user-specific deployment constraints, while classical AutoML frameworks cannot handle open-ended, natural language requirement specifications.

The core tension lies in the gap between unstructured, constraint-heavy user intent and the fragmented, heterogeneous documentation of rapidly expanding RSFMs. Core idea: systematically construct the first structured schema-guided database of 160+ RSFMs (RS-FMD), and build Remsa, a constraint-aware modular LLM agent that establishes a closed-loop selection pipeline integrating intent parsing, dense retrieval, rule-based hard constraint filtering, in-context ranking, interactive multi-turn clarification, and transparent explanation.

Method

Overall Architecture

Remsa adopts a modular, decision-centered agent architecture. Given an unstructured natural language query expressing remote sensing objectives and operational constraints, Remsa outputs a ranked set of top-\(k\) candidate RSFMs accompanied by confidence scores and transparent justification reports. The core of Remsa comprises an LLM agent core (an Interpreter and a Task Orchestrator) coordinated with a suite of decoupled external tools. The Interpreter parses free text into a schema-guided specification covering mandatory fields (e.g., application target, required data modality) and optional constraint fields (e.g., compute budget, downstream fine-tuning volume, performance priorities). The Task Orchestrator monitors the task state and dynamically invokes tools across retrieval, constraint filtering, in-context ranking, clarification, and explanation, while a vector-based Task Memory mechanism preserves user preferences across sessions.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["User Natural Language Query & Deployment Constraints"] --> B["Structured Metadata Extraction & Dual-Confidence Human Verification<br/>RS-FMD database construction and confidence check"]
    A --> C["Intent Interpretation & Task-Aware Dynamic Orchestration Loop<br/>Parse mandatory & optional constraints with adaptive scheduling"]
    B -.-> D["Structured Dense Retrieval<br/>Sentence-BERT & FAISS high-recall candidate search"]
    C --> D
    D --> E["Rule-Based Hard Constraint Filtering & Few-Shot In-Context Ranking<br/>Deterministic hard-constraint elimination & LLM ranking"]
    E -->|Low confidence or missing constraints| F["Progressive Multi-Turn Clarification & Memory Augmentation<br/>Targeted questions to refine task specification"]
    F --> C
    E -->|Termination conditions met| G["Explainable Recommendation Report Generation<br/>Synthesize selections, rationales, and open-source assets"]

Key Designs

1. Structured Metadata Extraction & Dual-Confidence Human Verification: Unifying 160+ RSFMs into RS-FMD

Addressing the issue of scattered, unstructured model documentation across papers, model cards, and codebases, this design establishes the Remote Sensing Foundation Model Database (RS-FMD). RS-FMD codifies over 160 models under a unified schema including supported sensor modalities, architectural families, pretraining datasets, parameter counts, open-weight repositories, and reported benchmark metrics. To populate the database accurately at scale without overwhelming manual effort, a semi-automated iterative extraction pipeline is paired with field-level uncertainty estimation. To eliminate hallucination risks, a dual confidence scoring formula integrates normalized generation log-probabilities with multi-sampling self-consistency:

\[ \text{Confidence} = w_{\text{logp}} \cdot \text{NormalizedLogProb} + w_{\text{cons}} \cdot \text{SelfConsistency} \]

Log-probabilities are normalized via a sigmoid function at temperature \(\tau=0.5\) to preserve sensitivity in moderate-confidence regimes, with empirical weights set to \(w_{\text{logp}}=0.7\) and \(w_{\text{cons}}=0.3\). Fields scoring below the threshold \(\theta=0.75\) are flagged for human expert inspection. This human-in-the-loop strategy concentrates expert effort strictly on ambiguous metadata entries, guaranteeing high fidelity while ensuring scalable maintenance.

2. Intent Interpretation & Task-Aware Dynamic Orchestration Loop: Constraint-Driven State Transitions

Traditional agent frameworks relying on unconstrained single-step action selection often falter when handling underspecified or constraint-heavy domain requirements. Remsa introduces a dedicated Interpreter and a dynamic Task Orchestrator. The Interpreter enforces a strict schema prompt to extract mandatory technical requirements (e.g., target application, sensor modality) and capture operational constraints (e.g., GPU memory limits, labeled sample availability). The Orchestrator manages an adaptive finite-state loop that evaluates constraint completeness, candidate set cardinality, and ranking confidence. It halts retrieval when mandatory constraints are absent, triggers interactive clarification when candidates are overwhelming or ambiguous, and activates a closest-match fallback when no candidate completely satisfies all hard constraints.

3. Rule-Based Hard Constraint Filtering & Few-Shot In-Context Ranking: Two-Stage Hybrid Candidate Pruning

Dense retrieval via Sentence-BERT encodes metadata fields with explicit structural prefix tokens (e.g., [APPLICATION], [MODALITY]) and searches candidates via FAISS cosine similarity to achieve maximum recall. However, similarity matching alone cannot strictly enforce complex engineering requirements. The Ranking Tool resolves this through a hybrid pruning pipeline: first, a deterministic rule engine discards candidates that violate hard physical constraints (e.g., rejecting an optical-only model when SAR data is specified, or excluding models exceeding GPU VRAM limitations). Second, an in-context LLM ranking promptโ€”grounded with expert-curated few-shot demonstrationsโ€”re-ranks the surviving candidates by balancing pretraining data diversity, reported benchmark evidence, and architectural transfer efficiency, outputting an ordered recommendation list with calibrated confidence scores.

4. Progressive Multi-Turn Clarification & Memory Augmentation: Conversational Grounding and Long-Term Preference Retention

Real-world non-expert users frequently formulate incomplete or ambiguous queries (e.g., omitting sensor ground sample distance or hardware specifications). The Clarification Generator Tool inspects unpopulated schema attributes and poses targeted, multiple-choice or short-answer clarification prompts (bounded by a maximum of 3 turns to prevent user fatigue). User responses are dynamically integrated into the task specification to refresh candidate ranking. Concurrently, a lightweight vector-based Task Memory indexes historical user sessions and environmental constraints, retrieving relevant operational context via cosine similarity to personalize future recommendations and accelerate convergence.

A Worked Example

Consider a user providing an underspecified requirement: "I need to perform flood surface water segmentation, but I only have an edge workstation equipped with a single RTX 4090 GPU." 1. Interpretation: The Interpreter extracts the application target as "surface water segmentation/change detection" and the compute budget as "single 24GB GPU," identifying "modality" and "spatial resolution" as missing mandatory attributes. 2. Clarification: The Clarification Generator prompts: "Is your input data Sentinel-1 dual-polarization SAR or Sentinel-2 optical imagery? Are multitemporal series available?" The user clarifies: "I am using Sentinel-1 dual-polarization SAR." 3. Retrieval & Hard Filtering: Dense retrieval retrieves 20 candidate models; the rule-based filter immediately eliminates optical-only vision models (e.g., MA3E, GeoText) and multi-billion-parameter clusters, leaving SAR-compatible candidates (e.g., OmniSat, SkySense-SAR). 4. In-Context Ranking & Explanation: The LLM ranks remaining candidates based on documented SAR water segmentation performance and single-GPU inference efficiency, recommending OmniSat and SkySense as top selections while providing code repository links and resource justification.

Key Experimental Results

Main Results

The evaluation benchmark comprises 100 realistic, natural language query scenarios verified by Earth observation and machine learning domain experts. Top-3 model recommendations generated by Remsa and three baseline systems were evaluated in a blind expert protocol across 7 criteria (Application Compatibility, Modality Match, Reported Performance, Efficiency, Popularity, Generalizability, Recency), scaled to 1โ€“100. The baselines include: - Remsa-Naive: Same tools and database, but relying on LangChain's unconstrained single-step agent orchestration; - DB-Retrieval: Pure FAISS dense retrieval over RS-FMD without LLM ranking or constraint reasoning; - Unstructured-RAG: Generic RAG providing unparsed model documentation directly to the LLM.

Performance comparison under the primary backbone (GPT-4.1) is summarized below:

System Avg Top-1 Score Avg Set Score Top-1 Hit Rate HQ Hit Rate (\(\ge 80\)) MRR
Remsa (Ours) 75.76 75.03 21.33% 40.00% 0.34
Remsa-Naive 72.67 72.00 20.00% 37.33% 0.29
Unstructured-RAG 71.23 68.39 13.33% 30.67% 0.24
DB-Retrieval 67.37 68.87 12.00% 17.33% 0.23

Robustness evaluation across different LLM backbones (DeepSeek-V3.2 and LLaMA-3.3-70B):

LLM Backbone System Avg Top-1 Score Avg Set Score Top-1 Hit Rate HQ Hit Rate MRR
DeepSeek-V3.2 Remsa (Ours) 75.35 73.81 18.67% 40.00% 0.30
DeepSeek-V3.2 Remsa-Naive 72.03 71.83 16.51% 36.89% 0.26
DeepSeek-V3.2 Unstructured-RAG 69.19 70.94 10.67% 24.00% 0.24
DeepSeek-V3.2 DB-Retrieval 67.37 68.87 12.00% 17.33% 0.23
LLaMA-3.3-70B Remsa (Ours) 73.39 70.34 14.67% 32.00% 0.26
LLaMA-3.3-70B Remsa-Naive 69.02 69.00 14.23% 29.47% 0.24
LLaMA-3.3-70B Unstructured-RAG 69.87 68.04 10.00% 26.67% 0.22
LLaMA-3.3-70B DB-Retrieval 67.37 68.87 12.00% 17.33% 0.23

Ablation Study

To dissect the influence of individual evaluation criteria on expert suitability scores, a leave-one-out sensitivity analysis was conducted on the full rubric under GPT-4.1:

Criteria Setting Avg Set Score Top-1 Hit Rate MRR Note / Impact
Full Scoring (All Criteria) 75.03 22.67% 0.38 Standard benchmark baseline
w/o Application Compatibility 73.32 21.33% 0.36 Noticeable drop; confirms task alignment is a core driver
w/o Modality Match 70.88 22.67% 0.36 Sharpest drop (-4.15 pts); highlights modality as a mandatory physical constraint
w/o Reported Performance 75.05 22.67% 0.38 Negligible change; literature evidence is largely captured via model family
w/o Efficiency 80.23 25.33% 0.38 Score increases; reflects that hardware constraints penalize high-capacity models
w/o Popularity + Recency 75.13 25.33% 0.39 Slight gain; confirms Remsa prioritizes technical fitness over citation hype
w/o Generalizability 75.10 22.67% 0.38 Minimal impact; pretraining scale is implicitly correlated with modality capabilities

Key Findings

  • Crucial Role of Structured Metadata Grounding: DB-Retrieval and Unstructured-RAG both fall short on High-Quality Hit Rate (\(\le 30.67\%\)). Grounding selections in structured RS-FMD metadata elevates Remsa's HQ Hit Rate to 40.00% and substantially improves MRR from 0.23 to 0.34.
  • Value of Dynamic Task Orchestration: Compared with Remsa-Naive, the adaptive orchestration loop boosts Avg Top-1 Score from 72.67 to 75.76 and MRR from 0.29 to 0.34, validating that iterative state management and clarification are vital for constraint reconciliation.
  • Model-Agnostic Generalization: Remsa maintains superior performance across commercial (GPT-4.1) and open-weights backbones (DeepSeek-V3.2, LLaMA-3.3-70B). Notably, DeepSeek-V3.2 matches GPT-4.1 with an identical 40.00% HQ Hit Rate.
  • Inference Latency Trade-Off: End-to-end execution latency averages 0.77s for DB-Retrieval, 11.9s for Unstructured-RAG, 22.7s for Remsa-Naive, and 31.7s for Remsa. The moderate time investment yields a significant margin in recommendation suitability and decision transparency.

Highlights & Insights

  • Shifting from Blind Downstream Tuning to Upstream Selection: Rather than running expensive, repetitive downstream fine-tuning across hundreds of candidate models, Remsa formalizes upstream model selection as an automated, constraint-aware reasoning stage.
  • Dual-Confidence Scoring for Grounded Knowledge Bases: Combining generation log-probabilities with multi-sample consistency provides a principled threshold (\(\theta=0.75\)) to focus human expert verification only where metadata ambiguity exists.
  • Pragmatic, Physics-Grounded Reasoning: Sensitivity analysis reveals that omitting popularity metrics (citations/stars) does not degrade performance, demonstrating that Remsa bases its decisions on physical sensor alignment, spectral capability, and hardware feasibility rather than popularity bias.

Limitations & Future Work

  • Acknowledged Limitations: The evaluation covers 100 expert-crafted queries with 3,000 rigorous blind scorings, but may still omit highly specialized or emerging niche earth observation tasks. Furthermore, ranking currently relies on in-context demonstrations rather than parameter fine-tuning or reinforcement learning from expert feedback.
  • Identified Limitations: Reported downstream performance metrics in RS-FMD are extracted from source papers that adopt heterogeneous validation splits and pre-processing protocols, introducing slight inherent variance across self-reported metrics.
  • Future Directions: Integrating lightweight execution sandboxes or surrogate performance predictors could enable rapid empirical verification on downstream sample slices; extending the framework to other sensor-heavy domains (e.g., multimodal medical imaging and autonomous driving) represents a natural next step.
  • vs Standard RS Benchmarks (e.g., GEO-Bench, GEO-Bench-2): Existing benchmarks assess downstream fine-tuning accuracy on static datasets; Remsa operates upstream to dynamically match open-ended user hardware, sensor, and operational constraints to optimal candidate architectures.
  • vs Geospatial Agents (e.g., GeoLLM-Squad, RS-Agent): Prior geospatial agents act as interactive execution assistants solving vision tasks (e.g., object detection, spatial VQA); Remsa is a meta-decision agent designed to optimize the selection of foundation models across the entire Earth observation lifecycle.
  • vs Classical AutoML (Auto-WEKA, Auto-sklearn): Traditional AutoML searches over bounded tabular hyperparameter grids; Remsa handles rich natural language requirements and multi-modal pretraining attributes of modern foundation models through structured LLM orchestration.

Rating

  • Novelty: โญโญโญโญโญ [Pioneering upstream constraint-aware foundation model selection in remote sensing, backed by the comprehensive RS-FMD database]
  • Experimental Thoroughness: โญโญโญโญโญ [3,000 blind expert scorings across 100 diverse scenarios, 4 system variants, and 3 major LLM backbones with detailed sensitivity ablations]
  • Writing Quality: โญโญโญโญโญ [Clean modular design, comprehensive technical formulation, clear pipeline diagrams, and transparent empirical analysis]
  • Value: โญโญโญโญโญ [Substantially lowers the entry barrier for applying cutting-edge foundation models to practical, resource-constrained Earth observation tasks]