Skip to content

GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/zxyreal/GEO-Detective
Area: LLM Safety
Keywords: Geolocation inference, Large vision-language models, Agents, Privacy leakage, Visual reverse search

TL;DR

This paper introduces GEO-Detective, an autonomous agent that mimics human reasoning and tool-use behaviors for image geolocation inference; by combining adaptive difficulty assessment, experience-augmented prompting, visual segmentation, and visual reverse search, it systematically demonstrates that agentic LVLMs expose severe location privacy risks in everyday shared images.

Background & Motivation

Users routinely share photos on social media that inadvertently capture subtle geographic cues, ranging from architectural styles and signage to environmental landmarks. Historically, uncovering private locations and conducting doxing attacks from raw imagery required specialized detective expertise and tedious manual investigation, limiting such privacy threats to targeted scenarios. However, the emergence of Large Vision-Language Models (LVLMs) has drastically lowered this barrier. Ordinary individuals can now leverage powerful multimodal reasoning to deduce geographic locations from everyday snapshots. Crucially, even when privacy-conscious users strip EXIF metadata, visual content itself frequently leaks sensitive location information, leaving users vulnerable to social engineering schemes such as spear phishing or real-world physical tracking.

Existing geolocation techniques fall into two major limitations: traditional classification and contrastive representation models (e.g., GeoCLIP) perform single-step spatial embeddings that lack fine-grained resolution, interpretability, and multi-turn iterative reasoning; conversely, standard LVLMs or generic tool-augmented agents are not optimized for geolocation. When standard LVLMs encounter complex images lacking conspicuous landmarks, they often transcode visual cues into simplistic text queries or prematurely surrender by predicting "unknown." Such brittle, single-shot paradigms significantly underestimate the true privacy exposure when an adversary utilizes iterative investigation strategies.

Inspired by how human geolocation experts operateโ€”first evaluating scene difficulty from visible cues, then dynamically deploying specialized external tools and cross-verifying hypothesesโ€”the authors build an autonomous agent to evaluate the upper bound of location privacy leakage. Core idea: develop a four-stage adaptive geolocation agent, GEO-Detective, that coordinates visual difficulty assessment, GeoCLIP-aligned experience-augmented prompting, feature segmentation, and visual reverse search within a closed-loop iterative refinement pipeline to uncover latent location privacy risks and evaluate practical defense countermeasures.

Method

Overall Architecture

GEO-Detective takes a user-shared image (stripped of EXIF metadata) as input and outputs a hierarchical location prediction (country, state/region, city) accompanied by a structured evidentiary chain of thought. The system operates across four successive stages: visual feature analysis, which calculates a heuristic difficulty score to categorize the image; strategy execution, which adaptively coordinates direct LVLM reasoning, experience-augmented prompting, feature segmentation, and visual reverse search; results synthesis, which resolves multi-source evidentiary conflicts into a unified location; and iterative refinement, which audits prediction completeness and self-corrects deficiencies through dynamic fallback before finalizing the prediction.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image (EXIF removed)"] --> B["Visual Feature Analysis<br/>8-factor weighted difficulty scoring (1-100)"]
    B --> C["Strategy Execution Scheduling<br/>Adaptive tool orchestration based on difficulty level"]

    C --> D1["Experience Augmented Prompting<br/>GeoCLIP semantic alignment & memory retrieval"]
    C --> D2["Geographic Segmentation & Reverse Search<br/>LVLM cropping code generation + visual web search"]

    D1 --> E["Results Synthesis<br/>Multi-source clue aggregation & conflict resolution"]
    D2 --> E

    E --> F{"Iterative Refinement & Self-Audit<br/>Verify hierarchical completeness & evidence validity"}
    F -->|Incomplete & Budget Left| C
    F -->|Satisfied / Budget Exhausted| G["Final Output: Hierarchical Location + Evidentiary Rationale"]

Key Designs

1. Visual Feature Difficulty Assessment: Guiding Adaptive Strategy Dispatch In open-world geolocation, recognizable global landmarks require minimal reasoning, whereas sparse rural roads necessitate deep investigation and external visual matching. Applying heavyweight tools uniformly across all images introduces computational overhead and spurious retrieval noise. To address this, the agent incorporates an empirical 8-factor heuristic scoring scheme. Starting from a base score of 50, points are adjusted based on visual cues: landmarks (+30), text visibility (+0 to +20), architectural style (+15), unique geographic topography (+15), image quality (-15 to +10), contextual clues such as vehicles/clothing (-10 to +10), scene type (+5 for urban, -5 for rural, -10 for indoor), and a composite cue bonus (+5 to +10). Scores map into five difficulty tiers: Easy (81โ€“100), Moderate (61โ€“80), Difficult (41โ€“60), Very Difficult (21โ€“40), and Extremely Difficult (1โ€“20), dictating whether the planner triggers lightweight reasoning or intensive external search.

2. Experience-Augmented Prompting: Steering Attention via Cross-Modal Alignment Naive prompting frequently causes LVLMs to attend to uninformative background regions (e.g., sky or crowds) rather than location-bearing signals. To align prompt phrasing with discriminative geographic visual elements, five domain categories are predefined: architectural, infrastructure, environmental, urban planning, and signage. The system extracts overlapping image patch embeddings using GeoCLIP's vision encoder and computes cosine similarities against candidate textual elements encoded by GeoCLIP's text encoder:

\[s_{\text{GeoCLIP}}(I, T) = \frac{f_{\text{img}}(I) \cdot f_{\text{text}}(T)}{\|f_{\text{img}}(I)\| \|f_{\text{text}}(T)\|}\]

The top-ranked elements guide the LVLM through three prompt optimization iterations. Prompts exhibiting higher image-to-text similarity than ground-truth label prompts are indexed into an experience memory bank. During inference, visually similar query images retrieve these optimized prompts, focusing LVLM attention directly onto rooftops, facades, and utility poles.

3. Geographic Feature Segmentation and Visual Reverse Search: Preserving Visual Granularity Standard agent frameworks convert visual cues into textual search queries, discarding high-frequency textures, spatial layouts, and regional typography, which often yields irrelevant search engine results. This design implements a two-pronged visual retrieval approach: an LVLM identifies geographically distinctive regions, predicts bounding coordinates, and generates executable Python cropping scripts evaluated against centrality, completeness, and boundary validity. Rather than converting crops to text, the system submits raw visual crops and full images directly to external image search engines (e.g., Google, Yandex). Retrieved candidates are filtered using GeoCLIP similarity thresholds to prune spurious matches, and the associated webpages are crawled for captions and metadata, preserving rich visual fidelity.

4. Evidence-Driven Results Synthesis and Iterative Refinement: Conflict Resolution and Fallback Auditing Because reverse image search provides external web documents while direct prompting yields internal conjectures, the agent frequently encounters conflicting geographic leads. The synthesis module adjudicates conflicts via a rule-based hierarchy: explicit place names carry top priority, followed by the quantity of independent corroborating sources, and finally consistency with visible visual cues. The preliminary hierarchy (country, state, city) is then passed to an autonomous self-audit module that inspects administrative completeness and evidentiary coherence without ground-truth access. If ambiguities or omitted tiers are diagnosed, the central planner triggers a dynamic fallback (e.g., escalating from prompt-based analysis to focused segment search) and iterates until confidence standards are satisfied or resource limits expire.

Key Experimental Results

Main Results

GEO-Detective was evaluated on the standard IM2GPS3K benchmark (3,000 images), the MP16-Pro test set (1,000 images), and the contamination-free DOXBENCH dataset (500 images across California cities). Metrics include localization accuracy at various distance thresholds (@1km, @25km), administrative accuracy (Country, State, City), and the "Unknown" prediction rate.

Performance comparisons on IM2GPS3K against established geolocation paradigms are summarized below:

Type Method @1km Accuracy (%) @25km Accuracy (%)
Agent GEO-Detective (Ours) 11.3 47.5
Agent GeoMiner (ICLR'26) 10.8 46.7
Agent smileGeo (KDD'25) 10.9 38.2
LVLM G3 (NeurIPS'24) 16.6 40.9
LVLM Img2Loc (SIGIR'24) 15.3 39.8
LVLM GeoReasoner (ICML'24) 9.9 33.8
Trained Model PIGEON (CVPR'24) 11.3 36.7
Trained Model GeoCLIP (NeurIPS'23) 14.1 34.5

On MP16-Pro, unknown prediction rates for o3 and GEO-Detective across image difficulty levels are compared below:

Difficulty Level (# Samples) o3 Baseline Unknown Rate (%) GEO-Detective Unknown Rate (%) Absolute Reduction (%)
Easy (269) 21.9 18.6 -3.3
Moderate (281) 29.2 26.3 -2.9
Difficult (349) 45.8 22.6 -23.2
Very Difficult (97) 55.7 28.9 -26.8
Extremely Difficult (4) 75.0 50.0 -25.0

Ablation Study

On MP16-Pro using OpenAI o3, the contributions of Experience-Augmented Prompting (EAP), Reverse Search (RS), Segmentation (Seg), and their combinations were evaluated for country-level accuracy across difficulty levels (unit: %):

Difficulty Base + EAP + RS EAP + RS Base + Seg + RS GEO-Detective (Full)
Easy (269) 74.0 78.4 (+4.4) 76.6 (+2.6) 77.0 (+3.0) 74.7 (+0.7) 77.3 (+3.3)
Moderate (281) 60.1 61.6 (+1.5) 50.9 (-9.2) 57.3 (-2.8) 50.5 (-9.6) 57.7 (-2.4)
Difficult (349) 35.5 35.2 (-0.3) 37.8 (+2.3) 40.4 (+4.9) 32.1 (-3.4) 40.1 (+4.6)
Very Difficult (97) 25.8 15.5 (-10.3) 26.8 (+1.0) 33.0 (+7.2) 15.5 (-10.3) 28.9 (+3.1)
Ext. Difficult (4) 0.0 0.0 (0.0) 0.0 (0.0) 0.0 (0.0) 0.0 (0.0) 0.0 (0.0)

To explore mitigation, four defenses (original, watermark stating prohibition, visual prompt injection VPI, trigger symbols, and EXIF alteration) were tested against baseline o3 and GEO-Detective:

Inference Setup Defense Mechanism Country Acc. (%) State Acc. (%) City Acc. (%) Unknown Rate (%)
o3 Baseline Original (No Defense) 50.0 34.0 27.0 33.0
o3 Baseline Watermark 6.0 5.0 4.0 94.0
o3 Baseline Visual Prompt Injection (VPI) 39.0 31.0 27.0 15.0
o3 Baseline Trigger-based 49.0 33.0 29.0 14.0
o3 Baseline EXIF Modification 52.0 38.0 27.0 30.0
GEO-Detective Original (No Defense) 51.0 34.0 30.0 21.0
GEO-Detective Watermark 10.0 6.0 5.0 84.0
GEO-Detective Visual Prompt Injection (VPI) 41.0 32.0 27.0 17.0
GEO-Detective Trigger-based 53.0 37.0 29.0 18.0
GEO-Detective EXIF Modification 51.0 34.0 32.0 23.0

Key Findings

  • Substantial Gains in Challenging Regimes: On Difficult and Very Difficult test images where standalone LVLMs struggle, GEO-Detective dramatically cuts unknown prediction rates from 45.8% and 55.7% down to 22.6% and 28.9%, respectively.
  • Synergy Between Visual Reverse Search and Experience Prompting: While external visual retrieval provides grounding clues for unfamiliar sites, experience prompts concentrate model attention on discriminative architectural cues, producing a +7.2% country accuracy leap on Very Difficult images.
  • Watermarks Deter Inference, While Agents Resist Minor Perturbations: Explicit textual watermarks prohibiting geolocation trigger ethical refusal safety bounds in LVLMs (unknown rate surges to 84%โ€“94%). In contrast, subtle visual disruptions (VPI and triggers) fail to deceive the agent, which exhibits higher robustness than raw baselines.

Highlights & Insights

  • Heuristic-Driven Staged Tool Activation: Rather than routing every query through compute-heavy search tools, calibrating task difficulty prior to tool orchestration balances operational efficiency while digging deeper into ambiguous images.
  • Contrastive Learning-Guided Prompt Evolution: Leveraging GeoCLIP similarity to systematically optimize and store prompts automates prompt engineering and aligns multimodal representations without manual intervention.
  • Privacy Threat Escalation via Agentic Tool Use: The study demonstrates that safety alignment inside raw LVLMs is insufficient; equipping models with visual retrieval tools significantly amplifies location privacy exposure.

Limitations & Future Work

  • Retrieval Drift in Moderate Scenarios: In moderately complex scenes with ambiguous regional styles, reverse search can retrieve visually similar yet geographically misleading candidates, causing slight accuracy drops (e.g., country accuracy sliding from 60.1% to 57.7%). Stronger spatial verification filters are required.
  • Fundamental Bounds on Information-Deprived Images: On Extremely Difficult images depicting generic nature or close-up textures lacking geographic markers, accuracy remains 0%, underscoring fundamental physical entropy limits in visual localization.
  • Need for Semantic-Level Visual Privacy Protections: Because agentic reasoning bypasses superficial pixel perturbations, future defensive research must explore adversarial inpainting and irreversible visual feature obfuscation.
  • vs GeoCLIP / PIGEON: While coordinate-regression models provide fast single-step embeddings, they cannot generate step-by-step interpretable reasoning or query real-time external indices. GEO-Detective achieves 47.5% @25km accuracy on IM2GPS3K, outperforming GeoCLIP (34.5%) and PIGEON (36.7%).
  • vs GeoMiner / smileGeo: Existing multi-agent geolocation systems rely predominantly on text-based web queries or agent debates. GEO-Detective leverages direct image-to-image reverse search and contrastive prompt optimization, yielding higher precision on difficult imagery and markedly lower unknown prediction rates.

Rating

  • Novelty: โญโญโญโญโ˜† (Systematically operationalizes human-like adaptive detective reasoning and visual search into an agentic framework)
  • Experimental Thoroughness: โญโญโญโญโญ (Rigorous cross-model evaluations across MP16-Pro, IM2GPS3K, DOXBENCH, along with defense stress-testing)
  • Writing Quality: โญโญโญโญโญ (Well-structured methodology, crisp experimental figures, and clear narrative progression)
  • Value: โญโญโญโญโญ (Provides timely empirical evidence on the escalating location privacy risks posed by multimodal agents)