Towards High-Resolution Visual Perception via Hierarchical Entity Exploration¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Object Detection / Multimodal VLM
Keywords: high-resolution visual perception, multimodal large language models, entity exploration, hierarchical tree search, training-free
TL;DR¶
To tackle entity fragmentation and background interference caused by conventional geometric image partitioning in high-resolution visual perception, this paper proposes Hierarchical Entity Exploration (HEE)โa training-free, model-agnostic framework that structures visual search around object-level entities, clusters them into coherent candidate subregions, and leverages dual scoring with confidence-guided backtracking to achieve superior accuracy and efficiency.
Background & Motivation¶
Current multimodal large language models (MLLMs), such as Qwen2.5-VL, InternVL2.5, and LLaVA-OneVision, operate under fixed visual pretraining input resolutions (e.g., 448ร448 or 336ร336). When confronting 4K or 8K ultra-high-resolution (HR) inputs, directly downsampling or resizing uniformly compresses fine-grained structures, triggering severe geometric distortion and dramatic performance degradation in detail-dependent tasks like visual grounding, small-object identification, and dense document OCR. To bridge this resolution gap, existing literature bifurcates into two main paradigms: training-based methods (e.g., DeepEyes) that fine-tune models to actively invoke visual zoom-in tools via reinforcement learning, which are hindered by prohibitive training costs and poor cross-architecture transferability; and training-free heuristic methods (e.g., ZoomEye, RAP) that recursively crop or retrieve uniform grid patches across the image canvas.
However, existing training-free approaches suffer from a fundamental flaw: their spatial partitioning boundaries are purely geometric and entirely agnostic to semantic object boundaries. This geometric blindness induces two severe bottlenecks: entity fragmentation, where visual targets straddling artificial grid boundaries are sliced into disjoint, incomplete components; and background interference, where geometric patches inevitably swallow vast expanses of irrelevant background clutter. Through a rigorous hierarchical decoupling analysis, the authors uncover a foundational insight: without applying any zoom-in enlargement, progressively removing background pixels around the ground truth leads to monotonic accuracy gains across all base MLLMs, whereas padding targets with semantic-free background pixels causes sharp performance drops. This confirms that the primary dividend of image cropping stems from background elimination rather than mere target upscaling.
Recognizing the contradiction between rigid geometric slicing and continuous physical object semantics, the authors shift the exploration unit from arbitrary coordinate grids to complete semantic entities. Core idea: recast high-resolution perception into query-guided hierarchical entity tree search, utilizing a frozen detector to extract complete visual objects and cluster them into coherent semantic nodes, while employing dual scoring and confidence-guided backtracking for adaptive, training-free perception.
Method¶
Overall Architecture¶
The Hierarchical Entity Exploration (HEE) framework reformulates static high-resolution image perception into a dynamic, query-oriented hierarchical tree search. The end-to-end workflow comprises four tightly coupled stages: first, entity-driven node partitioning invokes a frozen open-world detector to extract complete object-level entities within the current view and clusters them into coherent semantic candidate subregions; second, a dual scoring mechanism evaluates both cross-modal semantic alignment and model-level evidence sufficiency for each candidate; third, adaptive recursive exploration drives the search deeper along the highest-scoring candidate branch; finally, confidence-guided backtracking restores previous decision layers when a deep branch fails to reach sufficient confidence, re-evaluating alternative candidates to prevent irreversible search failures.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: High-Resolution Image & Query"] --> B["Entity-Driven Node Partitioning<br/>Detect entities and cluster into semantic subregions"]
B --> C["Dual Scoring Mechanism<br/>Combine cross-modal similarity & VLM confidence"]
C --> D{"Confidence score exceeds threshold tau"}
D -->|Yes / Terminate| G["Generate Final Answer"]
D -->|No & Depth < Max| E["Adaptive Recursive Exploration<br/>Recursively expand top candidate subregion"]
E --> B
D -->|Branch Exhausted / Low Scores| F["Confidence-Guided Backtracking<br/>Return to parent level to re-evaluate alternatives"]
F --> C
Key Designs¶
1. Entity-Driven Node Partitioning: Constructing Semantic Regions Free of Entity Fragmentation and Background Clutter The central pitfall of uniform grid partitioning is the artificial splitting of continuous objects and the inclusion of extensive background noise. HEE bypasses this by applying a frozen object detector \(\mathcal{D}\) (defaulting to DINO-X, with YOLO-World as an open-vocabulary alternative) on region \(\mathcal{I}^{(l)}\) at level \(l\) to yield discrete semantic entities \(\mathcal{E}^{(l)} = \{e_i^{(l)}\}_{i=1}^{N_l}\), where each entity is represented by its bounding box, category label, and semantic feature embedding. To aggregate fine-grained, scattered boxes into meaningful conceptual candidates, HEE performs K-Means clustering in the semantic embedding space to produce \(K\) semantic clusters \(\mathcal{C}^{(l)} = \{C_k^{(l)}\}_{k=1}^K\). For each cluster, its member bounding boxes are merged into a tight bounding box, yielding the explicit entity candidate set \(\mathcal{R}_{\text{entity}}^{(l+1)} = \{r_k^{(l)}\}_{k=1}^K\). Furthermore, to guarantee full spatial recall against potential detector misses, the framework computes residual regions via pixel-wise mask difference: \(\mathcal{R}_{\text{residual}}^{(l+1)} = \text{MaskDiff}(\mathcal{I}^{(l)}, \bigcup_{k=1}^K r_k^{(l)})\). The union \(\mathcal{R}^{(l+1)} = \mathcal{R}_{\text{entity}}^{(l+1)} \cup \mathcal{R}_{\text{residual}}^{(l+1)}\) forms the candidate search space for the subsequent level. This design ensures that every candidate tightly encapsulates whole semantic targets with a significantly higher foreground ratio while safeguarding full visual coverage.
2. Dual Scoring Mechanism: Harmonizing Vision-Language Matching with Model Evidence Sufficiency Once candidate subregions are formed, the system must determine which region holds the greatest promise and whether a region already possesses sufficient evidence to answer the question without redundant expansion. Relying solely on vision-language encoders (such as CLIP) biases the search toward superficial global keyword alignment while ignoring complex relational queries; conversely, querying the large MLLM alone on ambiguous crops often prompts hallucinated confidence. HEE devises a balanced dual scoring function that integrates a frozen vision-language matching encoder (SigLIP) with the target MLLM: $\(S(r) = \lambda s^{\text{vlm}}(r) + (1-\lambda) s^{\text{sim}}(r)\)$ where \(s^{\text{sim}}(r) = S(q, r)\) denotes the text-visual semantic similarity between question \(q\) and region \(r\), and \(s^{\text{vlm}}(r) = M(r, q)\) represents the target MLLM's internal confidence that region \(r\) provides adequate evidence to answer \(q\). Extensive tuning reveals that setting \(\lambda = 0.5\) strikes an optimal balance. This formulation simultaneously leverages cross-modal alignment to discard large irrelevant expanses and taps into the MLLM's reasoning faculties to verify evidence sufficiency.
3. Adaptive Recursive Exploration: Coarse-to-Fine Dynamic Focus Guided by the dual scoring mechanism, HEE maintains a globally ranked priority queue of candidate nodes. At each exploration layer, if the highest candidate score \(\max_{r} S(r)\) surpasses the confidence threshold \(\tau\) (fixed at 0.5), the search terminates immediately, and the selected region is fed into the MLLM for final answer generation, curbing needless token and inference costs. Otherwise, HEE selects the top-scoring candidate \(r^* = \arg\max_r S(r)\) and recursively executes entity partitioning on it to uncover finer sub-nodes. Remaining candidates at the current level are retained in memory for potential backtracking. The recursive down-sampling terminates adaptively when: (i) confidence exceeds \(\tau\), (ii) the detected entity count drops below a minimum threshold \(\eta\) (e.g., only 1โ2 atomic entities remain), or (iii) the maximum exploration step count \(T_{\max} = 50\) is reached. This enables the model to swiftly zoom from a macro view down to a sub-object detail in minimal steps.
4. Confidence-Guided Backtracking: Mitigating Local Traps and Single-Path Failures Greedy single-path zoom-in pipelines invariably suffer from irreversible mistakes if an ambiguous or deceptive visual cue triggers a false-positive branch selection at an early level. HEE incorporates confidence-guided backtracking to build an adaptive verification loop. When recursive search reaches the terminal leaf level and all detected entities fail to achieve confidence above the threshold (i.e., \(\max_i S(e_i) < \tau\)), the model flags the active branch as unpromising or misleading. It backtracks to the previous decision layer, retrieves the next-best candidate (e.g., Top-2, Top-3) from that layer's stored ranking, and initiates re-partitioning and evaluation. This error-recovery capability mimics human visual scanning, substantially suppressing error propagation in dense, cluttered environments.
Key Experimental Results¶
Main Results¶
On Visual Probe (a rigorous 4K benchmark encompassing Easy, Medium, and Hard splits), HR-Bench 4K, and HR-Bench 8K (covering fine-grained single-instance FSP and cross-instance FCP tasks), HEE delivers substantial performance gains across open-source MLLM families, consistently outperforming training-free competitors like RAP and ZoomEye.
| Benchmark Dataset / Split | Base Model | Baseline Acc. (%) | HEE (Ours) (%) | Gain (ฮ) | Reference Methods (RAP / ZoomEye) |
|---|---|---|---|---|---|
| Visual Probe (Easy) | Qwen2.5-VL-7B | 39.1 | 67.4 | +28.3 | 52.5 / 55.3 |
| Visual Probe (Medium) | Qwen2.5-VL-7B | 26.0 | 47.0 | +21.0 | 31.7 / 35.5 |
| Visual Probe (Hard) | Qwen2.5-VL-7B | 23.9 | 50.0 | +26.1 | 34.0 / 41.5 |
| Visual Probe (Easy) | InternVL2.5-8B | 49.6 | 60.3 | +10.7 | โ |
| Visual Probe (Hard) | InternVL2.5-8B | 17.0 | 41.5 | +24.5 | โ |
| Visual Probe (Easy) | LLaVA-ov-7B | 36.2 | 61.0 | +24.8 | 56.7 / 61.0 |
| Visual Probe (Hard) | LLaVA-ov-7B | 13.4 | 38.6 | +25.2 | 31.1 / 31.1 |
| HR-Bench 4K (Overall) | Qwen2.5-VL-7B | 66.5 | 73.9 | +7.4 | 71.1 / 71.4 |
| HR-Bench 8K (Overall) | Qwen2.5-VL-7B | 62.1 | 71.1 | +9.0 | 69.4 / 67.9 |
| HR-Bench 8K (FSP) | InternVL2.5-8B | 59.5 | 81.8 | +22.3 | โ |
| HR-Bench 8K (Overall) | LLaVA-ov-7B | 57.3 | 67.9 | +10.6 | 60.3 / 66.8 |
On the real-world benchmark MME-RealWorld, deploying HEE on Qwen2.5-VL-7B boosts Remote Sensing accuracy from 39.22% to 46.39% (+7.17%), Monitoring from 38.52% to 46.02% (+7.50%), and lifts overall average accuracy from 59.99% to 63.38% (+3.39%).
Furthermore, an efficiency comparison on HR-Bench 4K using Qwen2.5-VL-7B highlights HEE's marked computational advantages:
| Method | Throughput (samples/min) โ | Total Time (min) โ | Accuracy (%) โ |
|---|---|---|---|
| RAP | 1.6 | 126 | 71.1 |
| ZoomEye | 4.3 | 46 | 71.4 |
| HEE (Ours) | 8.3 | 24 | 73.9 |
Ablation Study¶
Ablation experiments conducted on Visual Probe using Qwen2.5-VL-7B rigorously isolate the contribution of each architectural component:
| Configuration Variant | Easy Acc. (%) | Medium Acc. (%) | Hard Acc. (%) | Average Acc. (%) | Note |
|---|---|---|---|---|---|
| HEE Full Model | 67.3 | 47.0 | 50.0 | 54.7 | Full pipeline with entity exploration, dual scoring, & backtracking |
| w/o Entity Exploration (Uniform Grid) | 61.7 | 44.7 | 37.7 | 48.0 | Plummets by 12.3% on Hard; severe entity fragmentation |
| w/o Backtracking | 60.3 | 44.8 | 41.5 | 48.9 | Unable to recover from suboptimal branch commitments |
| Entity Clusters \(K=2\) | 61.7 | 45.5 | 42.5 | 49.9 | Overly broad clusters retain excess background |
| Entity Clusters \(K=8\) | 62.4 | 45.1 | 37.7 | 48.4 | Over-fragmentation bloats candidate search paths |
| w/o Semantic Similarity \(S(q, r)\) | 64.5 | 45.9 | 43.4 | 51.3 | Relies exclusively on VLM confidence; lacks early grounding |
| w/o VLM Confidence \(M(r, q)\) | 56.0 | 42.9 | 21.7 | 40.2 | Severe collapse; cross-modal matching cannot gauge evidence sufficiency |
| Score Weight \(\lambda = 0.2\) | 65.9 | 44.8 | 41.5 | 50.7 | Skewed toward vision-language feature matching |
| Score Weight \(\lambda = 0.7\) | 66.7 | 46.6 | 46.2 | 53.2 | Skewed toward VLM confidence score |
Key Findings¶
- Entity exploration is indispensable in cluttered, fine-grained scenes: Replacing entity-driven decomposition with uniform grid slicing causes a dramatic 12.3% drop on the Hard subset (50.0% down to 37.7%), proving that preserving organic object boundaries is paramount when targets are tiny and distractors are dense.
- Model confidence is the primary gatekeeper for termination: Removing the VLM confidence term \(M(r, q)\) degrades the average score from 54.7% to 40.2% (with Hard plunging to 21.7%), demonstrating that while text-image alignment locates rough candidates, only the reasoning model can reliably judge whether a crop holds conclusive proof.
- Rapid unimodal step convergence: Step distribution analysis indicates that over 60% of test queries reach termination within 8 to 15 exploration steps, with Easy and Medium peaking around 10 steps. Because semantic clustering condenses trivial candidate regions, HEE processes 8.3 samples/minโ5.2ร faster than RAP and nearly 2ร faster than ZoomEye.
Highlights & Insights¶
- Paradigm Shift via Hierarchical Decoupling: Controlled background masking experiments disprove the long-held assumption that cropping works primarily by upscaling targets; instead, the primary performance gain stems from background removal. This critical insight exposes the fatal flaw of geometry-based grids and motivates entity-level parsing.
- Mask-Difference Residuals for Zero-Blindspot Recall: By pairing semantic entity clustering with a complementary \(\text{MaskDiff}\) residual pool, HEE retains full scene coverage. Empirical analysis reveals that 3.55% of correct predictions succeed even when the detector misses the primary target, capitalizing on background context.
- Model-Agnostic, Plug-and-Play Efficiency: Without updating model weights or engineering complex tool-calling fine-tuning, HEE seamlessly pairs with diverse MLLMs (Qwen2.5-VL, InternVL2.5, LLaVA-OneVision), yielding state-of-the-art accuracy at a fraction of existing inference latencies.
Limitations & Future Work¶
- Visual Ambiguity Beyond Localization: Failure case analysis indicates that in 24.82% of errors, HEE successfully isolates the ground-truth entity crop, yet the underlying MLLM misclassifies the object due to subtle intra-category visual ambiguity (e.g., confusing a kangaroo sculpture with an emu emblem).
- Confidence Misalignment: In a small fraction of queries, the search path encounters the correct candidate but veers into a hallucinated, falsely confident branch, terminating prematurely.
- Future directions include incorporating lightweight test-time confidence calibration and scaling hierarchical entity exploration to high-resolution long video understanding.
Related Work & Insights¶
- vs ZoomEye: ZoomEye performs recursive quadtree zooming over uniform spatial grids, frequently cleaving objects at boundary intersections and retaining substantial background noise; HEE seeds regions from detected object boundaries, outperforming ZoomEye by 8.5% on Visual Probe Hard.
- vs RAP: RAP applies fixed-scale patch retrieval across the entire canvas and reconstructs a flattened multi-crop mosaic, incurring prohibitive token overhead and a 126-minute latency; HEE focuses computation hierarchically on top candidates, slashing runtime to 24 minutes while gaining +2.8% accuracy.
- vs DeepEyes: DeepEyes relies on compute-intensive reinforcement learning to train models to use zoom-in tools; HEE is entirely training-free, easily portable across model architectures, and avoids catastrophic forgetting.
Rating¶
- Novelty: โญโญโญโญโญ Decouples cropping mechanics to prioritize background removal and introduces entity clustering tree search with residual coverage.
- Experimental Thoroughness: โญโญโญโญโญ Rigorous validation across 4K/8K benchmarks, real-world tasks, three distinct MLLM families, and detailed ablation/efficiency analyses.
- Writing Quality: โญโญโญโญโญ Exceptionally clear narrative arc connecting motivation experiments, mathematical formulation, and empirical findings.
- Value: โญโญโญโญโญ Highly practical, training-free, and resource-efficient solution for deploying MLLMs in high-resolution computer vision.