Skip to content

ProCompNav: Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries

Conference: NeurIPS2026 Oral
arXiv: 2605.06223
Code: https://github.com/tree-jhk/procompnav
Area: Robotics & Embodied AI
Keywords: instance navigation, disambiguation, comparative judgment, binary feedback, multi-view memory

TL;DR

ProCompNav collects same-category objects in an unknown environment, then recursively identifies distinguishing attributes and asks users yes/no questions, achieving success rates of 23.7%, 28.1%, and 17.0% across the three simulated CoIN-Bench splits while substantially shortening user-simulator responses.

Background & Motivation

Instance navigation requires finding the particular chair intended by the user, not just any chair. Prior methods often assume a sufficiently detailed description upfront, whereas a natural request may simply be “find the cabinet.” AIUTA allows clarification during navigation, but primarily judges each encountered candidate against accumulated descriptions. If an incorrect cabinet shares the target’s color, handles, and material, additional descriptions may not provide evidence that distinguishes it from the target, and the robot can still stop prematurely.

Collecting multiple candidates before matching reduces commitment to the first distractor, but does not automatically resolve ambiguity. The paper’s Pooled Independent Matching baseline asks open-ended questions for candidates sequentially and scores them afterward; the accumulated information may still concern shared properties and require repeated descriptions from the user. The missing capability is not obtaining more target details, but determining which details are worth asking about given the candidates already observed. Questions cannot assume that the agent shows images to the user: CoIN allows only language interaction, so users must answer from their own knowledge of the target.

ProCompNav therefore changes the decision from “does this object resemble the target?” to “which attribute differentiates the current candidates?” Each round does not need an attribute unique to the target; an attribute-value pair that divides the current candidates into two non-empty groups suffices to eliminate one group with one answer. Core idea: build a comparable candidate pool, then recursively ask binary questions about differences between candidates, isolating the target through short feedback rather than requesting a complete description for each candidate.

Method

Overall Architecture

The input is an open-vocabulary category and RGB-D observations obtained during exploration; the output is a selected instance and a stop action near it. The robot has a 30° field of view and can move forward, turn left, turn right, stop, or ask a language question; the number and locations of same-category instances are unknown. The system first builds a “Multi-view Candidate Pool,” then cycles through “Similar Core Selection,” “Discriminative Attributes and Group Refinement,” and “Binary Pruning and Re-exploration.” The latter three steps constitute Recursive Comparative Judgment (RCJ).

Comparison normally starts after collecting 5 candidates. On CoIN, it also starts at step 400 if the threshold has not been reached, preventing exploration from consuming the entire 500-step budget. Questions are based on differences within the current pool rather than a fixed color or material questionnaire. An answer retains the consistent group; a single remaining candidate becomes the navigation target. Feedback indicating that no pooled candidate is suitable triggers exploration again. This loop handles some cases where the target has not entered the pool, but does not recover a target already removed by an incorrect answer.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Category request + RGB-D"] --> B["Multi-view Candidate Pool"]
    B --> C["Similar Core Selection"]
    C --> D["Discriminative Attributes<br/>and Group Refinement"]
    D --> E["Binary Pruning<br/>and Re-exploration"]
    E -->|At least two candidates| C
    E -->|No and empty remainder: add one candidate| B
    E -->|One candidate| F["Navigate to instance and stop"]
    U["User answer"] --> E

The diagram shows CoIN inference data flow, not a training network: user feedback performs online pruning rather than updating model parameters. TextNav prohibits questions, substitutes the target description for user feedback, and adds final candidate verification; this diagram should not be interpreted as an interactive TextNav procedure.

Key Designs

1. Multi-view Candidate Pool: collect comparable instances before committing to the first detection

The system uses VLFM as its exploration backbone but does not terminate at the first category detection. Each detected region is back-projected into a 3D point cloud and compared with accumulated clouds of existing candidates. Appendix B computes point-neighbor fractions in both directions with a 0.03 m neighborhood radius, then takes their maximum. A maximum overlap of at least 0.3 assigns the detection to an existing instance; otherwise, a new candidate is created. This is a geometric-overlap association rule, not a guarantee of correct instance identity.

Each candidate stores multiple RGB views. DINOv2 embeddings are clustered with KMeans, with centroid-nearest views from 6 clusters used by default to form a collage; the multimodal model then generates a unified description. Descriptions include appearance, nearby objects, and spatial relations, allowing context such as “a red box nearby” to support comparison. Multi-view processing is intended not merely to increase image count, but to reduce missing attributes caused by context invisible from one viewpoint. Association errors can nevertheless split one physical object into several candidates.

Exploration also uses two auxiliary heuristics: when a recent-position stagnation score reaches 0.9 for 5 consecutive steps, frontiers in the same grid cell are temporarily blacklisted; when surrounding openness is at least 0.1 and the agent is at least 1.0 m from its previous rotation point, it rotates 360° in place. These address local loops and missed detections under a narrow field of view. They are not additional RCJ reasoning stages and do not guarantee exhaustive coverage.

2. Similar Core Selection: anchor a describable contrast with two similar candidates

With more than two candidates, directly asking a model to summarize all differences may yield heterogeneous attributes that are difficult to turn into a question. RCJ first separates the active pool into a core and a remainder. Pairwise similarity averages cosine similarities of normalized text and visual embeddings. It repeatedly removes the instance with the lowest total similarity to the others until two remain; these form the core, while removed candidates form the remainder.

Similar core candidates tend to share attributes, making it easier to extract a common property and then test whether it contrasts with the remainder. Greedy peeling guarantees only non-decreasing mean within-group similarity along the peeling path; it does not find a globally most-similar pair, maximize information gain, or enforce a balanced partition. With exactly two candidates, peeling is skipped, each candidate forms a singleton group, and the attribute with the largest entailment-score difference is selected.

3. Discriminative Attributes and Group Refinement: turn proposed differences into a verifiable partition

The large language model (LLM) extracts shared attribute-value pairs from core descriptions, such as “a red box nearby,” rather than attribute names such as “color” alone. A Natural Language Inference (NLI) model takes a candidate description as the premise and “the instance has this attribute” as the hypothesis, scoring every candidate–attribute combination. Appendix C converts entailment, neutral, and contradiction logits into the following score. It measures support in the description, not image-level truth probability, and does not itself establish empirical probability calibration.

\[ s(d_i,a)=\sigma\!\left(\ell_E(d_i,a)-\max\{\ell_N(d_i,a),\ell_C(d_i,a)\}\right). \]

RCJ selects the attribute maximizing mean core support minus mean remainder support. This favors attributes shared within the core yet discriminative relative to the remainder. The LLM proposes natural-language hypotheses, while NLI compares them consistently. This is more controlled than a single prompt asking an LLM to select an attribute and partition the pool, but if descriptions omit or misstate context, NLI still verifies incorrect textual evidence.

\[ a_t^*=\arg\max_{a\in\mathcal A}\left(\mathbb E_{i\in G_c}[s(d_i,a)]-\mathbb E_{j\in G_r}[s(d_j,a)]\right). \]

The initial core is not necessarily the complete set of candidates possessing the attribute. The system therefore scans the remainder and moves candidates with support of at least \(\tau=0.9\) into the core. This prevents a “yes” answer from deleting a target that has the attribute but was not in the similar core. If either refined group is empty, the next attribute is tried in descending entailment-gap order. If none produces two non-empty groups, the last attribute is still queried. Consequently, guaranteed candidate elimination applies to valid partitions, not unconditionally to the entire system.

4. Binary Pruning and Re-exploration: retain the answer-consistent group and detect a potentially missing target

The system asks whether the target possesses the selected attribute, retaining the core after “yes” and the remainder after “no.” With at least two candidates remaining, it repeats core selection and attribute comparison; a single candidate identifies the target. Questions are constructed relative to the current pool, so no attribute needs to uniquely identify the target in one step. Several rounds of local distinctions can resolve ambiguity. Asking consecutive questions after collection also reduces repeated interruptions during exploration.

Without a valid split, “yes” can leave the pool unchanged. If the answer is “no” and the remainder is empty, the system interprets this as a possible missing target, explores to add one candidate, pre-prunes using target facts collected from earlier answers, and resumes RCJ. Re-exploration remedies some pool-absence cases rather than providing general rollback: the current method cannot recover a true target pruned by incorrect binary feedback. “I don’t know” in the noise experiments instead skips that attribute and tries a lower-ranked one; it is not equivalent to “no.”

Non-interactive TextNav retains the collect-and-compare approach but extracts a fixed attribute set from the detailed goal description upfront. It always retains the attribute-supported core rather than asking users. When a single candidate remains or the remainder is empty, a text-only LLM verifies the candidate description against the complete goal. With multiple core candidates and an empty remainder, it verifies the candidate with the highest attribute support. Rejection triggers exploration again. Thus, TextNav supports candidate-relative comparison under detailed descriptions, not real-user clarification of ambiguous requests.

A Worked Example

The request in Figure 5 is “find the dresser.” The robot collects several dressers instead of stopping because one matches the target’s color and handles. The first question asks whether a red box is nearby. A “yes” eliminates three distractors at once. A second question asks whether a TV is on top, isolating the target among the remaining candidates.

Each question needs only a local distinction. The red box need not be unique to the target, and the TV need not appear in the initial request; candidate comparison reveals that these contextual properties are worth asking about. If the TV in the second round is visible only from a side view, whether the multi-view description preserves that evidence directly affects subsequent pruning.

Loss & Training

This is a zero-shot inference framework requiring no task-specific training, with no new navigation loss, policy training, or user-feedback fine-tuning. Text-only and multimodal modules share Qwen3-VL-8B, while NLI uses DeBERTa-v3-large. CoIN allows at most 500 steps and a question budget of 4; TextNav allows 1000 steps and starts comparison no later than step 600. Both tasks define success as executing stop within 1 m of the target.

The main results depend jointly on pretrained models, the exploration backbone, and rules. Zero-shot does not mean the absence of external models or computational costs. Appendix M measures per-episode costs on two 24 GB RTX 3090 GPUs; model inference and exploration time both contribute to system overhead.

Key Experimental Results

Main Results

CoIN-Bench provides only the category initially. The user simulator sees the target image, whereas the robot receives language answers only. SR is the fraction of successful episodes; SPL weights success by the ratio of shortest to actual path length, measuring navigation efficiency. RL counts total user-simulator response tokens per episode, while NQ averages questions over episodes with interaction. RL is not human word count, and NQ is not an unconditional average over every episode.

SR/SPL below are percentages; each cell lists SR / SPL / RL / NQ. SR and SPL come from Table 2, while RL and NQ come from Table 1. AIUTA* is a reproduction using the same language models, and Pooled is the independent-matching baseline sharing candidate-pool construction.

Method Val Seen Val Seen Synonyms Val Unseen
AIUTA* 10.5 / 4.5 / 109.5 / 1.2 15.3 / 8.4 / 129.2 / 1.2 8.9 / 4.0 / 122.8 / 1.3
Pooled Independent Matching 17.5 / 5.1 / 460.2 / 3.6 22.0 / 8.1 / 519.2 / 3.7 13.3 / 5.1 / 467.7 / 3.4
ProCompNav 23.7 / 7.0 / 4.2 / 2.2 28.1 / 8.5 / 4.3 / 2.2 17.0 / 6.2 / 4.2 / 2.3

On Val Seen, ProCompNav gains 13.2 SR percentage points over AIUTA, reported as approximately 126% relative improvement, and 6.2 points over Pooled. It asks fewer questions than Pooled but more than AIUTA. The benefit is therefore not fewer questions than every baseline, but lower expressive burden through shorter answers and more reliable disambiguation. Non-interactive baselines receiving detailed descriptions and interactive methods receiving categories have different input conditions, which a ranking of Table 2 alone would obscure.

In TextNav Table 3, ProCompNav has SR/SPL of 28.5/6.9, Context-Nav 26.2/9.1, UniGoal 20.2/11.4, and 3D-Mem* 14.1/9.6. ProCompNav has the highest SR but lower SPL than all three. Its SR gain over Context-Nav is 2.3 percentage points, approximately 8.8% relative, not an across-the-board improvement in navigation efficiency.

Ablation Study

The following CoIN-Bench Val Seen ablations come from Table 5; SR/SPL are percentages.

Config SR SPL RL NQ
Full model 23.7 7.0 4.2 2.2
Without multi-view aggregation 23.1 6.3 4.2 2.1
LLM-only attribute selection, no NLI 20.6 5.5 4.2 2.2
KMeans instead of similar core selection 20.5 6.5 4.1 2.2
Without group refinement 21.8 6.0 4.2 2.1
Single-prompt attribute selection and grouping 20.9 5.9 4.7 2.5

Removing core selection and NLI lowers SR by 3.2 and 3.1 percentage points, respectively, among the larger drops in this table. Removing group refinement costs 1.9 points. The multi-view ablation loses only 0.6 points, so it should not be presented as the sole source of the main improvement. Appendix I raises the pool threshold from 4 to 5 to 6: SR is 19.4, 23.7, and 25.5; SPL is 7.1, 7.0, and 6.1; mean steps are 212.7, 257.0, and 299.3. More candidates trade exploration cost for success.

Both human studies in Table 6 and Appendix O are within-subject web studies with 20 participants each, not end-to-end human evaluations of real robot navigation episodes.

Study and condition Median response time per question Participant preference
Study 1: open-ended questions 9.25 s 0/20
Study 1: binary questions 2.69 s 20/20
Study 2: interleaved questions 6.27 s 6/20
Study 2: batched questions 5.28 s 14/20

In Study 1, each participant answers 13 binary and 27 open-ended questions, with both conditions delivered in batches. Results therefore also involve differences in question content and count, rather than a strictly controlled change of response format alone. Study 2 uses three binary questions in each condition, with different waiting intervals simulating interleaved or batched delivery and no navigation required. The response-time difference for batching is not significant (\(p=0.55\)); the one-sided exact binomial test for preference gives \(p=0.058\). The 70% preference is descriptive and should not be called significant superiority over interleaving.

Key Findings

  • Incorrect feedback is more dangerous than “I don’t know.” In Table 4, 20% Flip lowers SR from 23.7 to 17.1, while 20% IDK yields 21.2. Binary pruning is efficient, but a wrong answer directly deletes the target’s group.
  • Human open-ended responses are much shorter than simulator responses: 89.1% of Korean responses contain at most three whitespace-delimited words, with medians of one word and three characters. They are not directly comparable to token-level simulator RL.
  • Most failures arise before final comparison. Of 634 failures, 436 lack the target in the pool, 160 subsequently prune it, and 38 fail during final navigation. Candidate splitting and caption–attribute placement errors account for 72.5% of the 160 pruning failures.
  • Numerical and reference discrepancies are preserved: Table 1 lists Val Seen AIUTA*/Pooled RL as 109.5/460.2 and NQ as 1.2/3.6, whereas the Original column of Table 4 gives 113.7/439.7 and 1.3/4.0. The paper does not explain this difference, so the values are not merged here. The main text and Appendix P refer to the failure-attribution table as Table 4, but it is Table 7 in the current version.

Highlights & Insights

  • A clarification question’s value depends on candidate differences, not target-description length. Replacing “collect facts” with “eliminate a candidate group” could transfer to interactive retrieval or instance selection for robotic grasping, provided the candidate pool covers the target.
  • LLM hypothesis generation and consistent NLI scoring enable structured filtering through natural-language attributes. Group refinement is especially important because geometric or embedding-based grouping does not necessarily coincide with attribute membership.
  • User burden includes interruption timing as well as response length. The paper studies batching separately from binary response format, but evidence is stronger for binary questions, while conclusions about batching require statistical restraint.

Limitations & Future Work

  • Navigation is evaluated only in simulation, without real robot deployment evidence. Web-based human studies also do not replace joint evaluation of movement, observation, and interaction in a real environment.
  • Candidate recall and cross-view merging are major bottlenecks. In Appendix N, only 50.9% of the 379 episodes whose pools contain the target merge its views into one candidate; the remaining 49.1% scatter them across candidates, directly affecting description completeness and question count.
  • The method cannot recover a target removed by incorrect feedback. Soft retention, answer confidence, contradiction detection, and rollback over candidate history are future directions, not existing capabilities.
  • A fixed threshold of 5 candidates may waste exploration when few same-category objects exist. Coverage or candidate uncertainty could instead determine when comparison starts.
  • Failure attribution uses extra target identities and post-hoc vision-language model (VLM) labels for diagnosis, not causal analysis. Each episode receives one primary label, and invisible versus unjudgeable attributes may be conflated. One observed NLI misrouting does not establish that NLI is generally error-free.
  • vs AIUTA: AIUTA interleaves exploration with open-ended clarification and independently matches candidates; ProCompNav collects first, then constructs binary questions through candidate contrasts. It reduces expressive burden and premature commitment, but requires waiting for the pool and makes incorrect binary answers more consequential.
  • vs Pooled Independent Matching: Both share pool construction. Pooled asks questions candidate by candidate, then scores candidates against the complete fact set. This control better isolates RCJ’s contribution and shows that collection alone does not yield the strongest disambiguation.
  • vs Context-Nav / UniGoal: These methods match detailed goals using spatial predicates or graph-structure scores; ProCompNav selects attributes discriminative relative to current candidates. TextNav supports this comparison strategy, but lower SPL shows that the gain is not free.
  • vs 3D-Mem: 3D-Mem builds visual memory during exploration and uses a VLM to select the target; ProCompNav emphasizes explicit attribute differences and binary elimination. Tables 2/3 mark 3D-Mem* as Comp in the Judge column, so a lack of explicit discriminative-attribute extraction should not be misrepresented as no candidate comparison at all.

Rating

  • Novelty: 4/5 — Combines candidate-relative comparison, attribute verification, and binary pruning to address same-category distractors explicitly.
  • Experimental Thoroughness: 4/5 — Includes two tasks, ablations, feedback noise, and human studies, but lacks real robot validation and some interaction conclusions are limited by simulated users and statistical boundaries.
  • Writing Quality: 4/5 — Mechanisms and failure analysis are clear, but table-reference errors and cross-table baseline discrepancies remain.
  • Value: 4/5 — Provides a reusable approach to low-expressive-burden instance navigation, with practical value still dependent on candidate recall, caption fidelity, and recovery from incorrect feedback.