Skip to content

Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

Conference: NeurIPS2026
arXiv: 2609.37655
Code: https://github.com/yingjiayu12/Exemplar2VQA
Area: Multimodal VLM
Keywords: spatial visual question answering, exemplar-driven synthesis, multi-agent coding, geometric utilities, sim-to-real transfer

TL;DR

Exemplar2VQA compiles spatial QA exemplars into reusable simulator capture and geometric annotation programs, using four-role collaboration and execution feedback to reduce generation errors; approximately 10K synthetic examples raise the author-reported average score of Qwen2.5-VL-3B on the multiple-choice subset of VSI-Bench from 35.3 to 42.9.

Background & Motivation

Vision-language models (VLMs) can recognize objects without reliably judging cross-view positional relationships, distances, or occlusion. Manually annotating multi-view spatial QA requires repeatedly checking instance correspondence and geometry, making it expensive. Asking a VLM to generate answers directly instead hands a problem requiring coordinate transformations and instance matching to unreliable language prediction. LLaVA- and ShareGPT4V-style semantic description synthesis does not inherently guarantee geometric consistency for these labels.

Simulators provide explicit object bounding boxes, camera poses, and scene metadata, allowing answers to be computed programmatically. However, fixed rules struggle to accommodate new question types. Switching from object left/right relationships to directions viewed from another object's perspective requires more than changing the wording: reference frames, instance filtering, and viewpoint acquisition logic must also change. Putting everything into one prompt can further entangle task interpretation, code generation, and debugging.

The paper uses a coding model to generate question-type-level programs rather than guess answers individually. Manually engineered geometric APIs handle numerical computation, while simulator and Python execution feedback help repair scripts. Core idea: turn a concrete spatial QA exemplar into executable, reusable capture and annotation logic, then instantiate training examples across simulated scenes, separating linguistic generalization from geometric computation.

Method

Overall Architecture

Inputs comprise a spatial QA exemplar, camera acquisition requirements, and available 3D simulated environments; outputs pair images or videos with questions and answers. The camera track obtains observations satisfying the requirements and their metadata. The QA track reads that metadata and generates and executes Python scripts invoking geometric APIs. Both tracks use Architect, Coder, Reviewer, and Refiner roles, all instantiated from the same Qwen3-Coder-30B-A3B-Instruct model. Specialized system prompts, isolated contexts, and tool access distinguish the rolesโ€”not four different models.

The diagram separates offline label generation, offline supervised fine-tuning (SFT), and test-time inference. The 3D metadata is privileged information used to synthesize labels, not an input to the test-time VLM. The evaluated model must still predict answers from real images or videos and questions.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    I["Capture requirements + simulator"] --> C["Camera-track collaboration<br/>Structured intent to capture"]
    C -->|Execution feedback: Reviewer / Refiner| C
    E["Spatial QA exemplar"] --> Q["QA-track collaboration<br/>Template to annotation program"]
    C -->|Observations and 3D metadata| Q
    Q -->|Execution feedback: Reviewer / Refiner| Q
    Q --> G["Geometric API orchestration<br/>Computation and ambiguity filtering"]
    G --> R["Script reuse and distribution control<br/>Batch label instantiation"]
    R --> D["Synthetic images or videos + QA"]
    D -->|Offline SFT supervision| V["Fine-tuned VLM"]
    T["Test images or videos + question<br/>No 3D metadata"] -->|Inference input| V
    V --> A["Predicted answer"]

Geometric API orchestration is a computation step inside the program generated by the QA track, not a third independent agent track. Feedback loops support retries and repairs but do not guarantee success: the four-role configuration achieves only 92.0% QA-generation success and 84.0% camera-task success.

Key Designs

1. Camera-track collaboration: specify viewpoint intent before executing a checkable capture program

Camera requirements cannot remain a vague instruction to circle an object. CameraArchitect converts natural-language requirements into structured JSON specifying target-object filters, trajectory type, angles, and constraints such as visible pixels. For objects of different sizes, the object-relative camera radius combines a fixed clearance with half the largest bounding-box dimension. This avoids using the same unsuitable radius for both a bed and a small object; tasks without a specific target instance bypass this adjustment.

\[ r(o)=r_{\mathrm{base}}+\frac{1}{2}\max(D_o) \]

Here, \(D_o\) denotes the target bounding box's length, width, and height, and \(r_{\mathrm{base}}\) is a predefined clearance. CameraCoder converts the JSON intent into Python, computes camera poses, and invokes AI2-THOR to obtain multi-view observations. Thus, the Architect produces structured pose intentโ€”not a complete simulator-validated trajectory or the training images themselves.

CameraReviewer diagnoses problems from execution records, errors, and violated constraints; CameraRefiner then patches the navigation code. Collisions, wall-obstructed views, and out-of-bounds rendering require environmental feedback rather than syntax inspection alone. Separating diagnosis from repair gives the repair role focused error descriptions instead of requiring simultaneous revision of task interpretation and geometric implementation. An executable trajectory can nevertheless fail to satisfy the original observation intent.

2. QA-track collaboration: recover cross-scene instantiation logic from one question exemplar

QA Architect receives a concrete exemplar, parameterizes object names, numerical conditions, and image identifiers, and extracts the question format and required comparison. For a question asking an object's direction from another object's perspective, substituting nouns is insufficient: the relationship between the reference viewpoint and target instance must remain intact. The resulting template specifies a generation program rather than inviting unconstrained linguistic paraphrases.

QA Coder writes Python using that specification and the captured scene metadata. The program filters eligible object combinations, invokes geometric functions, and fills in questions and answers. QA Reviewer and QA Refiner address syntax errors, type mismatches, incorrect API use, and logical defects within a Python execution environment. Once checked, the script iterates over candidate scene combinations to generate many QA instances rather than calling an LLM once per QA pair.

The dependency boundary matters: the camera track handles AI2-THOR observations and physical constraints, while the QA track handles corresponding metadata and question logic. An absence of exceptions does not establish label validity. The authors define generation success as mathematically and semantically correct final outputs; the appendix explicitly includes failures where execution succeeds but produces incorrect answers.

3. Geometric API orchestration: move coordinate computation out of language prediction and reject ambiguous relationships

The manually designed API library includes functions for projecting 3D object information onto the XZ plane, computing planar angles, testing directional relations, calculating projected distances, and measuring 3D bounding-box distances between categories. QA Coder organizes call dependencies according to the task. For object-centric directional reasoning, for example, it first obtains planar footprints through get_object_xz_points, then evaluates the relationship through is_spatial_relation_satisfied. A fluent explanation cannot substitute for this computation chain.

Quality control also uses heuristics such as is_angle_ambiguous to reject layouts near directional boundaries, reducing unstable labels such as โ€œleft or front.โ€ The original API example gives an angle-ambiguity threshold of 10 and a directional-relation parameter th=0.15, but does not fully explain the latter's units or implementation. The note therefore does not expand it into an exact geometric predicate formula.

Determinism is conditional on correct code, metadata, and API implementation: a fixed program can compute an explicit result from the same geometric input. This does not make LLM semantic parsing, tool selection, or code writing deterministic. Although the paper uses strong phrases such as eliminating hallucinations entirely, its residual errors and success rates support the narrower conclusion that the framework reduces reliance on implicit spatial guessingโ€”not that every label is guaranteed correct.

4. Script reuse and distribution control: shift cost from per-example calls to question-type-level compilation

A new spatial task initially incurs dual-track planning, code generation, and debugging costs. Once its scripts are usable, switching scenes or object combinations and filling templates largely requires program execution. On two RTX A6000 GPUs, the authors report approximately 3 minutes for camera-track code synthesis and 2 minutes for QA-track synthesis. These figures are not total dataset-generation times and do not include all subsequent scene rendering and instantiation.

Appendix A.6 separates the subsequent costs: camera execution, rendering, and saving observations take about 15 seconds per scene, while QA instantiation across 100 scenes takes about 25 seconds in total. Multi-object questions may use brute-force enumeration with complexity as high as \(\mathcal{O}(M^4)\), where \(M\) is the object count. The authors explain current feasibility through the typically small observable count per room, \(M\approx20\); this does not establish that larger object counts remain inexpensive.

After generating a candidate pool, targeted random sampling by question type can maintain approximately balanced proportions or configure chosen proportions. ProcTHOR and Holodeck are proposed for expanding room counts and layouts, but the validated training scales remain approximately 6K/10K. โ€œInfinite generationโ€ describes an aspiration based on program reuse, not demonstrated infinite independent scenes, semantic coverage, or verified labels.

A Worked Example

The following illustration is constructed from the method, not a particular reported sample or measured trajectory. A user provides a multiple-choice exemplar asking the direction of object B from object A's viewpoint and requests observations around a target. The camera Architect first specifies target filtering, an orbit trajectory, and visibility constraints; the Coder then generates an AI2-THOR script.

If a wall obstructs an observation, the Reviewer diagnoses why that viewpoint violates the requirements, and the Refiner repairs the camera code before another execution. After successful capture, QA Architect parameterizes A, B, and the image identifiers. QA Coder iterates over instance combinations with valid metadata, projects them onto the XZ plane, and invokes the relational API using the reference direction.

A candidate near an ambiguous angular boundary is filtered out. Otherwise, the program inserts object names, options, and the computed answer. The example enters the offline training set; the test-time VLM receives only observations and a question, without calling the scene metadata again to compute its answer. Moving the camera around a static scene can produce a video input, but does not create supervision for continuous object motion or dynamic tracking.

Loss & Training

The paper introduces neither a new VLM architecture nor a dedicated loss; its main contribution is data generation. Approximately 6K MMSI-format examples train Qwen2.5-VL-7B through LoRA and a full-SFT comparison. Approximately 10K VSI-format examples train Qwen2.5-VL-3B through full SFT. The former uses multiple images, whereas the latter organizes multi-view observations into continuous rotation videos. Question types and output formats mirror the target benchmarks, so the primary experiments cannot all be described as tests on unseen question types.

Appendix A.1 specifies 3B full SFT for 3 epochs at learning rate 1e-5, with per-device batch 4 and gradient accumulation 2. The 7B LoRA configuration uses 3 epochs, learning rate 1e-4, rank 16, alpha 32, per-device batch 2, and gradient accumulation 4. Both have effective batch 32 across four H200 GPUs and use cosine scheduling, warmup ratio 0.05, weight decay 0.01, DeepSpeed ZeRO-2, and NEFTune noise alpha 5. The maximum 7B image resolution is 1,003,520 pixels. The appendix does not provide equally detailed dedicated hyperparameters for 7B full SFT, so its learning rate cannot simply be inferred from LoRA.

Zero-shot transfer to SpaCE-10, ViewSpatial-Bench, and Spatial-Obj uses the 7B model fine-tuned only on MMSI synthetic data. OSR-Bench instead receives another approximately 6K-example dataset adapted to its exemplars and a separate 7B LoRA training run. It is a different adaptation experiment, not zero-shot evaluation of the same model on every benchmark. Some generated QA pairs were also manually spot-checked before training. Manual API engineering, exemplar provision, and spot-checking should not be obscured by the absence of per-example manual annotation.

There is an internal output-format inconsistency: appendix prose says MMSI outputs only an answer letter, while its training and evaluation templates actually show <answer>A. Above</answer>, including the letter and text. This conflict is preserved rather than claiming that the authors implemented either protocol unambiguously.

Key Experimental Results

Main Results

The following table records scores from main-text Tables 1โ€“4. Gains are absolute percentage points relative to the same-sized base model. VSI's 42.9 covers only four multiple-choice tasks, not the full VSI-Bench. Avg./Overall retain the authors' aggregates rather than being recalculated from displayed subcategories.

Dataset and scope Training and evaluation configuration Base Ours Gain
MMSI-Bench Overall Approximately 6K; 7B LoRA 25.9 28.2 +2.3
MMSI-Bench Overall Same approximately 6K; 7B full SFT 25.9 28.0 +2.1
VSI-Bench multiple-choice Avg. Approximately 10K; 3B full SFT 35.3 42.9 +7.6
SpaCE-10 Overall 7B trained on MMSI synthetic data; zero-shot 33.3 42.2 +8.9
ViewSpatial-Bench Overall 7B trained on MMSI synthetic data; zero-shot 36.9 45.4 +8.5

On MMSI, LoRA improves Cam.-Cam. from 24.7 to 33.3, a gain of +8.6 points, but the dynamic Obj. category without directly synthesized supervision drops from 39.5 to 25.0. Another unsupported category, semantic projection Appr., rises from 18.2 to 27.3; this does not establish generator support for that semantic task. VSI's untrained Route Plan* rises from 28.9 to 35.1, representing zero-shot subtask transfer after static geometric supervision.

Appendix A.3 adds important boundaries: dedicated OSR adaptation raises Avg. from 28.3 to 36.5 using four sparse surround-view training images and panoramic test inputs. Strictly zero-shot Spatial-Obj Overall rises only from 69.71 to 70.50. Fine-tuning InternVL2 on the same synthetic data raises VSI Avg. from 28.2 to 35.6 for 2B and from 36.7 to 46.1 for 8B. This supports the value of the data for other backbones, not equal gains across all tasks.

Ablation Study

Table 5 evaluates 100 seed exemplars per configuration. Success requires mathematically and semantically correct final outputs. Merged roles share context; separated roles have independent contexts. The main topology comparison uses the Qwen3-Coder-30B generator.

Config QA generation success (%) Camera trajectory success (%)
1 agent: all roles merged 34.0 60.0
2 agents: Architect; Coder+Reviewer+Refiner 72.0 74.0
3 agents: Architect; Coder; Reviewer+Refiner 85.0 80.0
4 agents: all four roles separated 92.0 84.0
4 agents: replacing generator with Qwen3.5-9B 48.0 56.0

Separating planning first increases QA success by 38.0 points, and the full four-role system gains 58.0 points over the monolithic configuration. The corresponding total camera-track gain is 24.0 points. The smaller-model comparison changes both scale and model family/coding specialization. It shows that generator replacement affects reliability, but does not isolate parameter count as a causal factor.

Table 6 uses the model named Gemini3.5-Flash by the authors to directly annotate 10K synthetic VSI multiple-choice training examples, then fine-tunes the same 3B base model under identical settings. The name is retained as reported; its specific model version was not independently verified.

Annotation/training configuration Rel. Dist Rel. Dir Route Plan Appr. Order Avg.
Qwen2.5-VL-3B base 34.7 42.6 28.9 35.0 35.3
Fine-tuned on direct LLM annotation 31.7 40.4 34.5 6.3 28.2
Fine-tuned on Exemplar2VQA program annotation 43.8 52.1 35.1 40.5 42.9

Key Findings

  • Direct annotation lowers Avg. by 7.1 points relative to the base model; program annotation exceeds direct annotation by 14.7 points. In particular, Appr. Order at 6.3 demonstrates that more synthetic labels do not automatically provide better supervision.
  • Appendix A.5 percentages describe the composition of failures: absent expected QA pairs fall from 47.0% to 37.5% of QA failures, and execution failures from 30.3% to 25.0%. Semantic misalignment accounts for 62.5% of remaining camera failures. These are not error rates over all 100 seeds.
  • SpaCE-10 Scene Quant. falls from 36.9 to 24.2, showing that instance bounding-box supervision need not help abstract functional-zone grouping. Overall cross-domain gains do not imply improvements in every spatial semantic category.
  • Appendix Table A4 reports MMSI Overall of \(29.33\pm0.99\) across three training seeds, but checklist item 7 says multiple-seed experiments were not conducted. These statements directly conflict, and the correspondence between A4 and the main-table 28.2/28.0 configurations is insufficiently clear. They cannot be combined as one statistical result.
  • Appendix Table A5 separately reports VSI's four numerical tasks improving in Avg. from 22.5 to 37.4. The main text nevertheless limits its primary comparison to multiple-choice tasks, and the extension lacks equally detailed dedicated training information. Its results are not combined with 42.9 into a full-benchmark claim.

Highlights & Insights

  • An exemplar constrains reference frames, instance filtering, and viewpoint requirementsโ€”not just language style. Compiling these into programs enables new question types to reuse the geometric foundation without paying model-inference costs per example.
  • Role isolation and geometric utilities address distinct problems: contextual interference and unreliable numerical computation. More agents without trustworthy computation can still repeat spatial mistakes; tools without clear semantics can still execute the wrong question.
  • A transferable direction is generating targeted data for weak spatial relations while controlling category proportions. Curriculum learning and arbitrary distribution control are presented mainly as mechanisms and future directions in the appendix, not completed curriculum-learning ablations.

Limitations & Future Work

  • The current generator targets static, object-centric geometry, not continuous object motion, complex temporal semantics, or semantic shape projection. Camera video inputs and improved Route Plan* transfer do not substitute for dynamic supervision coverage.
  • Synthetic labels depend on accurate simulator metadata and manually engineered APIs. Generating labels directly from real photographs would require upstream reconstruction or perception models, reintroducing their errors into annotation. Successful execution also does not rule out misunderstood templates.
  • Success rates of 92.0%/84.0% and remaining semantic failures bound generation reliability. Question-type-level semantic tests, geometric consistency assertions, and more systematic human audits are potential improvements, but the paper does not quantify them.
  • Room counts, object categories, and layout diversity limit effective data scale, while enumeration complexity constrains throughput in large scenes. โ€œInfinite dataโ€ and โ€œzero human laborโ€ both exceed what the experiments establish.
  • Generalization evidence consists of results on selected spatial VQA benchmarks, not robot navigation, manipulation, or physical deployment success. Statistical definitions, prompt protocols, and main-text/appendix conflicts require clarification through reproduction.
  • vs CLEVR-style fixed-rule synthesis: exemplars guide new task-program generation, improving question-type flexibility. Code and semantic correctness become review obligations rather than properties inherently guaranteed by fixed templates.
  • vs LLaVA / ShareGPT4V-style model annotation: these works motivate multimodal synthetic supervision, whereas this paper assigns spatial answers to geometric programs over metadata. The advantage depends on trustworthy scene information and does not make the framework a reliable annotator for arbitrary real images.
  • vs ViperGPT / PAL: all share the idea that models organize programs and programs perform computation. Here, however, the principal use is offline training-label generation; downstream test-time VLMs do not require program reasoning or privileged 3D metadata.
  • vs Code as Policies / Voyager: environmental feedback improves code in both settings, but this paper ultimately targets static spatial QA training data rather than real robot-control policies. Multimodal VLM is therefore a better classification than robot control.

Rating

  • Novelty: 4/5 โ€” Combines exemplar specifications, dual-track four-role collaboration, and geometric APIs into a reusable data generator.
  • Experimental Thoroughness: 3/5 โ€” Includes multiple benchmarks, topology ablations, and direct-annotation comparisons, but statistical and extended-training details have gaps.
  • Writing Quality: 3/5 โ€” The main mechanism is clear, but table references, output protocols, and random-seed declarations contain internal inconsistencies.
  • Value: 4/5 โ€” Offers a reusable approach to spatial supervision production, conditional on trustworthy metadata and effective semantic review.