EmbedCopilot: Evaluating Vision-Language Models for Hardware-Aware Embedded System Development¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/X-EASys/EmbedCopilot-Bench
Area: Multimodal VLM
Keywords: Vision-Language Models, Embedded Systems, Multimodal Benchmark, Hardware-Aware Copilot, Execution Success Rate
TL;DR¶
Introduces EmbedCopilot-Bench, the first multimodal benchmark for hardware-aware embedded development, pairing an iterative self-reflective LLM-as-a-Judge with emulator- and hardware-in-the-loop Execution Success Rate (ESR) evaluation to uncover critical pin-level visual grounding bottlenecks in state-of-the-art LVLMs.
Background & Motivation¶
Large language models and vision-language models (LVLMs) have driven transformative progress in general-purpose software programming, algorithmic problem solving, and plot-to-code generation, powering widely adopted developer assistants such as GitHub Copilot, Cursor, and Claude Code. However, embedded system and IoT development fundamentally diverges from pure software programming due to its deep entanglement with physical hardware. Developers do not simply synthesize abstract logic; they must wire breadboards, identify microcontroller general-purpose input/output (GPIO) pin allocations, interpret sensor schematics, and navigate multifaceted toolchain and operating system configurations. Text-only programming copilots remain blind to physical circuit states and are inherently incapable of reconciling hardware wiring discrepancies.
Existing vision-language benchmarks primarily evaluate static image recognition, OCR text transcription, mathematical diagram reasoning, or abstracted electronic circuit diagrams. None comprehensively capture the real-world closed loop where a developer supplies board wiring photographs alongside IDE screenshots and textual requirements to synthesize executable firmware. When building physical systems, models must jointly perform fine-grained pin localization, component identification, operating system UI navigation, and platform-specific firmware synthesis. The lack of an end-to-end multimodal benchmark has left the community without a reliable testbed to diagnose whether LVLMs can truly function as embedded copilots.
To bridge this divide between multimodal physical comprehension and embedded software engineering, this paper constructs a grounded benchmark derived from authentic step-by-step developer tutorials across diverse hardware platforms. Core idea: build EmbedCopilot-Bench from real-world tutorial videos covering 5+ mainstream embedded platforms and 27 peripheral modules, and establish a dual-track evaluation combining a 3-round iterative self-reflective LLM-as-a-Judge (Scoring-Refine) with emulator- and physical hardware-in-the-loop Execution Success Rate (ESR) to systematically expose LVLM bottlenecks in physical pin grounding and firmware execution.
Method¶
Overall Architecture¶
The EmbedCopilot-Bench pipeline systematically translates real-world embedded development workflows into rigorous multimodal diagnostic triplets and evaluates model predictions across both open-ended semantic quality and closed-loop execution fidelity. The workflow comprises three main phases: subtitle-guided task segmentation and data curation from tutorial videos, full-lifecycle benchmark construction across hardware and software dimensions, and hybrid evaluation uniting an iterative CoT LLM judge with an execution-based evaluator.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Multi-Platform Tutorial Videos & Environments<br/>5+ MCU platforms + 27 peripheral sensors"] --> B["Video Collection & Subtitle-Guided Task Segmentation<br/>Whisper transcription + GPT-5 step decoupling"]
B --> C["Multimodal QA Triplet Generation & Annotation<br/>Hardware Operation / Software Config / Code Generation"]
C --> D["Iterative Self-Reflective Judge Scoring-Refine<br/>Initial score → 2 reflection rounds → final reconciliation"]
D --> E["Dual-Criterion Execution Success Rate ESR<br/>Wokwi emulator + physical hardware-in-the-loop"]
E --> F["Output: Semantic Scores & Physical Diagnostics<br/>Strict criterion Estr vs pin-tolerant criterion Eper"]
Key Designs¶
1. Video Collection & Subtitle-Guided Task Segmentation: Decoupling cohesive subtasks from developer streams Authentic embedded tutorials feature instructors who concurrently explain operational intent, manipulate physical breadboards and wiring, and demonstrate computer IDE configurations. Naive temporal slicing disrupts the causal chain of hardware actions. The benchmark systematically collects tutorials across five dominant embedded architectures (Arduino Uno, Arduino Nano, ESP32 series, Raspberry Pi series, and STM32F401RE) paired with 27 diverse peripherals encompassing sensors (DHT11, HC-SR04, MLX90614), actuators (SG90/HS-311 servos, stepper motors), and displays. After transcribing verbal explanations via Whisper, GPT-5 analyzes the temporal subtitle flow to segment continuous videos into self-contained, goal-directed task units while preserving preceding operational summaries as contextual history to prevent ambiguities.
2. Multimodal QA Triplet Generation & Annotation: Covering the full hardware-software stack For each decoupled task, salient frames capturing breadboard wiring close-ups and development toolchain windows are extracted as visual context. GPT-5 generates concise natural language task queries, while transcribed expert instructions serve as ground-truth references. Human annotators subsequently correct Whisper transcription errors, refine ambiguous task queries, and mask explicit textual clues on printed boards to ensure models genuinely infer spatial pin topology rather than relying on trivial OCR shortcuts. The resulting 216 multimodal triplets are explicitly tagged across three core developer capability dimensions: hardware operation (breadboard wiring, pin jumping, polarities), software configuration (driver installation, COM port selection, OS flashing, IDE UI options), and code generation (bare-metal and framework-based C/C++/Python firmware).
3. Iterative Self-Reflective Judge Scoring-Refine: Stabilizing open-ended semantic evaluation Embedded guidance outputs exhibit high heterogeneity, ranging from low-level register manipulations to multi-step IDE dropdown selections, making exact-match string metrics ineffective. To suppress single-pass scoring drift and hallucinated evaluations, the benchmark adopts an iterative self-reflective scoring protocol termed Scoring-Refine spanning \(R=3\) passes. An LLM evaluator \(M\) evaluates model prediction \(\hat{y}\) against question \(q\) and reference \(g\) using a structured rubric covering correctness, grounding, and actionability, producing an initial score and explanation \((s_1, e_1)\). In subsequent rounds, self-reflection prompts \(p_{\text{ref}}\) instruct the judge to revisit its historical rationale, culminating in an aggregation prompt \(p_{\text{fin}}\) that reconciles discrepancies into a final normalized Likert score: $\(s_i = \frac{r_i - 1}{4}, \quad S = \frac{1}{N} \sum_{i=1}^N s_i \times 100\%\)$ This iterative procedure brings the mean absolute error between GPT-5 and human experts down to 0.062, sustaining high inter-judge Spearman rank correlations of 0.832–0.916 across diverse underlying LLM evaluators.
4. Dual-Criterion Execution Success Rate ESR: Disentangling pin localization from functional logic Superficial semantic scoring often conceals fatal firmware bugs that crash upon device deployment. For code-generating tasks, consecutive subtasks contributing to a unified functional milestone are grouped into meta-tasks. The framework instantiates mirroring virtual circuit schematics in the Wokwi emulator and deploys automated test harnesses, complemented by a physical hardware testbed for STM32CubeIDE and Raspberry Pi workflows adjudicated through double-blind human inspection. The evaluation establishes two complementary metrics: Strict (\(E_{\text{str}}\)), which demands exact pin correspondence with the hardware photo alongside successful target behavior, and Permissive (\(E_{\text{per}}\)), which tolerates pin indexing mistakes provided the control logic, timing, and peripheral protocols function flawlessly once the pin constant is corrected. The resulting differential: $\(\Delta_{\text{pin}} = E_{\text{per}} - E_{\text{str}}\)$ serves as a pure diagnostic proxy isolating spatial visual grounding deficits from core algorithmic coding proficiency.
Key Experimental Results¶
Main Results¶
Evaluation spans 4 leading proprietary LVLMs and 6 open-source model families across text-only (T) and multimodal with-image (+I) configurations. Table 1 reports the 5-run average LLM judge scores across dimensions alongside execution success rates.
| Model | Modality | Hardware (LLMhw) | Software (LLMsw) | Code (LLMcode) | Overall (LLMall) | Strict ESR (Estr) | Permissive ESR (Eper) |
|---|---|---|---|---|---|---|---|
| Claude-Sonnet-4.5 | T | 79.7±3.1 | 96.1±1.8 | 84.5±2.8 | 87.4±2.4 | 48.3% | 82.8% |
| Claude-Sonnet-4.5 | +I | 86.5±3.2 | 95.3±2.2 | 83.6±2.5 | 89.6±2.3 | 44.8% | 86.2% |
| GPT-5 | T | 74.5±3.3 | 95.5±2.1 | 77.8±3.0 | 83.6±2.5 | 17.2% | 27.6% |
| GPT-5 | +I | 88.2±3.4 | 96.9±1.4 | 81.6±3.0 | 89.2±2.3 | 31.0% | 58.6% |
| GPT-4o | T | 65.6±3.9 | 93.8±2.2 | 76.9±2.1 | 80.3±2.4 | 37.9% | 62.1% |
| GPT-4o | +I | 84.7±3.9 | 93.3±2.4 | 81.7±2.7 | 87.0±2.6 | 48.3% | 79.3% |
| Gemini-2.5-Pro | T | 71.7±2.9 | 92.7±2.2 | 79.1±2.8 | 82.4±2.4 | 31.0% | 44.8% |
| Gemini-2.5-Pro | +I | 82.7±2.9 | 93.3±1.6 | 82.2±2.7 | 86.6±2.3 | 41.4% | 55.2% |
| Mistral-Small-3-24B | +I | 72.2±3.4 | 88.2±2.4 | 78.5±1.3 | 80.8±2.2 | 37.9% | 82.8% |
| InternVL3-8B | +I | 65.0±4.0 | 83.5±3.3 | 70.5±2.8 | 74.4±3.0 | 13.8% | 31.0% |
| Qwen3-VL-8B | +I | 71.9±3.6 | 77.9±3.5 | 69.1±3.3 | 74.0±3.3 | 6.9% | 24.1% |
| Qwen2.5-VL-7B | +I | 59.0±3.3 | 73.8±4.3 | 66.1±4.2 | 67.3±3.6 | 20.7% | 34.5% |
| Qwen3-VL-4B | +I | 60.8±2.7 | 70.1±2.2 | 62.8±2.5 | 65.6±2.3 | 13.8% | 24.1% |
| Gemma-3-4B | +I | 52.0±3.4 | 71.6±3.6 | 58.9±2.0 | 62.2±2.6 | 6.9% | 17.2% |
| LLaVA-OneVision-7B | +I | 52.6±3.0 | 71.0±4.2 | 55.3±2.6 | 60.6±3.3 | 3.4% | 10.3% |
| Qwen3-VL-2B | +I | 47.1±1.6 | 61.9±1.3 | 53.9±1.0 | 54.9±0.9 | 0.0% | 0.0% |
Ablation & Judge Robustness¶
Ablation and robustness studies evaluate inter-judge agreement by benchmarking 6 alternative open-weight and proprietary models across 5,184 evaluated model outputs against GPT-5.
| Alternative Judge | Open-Weight? | Model-Level Spearman ρ | Sample-Level MAE (0–1) |
|---|---|---|---|
| Qwen3-235B | Yes | 0.909 | 0.145 |
| GLM-4.7 | Yes | 0.916 | 0.152 |
| DeepSeek-V3.2 | Yes | 0.832 | 0.206 |
| Claude-Sonnet-4.5 | No | 0.909 | 0.195 |
| Gemini-2.5-Pro | No | 0.888 | 0.163 |
| GPT-4o | No | 0.916 | 0.108 |
On the 172-sample human gold subset, GPT-5 achieves mean absolute differences of 0.058, 0.048, and 0.080 against three independent human raters (overall mean 0.062). Blind audits by 12 embedded systems experts yielded an explanation-score consistency rating of 4.7/5 (96.1% rated \(\ge 4\)).
Key Findings¶
- Visual context delivers substantial hardware gains: Incorporating board photographs boosts hardware scores markedly for capable models (e.g., GPT-4o's \(LLM_{\text{hw}}\) increases from 65.6 to 84.7, and Qwen2.5-VL-7B from 43.2 to 59.0). However, for smaller open-source models like Gemma-3-4B and LLaVA-OneVision-7B, image input introduces visual noise and fails to yield positive gains.
- Severe chasm between semantic scores and physical execution: While top models achieve high LLM-as-a-Judge code scores exceeding 80%, their strict physical execution success rate (\(E_{\text{str}}\)) tops out at only 48.3% (Claude-Sonnet-4.5 and GPT-4o), confirming that fluent syntax does not equal functional hardware compliance.
- Pin grounding is the primary bottleneck for top-tier models: Under permissive execution (\(E_{\text{per}}\)), Claude-Sonnet-4.5 surges to 86.2% and Mistral-Small-3-24B reaches 82.8%. The enormous Strict–Permissive gap reveals that top models grasp the required peripheral timing and protocol logic but suffer near-misses due to subtle pin misidentifications (e.g., mistaking analog pin A1 for A0 under oblique viewing angles).
- Failure divergence across model tiers: Advanced models are bottlenecked primarily by spatial pin localization, whereas lightweight open-source models remain below 25% even under permissive evaluation, plagued by fundamental failures in hardware control logic and API synthesis.
Highlights & Insights¶
- Execution-based hardware-in-the-loop evaluation: By coupling the Wokwi emulator with physical microcontroller rigs and formalizing the strict vs. permissive ESR metrics, the benchmark eliminates speculative semantic grading and pinpoints spatial physical misalignment with minimal noise.
- Iterative Scoring-Refine judge mechanism: The multi-pass reflection loop mitigates stochastic hallucinations in open-ended technical evaluation, attaining an expert-level consistency score of 4.7/5 and offering a robust template for engineering agent benchmarks.
- Need for structured board topology reasoning: Case studies demonstrate that humans deduce occluded pin numbers by counting relative header positions from fiducial board landmarks, highlighting that future copilots must incorporate board layout graph priors rather than treating circuits as unstructured 2D images.
Limitations & Future Work¶
- Static capture vs. dynamic real-world inspection: The benchmark currently evaluates static key frames; authentic debugging frequently requires active multi-angle camera inspection or interactive probing to resolve obscured jumper wire paths.
- Dataset scale across niche industrial protocols: While spanning 216 meticulously verified triplets across 5 major MCU architectures and 27 peripherals, scaling up to complex industrial buses (CAN, Modbus) and multi-board topologies remains an open challenge.
- Model-side architectural enhancements: Promising directions include fine-grained region-of-interest magnification (RoI-zoom) mechanisms and pinout schematic retrieval-augmented generation (RAG) to ground physical pin coordinates robustly.
Related Work & Insights¶
- vs EmbedBench / IoTPilot / AutoIOT: Prior embedded code generators rely on text prompts or static component knowledge graphs, operating without visual awareness of physical wiring or IDE screens; EmbedCopilot-Bench is the first to evaluate end-to-end multimodal perception linked to executable firmware.
- vs MMCode / Plot2Code / PCB-bench: MMCode and Plot2Code target algorithmic programming diagrams and scientific charts, while PCB-bench focuses on EDA layout design; EmbedCopilot-Bench addresses the physical microcontroller development loop spanning breadboard wiring, pin inference, and on-device execution.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Establishes the first multimodal, hardware-aware embedded development benchmark with realistic physical execution verification.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluation across 10 closed- and open-source models, dual-modality ablations, emulator/physical execution, and extensive human auditor calibrations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear exposition, well-structured taxonomy, and insightful error decomposition distinguishing pin mislocalization from logic failure.
- Value: ⭐⭐⭐⭐⭐ Serves as a vital diagnostic catalyst bridging computer vision, multimodal reasoning, and cyber-physical systems development.