ICLAgent: Integrated Circuit Footprint Geometry Labeling via LMM-empowered Multi-Agent Framework¶
Conference: ECCV 2026
Paper: ECCV Paper
Code: https://github.com/IC-LMM/ICLAgent
Area: Multi-Agent / Multimodal VLM
Keywords: IC Footprint, Multi-Agent Framework, Geometry Parameter Extraction, Dynamic Planning Reasoning, Large Multimodal Models
TL;DR¶
ICLAgent mirrors the cognitive workflow of PCB engineers by constructing a four-stage multi-agent framework encompassing diagram detection, topology planning, and parameter derivation, elevating IC footprint labeling accuracy to 79.0% IoU on a lightweight Qwen2-VL-7B backbone.
Background & Motivation¶
In printed circuit board (PCB) design and electronic manufacturing pipelines, engineers must extract pin footprint geometries (commonly designated as "Land Patterns" or "Suggested Pads") from component datasheets, identifying total pin counts, precise spatial coordinates, and physical dimensions for every pad. This process is essential for ensuring correct component placement and reliable electrical connectivity. However, modern component libraries are colossalโDigi-Key alone catalogues over 13 million electronic productsโand the industry still relies heavily on manual visual inspection of PDF drawings and manual data entry into electronic design automation (EDA) software, which is tedious, error-prone, and a major bottleneck for rapid prototyping.
Traditional EDA tools lack high-level visual understanding of datasheet schematics. Meanwhile, conventional optical character recognition (OCR) and generic object detection models can extract isolated text tokens or count visible bounding boxes, but fail to decipher the implicit geometric constraints inherent to engineering drawings. For instance, the center-to-center pitch between two opposite pin rows is rarely annotated directly, requiring human engineers to calculate it algebraically from inner and outer boundary dimensions. Recent end-to-end multimodal systems like LLM4-IC8K attempt to feed entire PDF pages into fine-tuned models to predict raw pin coordinates in a single shot. However, treating the task as a black-box regression makes models prone to hallucinations and shortcut learning, lacks engineering interpretability, and suffers severe coordinate drift on dense, high-pin-count symmetrical packages.
This paper's angle of attack is that human PCB engineers never solve footprint extraction in a single blind leap. Instead, they follow an explicit, disciplined multi-step reasoning progression: locating the relevant diagram on the page, categorizing the symmetrical layout to define an extraction checklist, calculating explicit and derived geometric parameters, and finally invoking CAD scripts to generate deterministic coordinates. Core idea: explicitly decouple and emulate the expert PCB engineer's cognitive workflow using a specialized multi-agent system comprising diagram detection, topological planning, and parameter reasoning agents, transforming brittle end-to-end regression into an interpretable and controllable parameter planning and derivation pipeline.
Method¶
Overall Architecture¶
ICLAgent structures the IC footprint labeling pipeline into four tightly coupled stages. The first three stages are handled by specialized LMM agents fine-tuned for their respective subtasks, while the final stage executes deterministic geometry generation via lightweight scripts: 1. Diagram Region Detection (Stage 1): Accurately identifies the normalized bounding box \((x, y, dx, dy)\) of the target IC footprint diagram within a complete PDF datasheet page, filtering out surrounding text, electrical rating tables, and irrelevant characteristic curves; 2. Footprint Classification & Parameter Planning (Stage 2): Classifies the package arrangement into one of four topological categories based on pin symmetry: 2-sides (e.g., SOP/SOIC), 4-sides (e.g., QFP/QFN), grid (e.g., BGA/LGA), or other (e.g., SOT/TO). It then evaluates a category-specific parameter checklist against visible annotations to dynamically formulate procedural derivation instructions for missing parameters; 3. Dynamic Parameter Reasoning (Stage 3): Aligns numeric labels with graphic dimension lines to extract explicit geometric values, and executes algebraic derivations for implicit parameters (such as pitch centers or array spacings) following the planning instructions; 4. Description Generation (Stage 4): Executes standardized CAD generation scripts taking structured parameters as input to compute the deterministic index, center coordinates, and bounding dimensions of all pins.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Datasheet PDF Page Input"] --> B["Stage 1: Diagram Region Detection<br/>YOLOv12 Candidate Proposal & Cross-modal Diagram Agent Fusion"]
B --> C["Stage 2: Footprint Classification & Parameter Planning<br/>4-way Topology Categorization & Missing Parameter Instruction Generation"]
C --> D["Stage 3: Dynamic Parameter Reasoning<br/>Explicit Label Extraction & Implicit Geometry Algebraic Derivation"]
D --> E["Stage 4: Description Generation<br/>Parameter-driven EDA Rule Scripts for Coordinate and Dimension Computation"]
E --> F["Output Standardized Pin Annotations<br/>Pin Count, Coordinates List, Pad Dimensions"]
Key Designs¶
1. Dual-Path Diagram Region Detection with Cross-Modal Fusion: Eliminating Dense Layout Interference
Component datasheets pack functional block diagrams, pinout schematics, timing waveforms, and mechanical package drawings onto single pages. Feeding entire dense pages directly into multimodal reasoning models causes severe visual attention dispersion. The Diagram Agent addresses this through a dual-path cross-validation mechanism: a specialized lightweight detector (YOLOv12) first scans the entire page to propose candidate regions of interest (ROIs) with associated confidence scores; the Diagram Agent then evaluates semantic context and spatial layout to cross-validate and re-weight candidates, fusing them into the final bounding box \((x, y, dx, dy)\). Cropping to this localized high-resolution region eliminates distraction from surrounding tabular text and irrelevant graphics before subsequent reasoning begins.
2. Topology Classification and Dynamic Parameter Planning: Expert-Scaffolded Structured Reasoning
Over 80% of real-world IC packages follow strictly regular geometric topologies. The Planning Agent categorizes footprints into dual-side packages (2-sides, 41.3%), quad-side packages (4-sides, 31.7%), grid array packages (grid, 7.3%), or asymmetric low-pin devices (other, 19.6%). For the three regular classes, the agent accesses an expert-defined checklist of potential parameters and identifies which standard measurements are absent. When an essential value (e.g., center-to-center row spacing) is omitted in favor of inner and outer edge bounds (\(D_{in}\) and \(D_{out}\)), the Planning Agent formulates an explicit chain-of-thought (CoT) derivation directive:
This planning step transforms unconstrained pixel regression into a bounded, verifiable parameter extraction plan that guides the downstream extractor.
3. Numeric-Symbol Alignment and Algebraic Derivation: Resolving Implicit Geometric Relationships
Given the cropped high-resolution diagram and the procedural planning prompt, the Parameter Agent binds callout lines and dimension ticks to their corresponding numeric values. It records explicit parameters (such as pin counts per edge, pad width \(dx\), and pad height \(dy\)) and performs multi-step arithmetic calculations to populate any missing standard parameters specified by the Planning Agent. For the irregular "other" category (< 20 pins), the agent switches to direct pin-level geometric regression, whereas for grid-array packages (e.g., BGAs with depopulated pin corners), it leverages matrix mask rules to adjust overall pin counts and coordinates accurately.
4. Rule-Driven Deterministic Description Generation Tools: Halting Token-Level Hallucination Accumulation
Requiring an autoregressive multimodal model to output hundreds of floating-point coordinates across lengthy token sequences inevitably leads to numeric drift, dropped signs, and format corruption. ICLAgent enforces a clear division of labor: large models handle high-level semantic perception and algebraic planning, while deterministic code handles coordinate expansion. Stage 4 invokes dedicated Python generator scripts (Tool1, Tool2, or Tool3 corresponding to 2-sides, 4-sides, and grid types) that take the compact parameter tuple from Stage 3 and compute exact Cartesian coordinates using standard analytic geometry formulas. This guarantees perfect mathematical consistency and zero formatting errors across large pin arrays.
A Worked Example¶
Consider a standard 48-pin dual-side (2-sides) SOIC package:
1. Stage 1: The Diagram Agent filters out electrical specifications on the upper page, isolating the footprint pad diagram in the bottom-right quadrant with bounding box \((x=0.55, y=0.62, dx=0.40, dy=0.32)\);
2. Stage 2: The Planning Agent observes symmetrical rows on two parallel sides, classifies the type as 2-sides, and inspects the 2-sides checklist. It identifies explicit annotations: row=2, column=24, column pitch 0.5mm, pad dimensions dx=0.3mm, dy=2.3mm. Detecting that center row spacing is unannotated while inner edge distance 4.6mm and outer edge distance 9.2mm are given, it issues the instruction: "Compute row spacing as \((4.6+9.2)/2\)";
3. Stage 3: The Parameter Agent reads the values and computes the derived spacing, outputting the standardized dictionary: {row: 2, column: 24, row_spacing: 6.9, column_spacing: 0.5, dx: 0.3, dy: 2.3};
4. Stage 4: Invoking Tool1, the deterministic script calculates coordinates for all 48 pins symmetrically around the origin (e.g., Pin 1 at \([-5.7, -3.4]\), Pin 2 at \([-5.2, -3.4]\), ..., Pin 48 at \([-5.7, 3.4]\)) and assigns dimensions, directly populating a standard EDA library entry.
Loss & Training¶
To train the specialized agents, the authors curated ICAgent-Instruct, derived and cleaned from ICGeo8K to yield 3,737 verified real-world samples (3,337 training, 400 held-out test). All three agents are instantiated from the open-source Qwen2-VL-7B foundation model and fine-tuned independently using supervised fine-tuning (SFT) with chain-of-thought (CoT) prompts. - The training objective is standard cross-entropy over target response tokens: $\(\mathcal{L}_{SFT} = -\sum_{t=1}^{T} \log P(y_t \mid y_{<t}, X_{img}, X_{prompt})\)$ - Resource footprint is remarkably modest: the entire multi-agent system fine-tunes in just 5 hours on 2 NVIDIA A100-40GB GPUs, and achieves an average inference latency of 13.5 seconds per datasheet page.
Key Experimental Results¶
Main Results¶
On the 400-sample ICGeoQA benchmark, models are evaluated across three core tasks: Task 1 (pin count error via MAE and RMSE), Task 2 (pin coordinate error \(d_{pin}\) in mm), Task 3 (pin dimension overlap \(IoU_{pin}\)), and the comprehensive metric \(IoU_{IC}\) representing overall footprint match.
| Method | Model Type / Scale | Overall (\(IoU_{IC}\) %) โ | Task 1 (MAE / RMSE) โ | Task 2 (\(d_{pin}\)) โ | Task 3 (\(IoU_{pin}\) %) โ |
|---|---|---|---|---|---|
| Manual EDA (Industry Baseline) | Human Manual Entry | 44.0 | - / - | 2.98 | 58.0 |
| Claude Sonnet 4 | Proprietary Frontier LMM | 10.9 ยฑ 0.7 | 2.03 / 12.55 | 3.62 ยฑ 0.23 | 12.4 ยฑ 0.4 |
| GPT-4o | Proprietary Frontier LMM | 11.1 ยฑ 0.4 | 8.21 / 23.04 | 4.01 ยฑ 0.02 | 45.6 ยฑ 0.3 |
| GPT-5 | Proprietary Frontier LMM | 40.6 ยฑ 1.1 | 0.59 / 4.13 | 4.18 ยฑ 0.22 | 65.4 ยฑ 0.1 |
| Gemini 2.5 Pro | Proprietary Frontier LMM | 55.7 ยฑ 0.7 | 0.57 / 4.78 | 5.58 ยฑ 0.48 | 70.1 ยฑ 0.3 |
| Gemini 3 Pro | Proprietary Frontier LMM | 68.6 ยฑ 0.6 | 0.44 / 3.46 | 3.56 ยฑ 0.20 | 82.5 ยฑ 0.3 |
| Qwen3-VL-235B | Open-source Massive LMM | 9.8 ยฑ 0.2 | 0.46 / 3.79 | 5.58 ยฑ 0.17 | 44.3 ยฑ 0.2 |
| LLM4-IC8K (Prior SOTA) | Domain-Tuned Single LMM | 71.6 ยฑ 0.5 | 0.35 ยฑ 0.07 / 2.81 ยฑ 0.08 | 1.11 ยฑ 0.02 | 88.0 ยฑ 0.3 |
| ICLAgent (Ours) | Specialized Multi-Agent (7B) | 79.0 ยฑ 0.2 | 0.37 ยฑ 0.05 / 3.25 ยฑ 0.14 | 1.07 ยฑ 0.03 | 95.6 ยฑ 0.2 |
Ablation Study¶
The authors ablate each stage of the reasoning workflow on ICGeoQA (Table 3), demonstrating the necessity of the multi-agent decomposition:
| Configuration | \(IoU_{IC}\) (%) โ | Relative Change | Observations & Analysis |
|---|---|---|---|
| Full stages (ICLAgent) | 79.0 ยฑ 0.2 | Baseline | Full four-stage collaborative pipeline |
| w/o Stage 1 (no region detection) | 72.3 ยฑ 0.3 | -6.7% (-8.5% rel.) | Background clutter and text distract downstream parameter extraction |
| w/o Stage 2 (no planning & classification) | 70.2 ยฑ 0.2 | -8.8% (-11.1% rel.) | Lacks structured parameter checklist, missing implicit parameter derivations |
| w/o Stage 1 & Stage 2 (Stage 3 only) | 68.2 ยฑ 0.4 | -10.8% (-13.7% rel.) | Single agent directly extracting parameters from raw pages suffers severe omissions |
| End-to-end regression baseline | 65.7 ยฑ 0.1 | -13.3% (-16.8% rel.) | Black-box regression directly to pin coordinates fails on dense layouts |
A fine-grained breakdown by package arrangement (Table 4) illustrates performance dynamics across geometric categories:
| Package Type | Sample Share (%) | Stage 1 (\(IoU_{diag}\) %) | Stage 2 (KPA %) | Stage 3 (PC %) | Overall (\(IoU_{IC}\) %) | Coordinate Error \(d_{pin}\) โ | Dimension \(IoU_{pin}\) % โ |
|---|---|---|---|---|---|---|---|
| 2-sides | 41.3 | 86.0 | 95.6 | 81.4 | 88.2 | 0.55 | 95.9 |
| 4-sides | 31.7 | 58.8 | 98.3 | 80.4 | 73.2 | 0.38 | 95.6 |
| grid | 7.3 | 97.2 | 97.8 | 89.3 | 84.4 | 0.18 | 98.0 |
| other | 19.6 | 45.7 | - | - | 68.3 | 3.50 | 94.5 |
Key Findings¶
- Decomposed domain workflows trump raw foundation model scaling: The 7B-parameter ICLAgent system achieves 79.0% IoU, outperforming general-purpose behemoths like GPT-5 (40.6%) and Gemini 3 Pro (68.6%) by large margins. Aligning agent interactions with domain engineering steps provides far stronger inductive bias than unconstrained model scaling;
- Topological planning acts as a critical hallucination barrier: Ablating Stage 2 planning incurs the single largest drop (-8.8% absolute IoU). The structured checklist forces the model to verify each geometric attribute explicitly rather than skipping implicit constraints;
- Symmetry-based generation yields sub-millimeter precision: On structured packages (2-sides, 4-sides, grid), script-based generation achieves pad dimension overlaps exceeding 95% and grid coordinate errors down to 0.18 mm. Conversely, on irregular "other" packages where the system falls back to direct regression, coordinate error climbs to 3.50 mm, highlighting the decisive advantage of structured generation over direct token regression.
Highlights & Insights¶
- Reframing dense spatial regression into discrete parameter extraction and code execution: Instead of forcing neural networks to memorize dozens of floating-point tuples, the framework restricts the LMM to perceiving graphic callouts and high-level algebra, delegating coordinate expansion to deterministic Python functions;
- High-fidelity emulation of human engineering practice: The four-stage division directly corresponds to how human PCB layout designers operate (find drawing, classify package, calculate missing pitches, run CAD script), making model outputs transparent, modular, and easy to audit;
- Accessible training and deployment footprint: Requiring only 5 hours on two A100 GPUs and running in 13.5 seconds per page, the pipeline offers immediate practical utility for semiconductor distributors and EDA vendors looking to digitize component libraries at scale.
Limitations & Future Work¶
- Unidirectional cascading error vulnerability: Because the pipeline is strictly sequential, errors in early stages propagate downstream. If Stage 1 produces a clipped bounding box (such as the 58.8% detection IoU on 4-sides packages) or Stage 2 swaps \(dx\) and \(dy\), the final CAD generation fails completely; incorporating verification loops or ReAct-style self-correction would improve robustness;
- Degraded performance on irregular and asymmetric packages: For SOT, TO, and arbitrary irregular pinouts ("other" category), the system lacks formal topological grammars and reverts to direct regression, leading to elevated coordinate error (3.50 mm);
- Reliance on hardcoded package categories: The four predefined layout classes cover common chips but may fail on novel 3D heterogeneously integrated chiplets or hybrid multi-die packages, requiring future research into open-vocabulary geometry graph synthesis.
Related Work & Insights¶
- vs LLM4-IC8K: While LLM4-IC8K pioneered the real-world ICGeo8K benchmark, its end-to-end black-box fine-tuning struggles with dense pin counting and implicit spatial relations; ICLAgent improves overall IoU by 10.3% on identical data while providing interpretable intermediate reasoning traces;
- vs EDA Copilots (e.g., PCBAgent, ChipGPT): Most existing EDA agent systems target high-level macro tasks like HDL logic synthesis, floorplanning, or chatbot assistants, neglecting the foundational micro-level component geometry modeling; ICLAgent fills this critical gap at the very front of the PCB manufacturing pipeline.
Rating¶
- Novelty: โญโญโญโญโ [Pioneering multi-agent planning and symbolic tool execution applied to technical IC footprint labeling]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive comparisons across proprietary LMMs, open-source models, domain baselines, and human benchmarks with complete stage ablations]
- Writing Quality: โญโญโญโญโญ [Exemplary clarity, well-structured figures and mathematical definitions, with transparent engineering motivations]
- Value: โญโญโญโญโญ [High industrial impact for automating EDA footprint generation, backed by released datasets and reproducible code]