RobustRDP: Advancing Reaction Diagram Parsing via Synthetic-to-Real Data Scaling and Robustness-Oriented Training¶
Conference: ECCV 2026
Paper: ECCV Paper
Code: https://github.com/jaydetang/RobustRDP
Area: Multimodal VLM
Keywords: Reaction Diagram Parsing, Vision-Language Model, Synthetic-to-Real Data Scaling, Prefix Perturbation, Direct Preference Optimization
TL;DR¶
RobustRDP addresses two core bottlenecks in chemical reaction diagram parsing—expert annotation scarcity and autoregressive error accumulation—by introducing a "sample-render-arrange" synthetic-to-real data scaling pipeline alongside a three-stage progressive training strategy that integrates region guidance, prefix perturbation, and Direct Preference Optimization (DPO), significantly surpassing previous SOTA baselines.
Background & Motivation¶
Chemical reaction diagram parsing aims to automatically convert complex chemical reaction pathways from literature and patent images into machine-readable structured representations, accurately identifying reactants, conditions, and products for each reaction. This capability serves as an indispensable cornerstone for constructing massive chemical databases, empowering computer-aided synthesis planning, and accelerating automated drug discovery. However, traditional cascaded pipelines rely heavily on separated object detection modules and hand-crafted geometric heuristic rules, suffering severe performance degradation caused by cross-stage error accumulation and brittle generalization across diverse publication layouts. While recent end-to-end approaches reformulate object localization and relation extraction into unified coordinate sequence generation using multimodal large language models (MLLMs), they remain hindered by severe data scarcity and generation instability.
Two primary bottlenecks impede progress in this domain. First, chemical reaction diagrams encompass heterogeneous visual elements—molecular structures, text labels, directional arrows, and intricate topological graphs. Due to high expert annotation costs, existing models predominantly depend on the RxnScribe benchmark containing merely 1,240 training diagrams, which is vastly insufficient for training data-hungry MLLMs. Second, end-to-end autoregressive coordinate generation exhibits complex sequential dependencies across successive reactions. Causal attention induces shortcut learning, wherein the model exploits preceding reaction outputs (e.g., prior products frequently serving as current reactants) instead of grounding directly on local visual features. Moreover, because training exclusively exposes the model to ground-truth prefixes (oracle prefix), minor inference errors in early tokens inevitably propagate and trigger catastrophic cascading sequence collapses.
To resolve these tensions, this paper tackles data deficiency and sequential fragility simultaneously: constructing procedural synthesis and semi-automated annotation pipelines to expand training resources, while decoupling temporal dependencies through structural spatial prompting, noisy prefix exposure, and explicit failure suppression. Core idea: develop a "sample-render-arrange" procedural synthesizer and a semi-automated detection labeling platform for synthetic-to-real data scaling, and introduce a three-stage progressive robustness-oriented training framework combining region-guided multi-task learning, prefix perturbation, and DPO failure mode suppression.
Method¶
Overall Architecture¶
RobustRDP adopts Qwen2.5-VL-3B-Instruct as its multimodal foundational backbone, augmenting its vocabulary with dedicated structural marker tokens—<rxn>, <rct>, <cnd>, <prd>, <mol>, and <txt>—to serialize two-dimensional diagram layouts into structured coordinate token streams. The system rests on two synergistic pillars: a synthetic-to-real data scaling pipeline that supplies diverse layout topologies for pretraining and fine-tuning, and a three-stage robustness-oriented training strategy ("Synthetic Pretraining \(\to\) Multi-Task SFT \(\to\) Preference Alignment DPO") that progressively fortifies the model's visual grounding and error recovery capabilities.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Reaction Diagram<br/>Single/Multi-line/Tree/Graph layouts"] --> B["Synthetic-to-Real Data Scaling<br/>Layout-driven synthesizer + semi-automated platform"]
B --> C["Pretraining Stage<br/>60k synthetic diagrams for structural markers & grounding"]
C --> D["Robustness-Oriented Multi-Task SFT<br/>VRP primary task + Region-guided RGRP + Prefix-perturbed PPRP"]
D --> E["Failure Mode Direct Preference Optimization<br/>DPO explicitly penalizing low-F1 erroneous sequences"]
E --> F["Structured Reaction Sequence<br/>Reactant/condition/product coordinates & entity categories"]
Key Designs¶
1. Synthetic-to-Real Data Scaling Scheme: Breaking Literature Annotation Bottlenecks
To overcome severe data scarcity driven by high expert labeling costs, this design combines procedural layout synthesis with semi-automated real-image annotation. The layout-driven synthesizer executes a "sample-render-arrange" workflow: molecular structures are sampled from PubChem, while reaction conditions (reagents, catalysts, temperatures) are sampled from curated text corpora; sampled molecules are rendered via Indigo with coordinate scaling and node shuffling, while text elements adopt diverse publication fonts and styles. Elements are subsequently arranged on canvases under directional arrows across four authentic topological layouts: Single-line (linear reactions), Multiple-line (wrapping pathways with row-spanning intermediates), Tree (convergent/divergent branches), and Graph (dense reaction networks), producing 60,000 synthetic diagrams with accurate bounding box and role annotations. To bridge the domain gap between synthetic visuals and real journal prints, an efficient annotation platform integrated with YOLO26 is developed; the detector generates candidate boxes so annotators only verify boxes and link chemical relationships, slashing manual labeling time by ~40% and yielding 3,000 newly annotated authentic diagrams to form enlarged training and test sets.
2. Robustness-Oriented Multi-Task SFT: Mitigating Shortcut Learning and Cascading Errors
Standard autoregressive next-token prediction causes the model to over-rely on ground-truth preceding outputs \(R_{<i}\) as shortcuts to infer \(R_i\), neglecting visual grounding on the image \(I\). Furthermore, exposure only to perfect training prefixes leaves the model incapable of recovering from its own prediction errors during test-time autoregression. To neutralize these vulnerabilities, multi-task SFT integrates two auxiliary objectives alongside Vanilla Reaction Parsing (VRP). First, Region-Guided Reaction Parsing (RGRP) feeds a local bounding box \(B_{roi}\) enclosing a single target reaction \(R_{roi}\) alongside image \(I\), forcing the model to decode solely the isolated reaction without contextual history: $$ \mathcal{L}{RGRP} = -\log P(R, I) $$ Second, Prefix-Perturbed Reaction Parsing (PPRP) introduces simulated inference errors—random coordinate jitter, entity omission, and redundant entity injection—into a subset of prefix reactions } \mid B_{roi\(\mathcal{P}\) to create a noisy prefix \(\tilde{R}_{<i}\). The model is supervised to accurately predict subsequent unperturbed reactions: $$ \mathcal{L}{PPRP} = -\sum, I) $$ Jointly training on VRP, RGRP, and PPRP simultaneously bolsters standalone local parsing accuracy and endows the generator with prefix error resilience.}} \log P(R_i \mid \tilde{R}_{<i
3. Direct Preference Optimization on Failure Modes: Suppressing Inherent Low-Quality Predictions
Because supervised fine-tuning relies solely on maximum likelihood imitation of positive demonstrations, it fails to explicitly penalize recurrent structural confusions, allowing the model to generate catastrophic misalignments with non-trivial probability. To actively suppress these failure modes, the pipeline introduces Direct Preference Optimization (DPO). A preference dataset \(\mathcal{D}_{DPO}\) of 14,169 triplets \((I, y_w, y_l)\) is curated by running inference with the SFT checkpoint over the training split; predictions scoring below a Hard Match F1 threshold of 0.8 serve as rejected responses \(y_l\), while ground-truth sequences act as winning responses \(y_w\). Using the frozen SFT model as the reference policy \(\pi_{ref}\), the target policy \(\pi_\theta\) optimizes the implicit reward margin: $$ \mathcal{L}{DPO} = -\mathbb{E}{DPO}} \left[ \log \sigma \left( \beta \log \frac{\pi\theta(y_w \mid I)}{\pi_{ref}(y_w \mid I)} - \beta \log \frac{\pi_\theta(y_l \mid I)}{\pi_{ref}(y_l \mid I)} \right) \right] $$ This preference optimization suppresses high-frequency error patterns in ambiguous layout boundaries, preventing local indecision from destabilizing full sequence decoding.
Loss & Training¶
The progressive training schedule operates across three stages using a cosine learning rate scheduler with a warmup ratio of 0.03: 1. Pretraining Stage: Trained for 1 epoch on 60,000 synthetic diagrams; the vision encoder and cross-modal projector are frozen while updating LLM parameters at learning rate \(1.0 \times 10^{-6}\) with global batch size 16. 2. Multi-Task SFT Stage: Jointly optimizes all objectives with \(\mathcal{L}_{SFT} = \mathcal{L}_{VRP} + \mathcal{L}_{RGRP} + \mathcal{L}_{PPRP}\); full-parameter fine-tuning is conducted for 1 epoch at learning rate \(1.0 \times 10^{-5}\) with global batch size 4. 3. DPO Stage: Optimized over 14,169 preference pairs for 1 epoch; only the LLM parameters are updated at learning rate \(3.0 \times 10^{-7}\) with global batch size 64 and preference margin hyperparameter \(\beta\).
Key Experimental Results¶
Main Results¶
Evaluation is conducted on two benchmarks: the established RxnScribe-test benchmark (138 images, 392 reactions) and the newly curated, more demanding RobustRDP-test benchmark (500 real literature images, 2,634 reactions). Metrics comprise Hard Match (strictly requiring exact matching across all reactants, conditions, and products) and Soft Match (evaluating molecular entities without role differentiation between reactants and reagents), evaluated at IoU threshold 0.5.
| Dataset | Model | Hard Match Precision (%) | Hard Match Recall (%) | Hard Match F1 (%) | Soft Match Precision (%) | Soft Match Recall (%) | Soft Match F1 (%) |
|---|---|---|---|---|---|---|---|
| RxnScribe-test | RxnDE | 4.1 | 1.3 | 1.9 | 19.4 | 5.9 | 9.0 |
| OChemR | 4.4 | 2.8 | 3.4 | 12.4 | 7.9 | 9.6 | |
| RxnScribe | 72.3 | 66.2 | 69.1 | 83.8 | 76.5 | 80.0 | |
| RxnIM | 74.7 | 69.7 | 72.1 | 86.9 | 82.8 | 84.8 | |
| RxnCaption-VL | 71.6 | 72.7 | 72.2 | 85.3 | 87.1 | 86.2 | |
| RobustRDP (Ours) | 81.0 | 82.4 | 81.7 | 90.2 | 91.3 | 90.8 | |
| RobustRDP-test | RxnScribe | 69.7 | 61.7 | 65.5 | 82.7 | 72.3 | 77.2 |
| RxnIM | 57.5 | 50.2 | 53.6 | 76.2 | 65.8 | 70.6 | |
| RobustRDP (Ours) | 82.7 | 82.4 | 82.5 | 91.1 | 90.8 | 91.0 |
Ablation Study¶
The impact of each design component is systematically assessed on the RxnScribe-test set via incremental subtraction:
| Config | DPO | Pretrain | RGRP | PPRP | RobustRDP-train | Hard Match F1 (%) | Soft Match F1 (%) | Note |
|---|---|---|---|---|---|---|---|---|
| Full model | ✓ | ✓ | ✓ | ✓ | ✓ | 81.7 | 90.8 | Full RobustRDP achieves peak performance |
| w/o DPO | – | ✓ | ✓ | ✓ | ✓ | 80.9 | 90.0 | Missing negative suppression drops Hard F1 by 0.8% |
| w/o Pretrain & DPO | – | – | ✓ | ✓ | ✓ | 78.8 | 88.5 | Absence of synthetic pre-adaptation drops Hard F1 by 2.1% |
| w/o RGRP | – | – | – | ✓ | ✓ | 77.6 | 87.8 | Shortcut bias degrades local parsing, Hard F1 drops 1.2% |
| w/o PPRP | – | – | ✓ | – | ✓ | 75.7 | 86.8 | Lack of error recovery drops Hard F1 by 3.1% |
| w/o RGRP & PPRP | – | – | – | – | ✓ | 74.6 | 85.9 | Vanilla VRP suffers severe error propagation (-4.2%) |
| RxnScribe-train only | – | – | – | – | – | 70.4 | 80.8 | Relying solely on legacy data collapses Hard F1 to 70.4% |
Key Findings¶
- Prefix perturbation (PPRP) serves as the primary safeguard against error cascading: Removing PPRP inflicts a substantial 3.1% decline in Hard Match F1 (78.8% \(\to\) 75.7%), demonstrating that autoregressive generation is exquisitely sensitive to prefix errors and that noisy prefix exposure is vital for test-time resilience.
- Synthetic pretraining grounds fundamental structural decoding: Eliminating synthetic pretraining decreases Hard Match F1 by 2.1%. Although synthetic data exhibits distribution shift, its diverse topological coverage primes the MLLM to coordinate spatial bounding boxes with special role tokens.
- Robustness advantage widens on challenging real-world benchmarks: On the newly curated RobustRDP-test benchmark, previous SOTA model RxnIM suffers an 18.5% collapse in Hard Match F1 (72.1% \(\to\) 53.6%), whereas RobustRDP retains a high score of 82.5%, proving its capability in parsing complex, highly branched reaction pathways.
Highlights & Insights¶
- Inverting autoregressive brittleness into supervised resilience: Instead of abandoning generative MLLMs for complex two-stage architectures, the authors perturb prefix sequences during training, effectively teaching the model self-correcting behavior during autoregressive generation.
- Spatial prompt regularization to break linguistic shortcuts: Imposing region-guided bounding boxes effectively forces the MLLM to disentangle visual evidence from prior contextual text, ensuring high local parsing accuracy for individual reaction steps.
- Scalable synthetic-to-real paradigm for scientific vision tasks: Coupling a rule-based chemical diagram generator with semi-automated YOLO-assisted human verification offers an actionable blueprint for high-precision scientific document extraction where manual labels are prohibitively expensive.
Limitations & Future Work¶
- Absence of deep chemical reaction mechanism priors: The parser relies entirely on visual connectivity and text-token heuristics without verifying chemical plausibility (e.g., atom balance, standard functional group transformations). In heavily occluded or ambiguous diagrams, the model may predict chemically invalid pathways.
- Two-stage decoupling from molecular structure recognition: RobustRDP yields spatial boxes and roles rather than full SMILES strings or parsed chemical text, still requiring downstream OCSR and OCR models. Unifying diagram parsing and molecular structure decoding into a single end-to-end framework represents a promising frontier.
Related Work & Insights¶
- vs RxnDE / OChemR: Cascaded heuristic approaches rely on brittle rule-based arrow associations that collapse under complex layouts, scoring below 5% Hard F1; RobustRDP unifies localization and role assignment in a single sequence decoding pass, setting a much higher benchmark.
- vs RxnScribe / RxnIM: RxnScribe is bounded by a lightweight Pix2Seq capacity, while RxnIM suffers from shortcut learning and cascading error sensitivity; RobustRDP expands real-world training scale and resolves error accumulation through RGRP and PPRP auxiliary supervision, lifting Hard Match F1 from 72.1% to 81.7%.
- vs RxnCaption-VL: RxnCaption-VL relies on an external detector to paint visual prompt boxes directly onto images, which introduces upstream detection error propagation; RobustRDP parses raw diagrams end-to-end without visual tampering, surpassing RxnCaption-VL by 9.0% Hard Match F1.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Well-motivated auxiliary multi-tasking and preference learning directly addressing shortcut learning and oracle prefix vulnerabilities in autoregressive parsing]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Introduces an extensive 500-image real literature benchmark, with exhaustive component-wise ablations across both strict and soft evaluation metrics]
- Writing Quality: ⭐⭐⭐⭐⭐ [Exceptionally clear problem framing, coherent methodology presentation, and rigorous empirical analysis]
- Value: ⭐⭐⭐⭐⭐ [Establishes a solid foundational parser and open-source benchmark for automated scientific literature and patent data extraction]