HighlightBench: Benchmarking and Diagnosing Markup-Driven Table Reasoning in Scientific Documents¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Project: HighlightBench
Area: Multimodal VLM / LLM Reasoning
Keywords: Markup-driven table reasoning, Multimodal large language models, Diagnostic benchmark, Structured reasoning, Counterfactual analysis
TL;DR¶
Addressing the evaluation blind spot where multimodal large language models cannot distinguish between perceiving visual cues and executing logical constraints, HighlightBench introduces a diagnostic benchmark of 446 tables and 3,283 questions across five decoupled task families, counterfactual perturbations, and an explicit DSL reference pipeline to expose how visual salience systematically misleads reasoning.
Background & Motivation¶
In scientific papers, technical reports, and data analysis workflows, tables frequently feature visual markups such as colored highlights, underlines, bold typography, arrows, or shaded cells to emphasize critical information. These visual markers are not merely decorative styling; they serve as explicit, functional directives that govern which rows, columns, or cells should be extracted, compared, filtered, or verified. For multimodal large language models (MLLMs), interpreting markup-conditioned tables introduces challenges distinct from standard TableQA. A model must not only parse text values and 2D grid coordinates, but also bind queried visual cues to the intended structural units, translate them into executable constraints over the table topology, and preserve these constraints throughout multi-step reasoning.
However, existing document and table evaluation benchmarks rely predominantly on end-to-end answer accuracy, creating a severe diagnostic blind spot. Recent investigations show that state-of-the-art models often output surface-correct answers without faithfully grounding the queried markup, relying instead on dataset regularities, linguistic priors, or opportunistic shortcuts. Conversely, when an answer is wrong, scalar accuracy fails to attribute whether the model failed to detect the markup, mis-bound the visual cue to the wrong cell coordinates, or blundered during downstream symbolic calculation. This conflation leaves the community unable to assess whether MLLMs genuinely reason under markup-conditioned rules.
This paper tackles the ambiguity by decomposing end-to-end execution into a transparent perception-conditioning-execution chain and introducing controlled counterfactual interventions. Core idea: formalize markup-conditioned table reasoning as a three-stage probabilistic chain over structural evidence binding, question-induced constraint generation, and topological execution, disentangling genuine constraint following from visual salience shortcuts via five decoupled task families and counterfactual marker relocations.
Method¶
Overall Architecture¶
HighlightBench formalizes each evaluation instance as a tuple \((I, M, Q, A)\), where \(I\) denotes the table image, \(M\) the set of embedded visual markups, \(Q\) the natural language question, and \(A\) the expected structured answer. To localize failure modes along the operational pipeline, the benchmark introduces two intermediate latent variables: \(z\), representing the structural evidence anchored by markups (such as marked cells, headers, rows, columns, or local neighborhoods), and \(c\), representing the reasoning condition induced by that evidence under the question context (such as filtered candidate subsets, directional relations, or comparison scopes).
The methodology rests on three core pillars: diagnostic task factorization, hybrid benchmark construction with counterfactual protocols, and an explicit reference pipeline. The reference pipeline parses an input table image into a graph-structured intermediate representation, resolves execution paths via two-stage routing, and executes deterministic domain-specific language (DSL) plans to produce structured answers alongside fine-grained diagnostic traces.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Table Image & Question<br/>(I, M, Q)"] --> B["Docgraph Generation<br/>Text content/topology/markup binding"]
B --> C["Two-Stage Routing<br/>Coarse operator routing & uncertainty resolution"]
C --> D["DSL Execution<br/>Filtering/extrema/aggregation/consistency checks"]
D --> E["Structured Output & Diagnostic Trace<br/>(Answer & Intermediate Traces)"]
Key Designs¶
1. Diagnostic Factorization and Five Task Families: Decoupling Perception, Conditioning, and Execution Standard evaluations directly compute \(P(A \mid I, M, Q)\), masking internal breakdowns. HighlightBench expands this into a full probabilistic factorization:
This formulation yields five complementary task families covering 21 distinct subtasks: - T1. Markup Grounding: Evaluates \(P(z \mid I, M)\) by probing whether models can accurately associate visual highlights, underlines, boldface, or arrows with corresponding cells, headers, or columns, including marker existence verification, marked cell counting, and style-conditioned extraction. - T2. Constrained Retrieval: Evaluates \(P(c \mid z, Q)\) by testing whether models can retrieve entries strictly within an explicit scope (such as color-grouped highlights, target delta columns, or coordinate intersections) without leaking values from visually salient yet out-of-scope regions. - T3. Local Relations: Evaluates local topological stability by requiring models to locate an anchor cell and extract its directional neighbors (top/bottom/left/right), testing whether structural integrity is preserved when transitioning from evidence \(z\) to execution. - T4. Aggregation & Comparison: Evaluates \(P(A \mid I, c, Q)\) by executing numerical comparisons, extrema selection (argmax/argmin), sorting, or summation over candidate subsets strictly bounded by markups. - T5. Consistency & Missingness: Evaluates robustness against table irregularities (e.g., dashes, N/A, partially missing entries, or incomplete grids), checking whether models respect missing evidence rather than hallucinating completions.
2. Hybrid Data Construction and Counterfactual Probing: Isolating Reading from Salience Bias The benchmark comprises 446 table images and 3,283 questions. The real-world split includes 266 scientific tables and 2,268 questions drawn from CV and NLP conference publications, retaining natural irregularities such as merged cells, multi-tier headers, and hand-drawn or varied markup styles (annotator agreement reached Cohen's \(\kappa = 0.96\)). The synthetic split includes 180 images and 1,015 questions generated via programmatic templates to provide controllable difficulty and distractor configurations.
To conclusively distinguish basic table reading from markup-induced reasoning failures, the authors establish a rigorous counterfactual protocol: holding table textual data and question text strictly identical while systematically shifting the visual markup across 7 variants: - No-markup baseline (\(n\)); - Congruent variants (\(H_c, U_c, B_c\)): placing highlight, underline, or bold styling on the ground-truth target cell; - Incongruent variants (\(H_i, U_i, B_i\)): placing the identical styling on a competitive distractor cell. Under this protocol, failed model predictions are mapped to three operational attribution categories: Marker Hijacked (incorrect prediction that lands precisely on the marked distractor, indicating salience overrode logic), Logical Failure (incorrect prediction without hitting the marker, reflecting pure arithmetic/reasoning errors), and Coordinate Error (correct numeric value but wrong associated entity/row label).
3. Explicit Reference Pipeline: Docgraph, Two-Stage Routing, and Deterministic DSL Execution To provide an interpretable diagnostic baseline, the authors construct a modular reference pipeline that replaces monolithic autoregression with explicit symbolic execution: - Docgraph Generation: Converts the raw image into an explicit topological graph of cell nodes and adjacency edges, binding detected markups as structured node attributes rather than latent image features. - Two-Stage Routing: Decomposes decision planning into a coarse router (determining query topology and primary operator types) and a fine uncertainty resolver (resolving ambiguous or weak visual evidence before dispatching execution plans). - DSL Execution: Compiles the routed strategy into deterministic domain-specific language statements executed over the docgraph (atomic filtering, extrema calculation, aggregation, and neighborhood lookup), eliminating syntax degeneration and calculation drift.
Loss & Training¶
As an evaluation benchmark, HighlightBench employs a standardized zero-shot and few-shot evaluation protocol across all evaluated open-source and proprietary models via VLMEvalKit. Models receive structured prompt templates requiring outputs formatted as strict JSON dictionaries. A dedicated post-processing evaluator extracts fields and computes Exact-Match (EM) accuracy against ground-truth annotations, preventing superficial textual differences from confounding evaluation accuracy.
Key Experimental Results¶
Main Results¶
Twelve leading multimodal models (nine open-source, three proprietary) and the reference pipeline (Ours) were evaluated across all task families. Exact-match accuracy (%) across both splits is presented below:
| Model Category | Model Name | Synthetic (T1) | Synthetic (T2) | Synthetic (T3) | Synthetic (T4) | Synthetic (T5) | Synthetic (All) | Real-world (T1) | Real-world (T2) | Real-world (T3) | Real-world (T4) | Real-world (T5) | Real-world (All) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Open-source | Gemma 3 4B | 4.3 | 10.2 | 2.4 | 20.3 | 23.3 | 12.4 | 22.2 | 9.9 | 2.5 | 10.1 | 16.5 | 13.7 |
| Open-source | InternVL3.5-8B | 29.0 | 81.2 | 17.6 | 24.7 | 71.0 | 40.5 | 54.2 | 53.1 | 23.0 | 36.6 | 67.7 | 46.3 |
| Open-source | Kimi-VL-A3B | 7.4 | 21.0 | 12.1 | 31.2 | 49.7 | 23.8 | 33.1 | 37.5 | 17.4 | 25.9 | 17.0 | 28.6 |
| Open-source | LLaVA1.5-7B | 8.1 | 0.7 | 0.0 | 0.0 | 1.0 | 2.7 | 3.4 | 1.5 | 1.2 | 0.2 | 0.0 | 1.4 |
| Open-source | MiniCPM-V 2.6 | 17.6 | 59.0 | 7.9 | 13.1 | 14.5 | 20.4 | 17.2 | 29.0 | 9.9 | 17.5 | 3.4 | 17.5 |
| Open-source | MiniCPM-V 4.5 | 26.2 | 74.3 | 13.3 | 28.8 | 59.2 | 36.9 | 56.2 | 50.0 | 15.5 | 35.9 | 35.7 | 42.8 |
| Open-source | Qwen2.5-VL-3B | 14.6 | 61.1 | 6.7 | 13.1 | 47.1 | 24.9 | 39.4 | 35.6 | 1.2 | 31.3 | 58.9 | 34.7 |
| Open-source | Qwen3-VL-8B | 37.0 | 74.3 | 18.8 | 31.5 | 74.4 | 44.2 | 56.3 | 57.1 | 24.2 | 36.3 | 78.0 | 48.4 |
| Open-source | Qwen3.5-9B | 27.7 | 28.5 | 26.7 | 45.5 | 60.5 | 37.7 | 55.3 | 50.4 | 23.6 | 42.8 | 70.1 | 48.8 |
| Closed-source | Claude Sonnet 4 | 32.7 | 81.9 | 40.0 | 37.5 | 67.1 | 47.6 | 56.7 | 59.6 | 42.2 | 60.7 | 68.2 | 58.7 |
| Closed-source | Gemini 2.5 Flash | 51.9 | 90.3 | 89.1 | 67.8 | 93.9 | 73.9 | 81.3 | 75.3 | 62.7 | 72.2 | 78.5 | 75.3 |
| Closed-source | GPT-4o mini | 13.0 | 15.1 | 13.3 | 30.1 | 32.3 | 21.1 | 21.3 | 16.3 | 8.5 | 9.4 | 37.7 | 16.4 |
| Reference Baseline | Ours (Reference Pipeline) | 53.4 | 88.9 | 73.3 | 81.5 | 78.9 | 72.3 | 72.9 | 74.2 | 70.8 | 82.2 | 79.8 | 77.1 |
Ablation Study¶
On the controlled counterfactual subset evaluated with Qwen3-VL-8B, table content was held constant while markup locations varied across seven configurations. The table below details accuracy and error breakdowns between control reading tasks and reasoning tasks:
| Evaluation Probe | Baseline (\(n\)) | Congruent High. (\(H_c\)) | Congruent Und. (\(U_c\)) | Congruent Bold (\(B_c\)) | Incongruent High. (\(H_i\)) | Incongruent Und. (\(U_i\)) | Incongruent Bold (\(B_i\)) | Empirical Behavioral Profile |
|---|---|---|---|---|---|---|---|---|
| Cell Retrieval Control (%) | 95.2 | 96.1 | 95.8 | 96.5 | 94.8 | 94.5 | 95.0 | Consistently stable (94.5%–96.5%), confirming OCR/reading is intact |
| Extremum Reasoning (%) | Reference | Large Gain | Large Gain | Large Gain | Severe Drop | Severe Drop | Severe Drop | Congruent markups boost accuracy; incongruent markups derail reasoning |
| Attribution: Marker Hijacked | - | - | - | - | Dominant (>60%) | Dominant (>60%) | Dominant (>60%) | Model predicts the marked distractor instead of calculating extremum |
| Attribution: Logical Failure | Arithmetic error | Minor share | Minor share | Minor share | Minor share (<30%) | Minor share (<30%) | Minor share (<30%) | Intrinsic calculation error uninfluenced by markup location |
| Attribution: Coordinate Error | Label mismatch | Minor share | Minor share | Minor share | Minor share (<10%) | Minor share (<10%) | Minor share (<10%) | Numerical value correct but bound to incorrect row header |
Key Findings¶
- Visual salience directly hijacks symbolic reasoning: When visual cues are placed on suboptimal distractors, Extremum Reasoning accuracy plummets, with more than 60% of all failures attributed to "Marker Hijacked". Because Cell Retrieval accuracy remains virtually flat across all variants (~95%), this degradation is conclusively caused by reasoning shortcuts rather than visual recognition failures.
- Asymmetric perceptual difficulty across markup types: Models exhibit pronounced negative bias on bold and underline recognition (easily identifying unmarked cells), whereas highlight backgrounds prove significantly harder and less stable to perceive, frequently confusing models when light shading is used.
- Task family bottlenecks: While models achieve relatively solid numbers on T2 (Constrained Retrieval), they falter heavily on T1 (Markup Grounding) and T3 (Local Relations). Furthermore, T4 (Aggregation & Comparison) magnifies upstream grounding errors, leading to compound failures during arithmetic reduction.
- Superiority of explicit symbolic pipelines in structured tasks: By separating graph grounding from deterministic execution, the reference pipeline achieves 70.8% on real-world T3 (vs. Claude Sonnet 4's 42.2%) and 82.2% on real-world T4 (vs. Gemini 2.5 Flash's 72.2%), proving that intermediate modular representation is far more resilient against visual hijacking.
Highlights & Insights¶
- Pure causal evaluation via counterfactual interventions: By freezing table values and textual questions while solely perturbing visual markup coordinates, HighlightBench isolates the causal impact of visual styling from semantic priors.
- Elevating visual marks to first-class logical constraints: Rather than treating highlights as passive visual features, the paper models them as formal set-theoretic filtering criteria, bridging vision and symbolic execution.
- Traceable error taxonomy for post-mortem diagnostics: The tri-fold attribution framework (Marker Hijacked vs. Logical Failure vs. Coordinate Error) provides an actionable diagnostic protocol for debugging document foundation models.
Limitations & Future Work¶
- Scope of visual styles: The benchmark currently focuses on conventional academic markup forms (highlights, underlines, bold text, clean arrows) and does not cover chaotic real-world industrial artifacts such as handwritten margin annotations, coffee stains, or freehand lasso circles.
- Graph construction sensitivity: The reference pipeline relies heavily on accurate initial table parsing; unbordered, nested, or merged multi-tier headers can degrade initial docgraph topology.
- Future directions: Incorporating counterfactual markup consistency objectives into multimodal instruction tuning and reinforcement learning (e.g., GRPO with constraint-penalized rewards) to suppress visual salience shortcuts.
Related Work & Insights¶
- vs TAPAS / TableBench: Traditional TableQA benchmarks evaluate comprehension over clean text grids without visual styling, failing to assess how visual directives interact with table topology.
- vs VIP-LLaVA / ControlMLLM: Existing visual prompting frameworks require human-drawn external prompts injected at inference time, whereas HighlightBench tests inherent comprehension of document-native markups whose semantic functions must be inferred autonomously.
- vs SO-Bench / ExtractBench: While structured output benchmarks emphasize format compliance, HighlightBench couples schema constraints with counterfactual visual interventions to diagnose the complete perception-to-execution chain.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneering formalization of markup-driven table constraints with a rigorous counterfactual protocol]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive benchmarking over 12 SOTA MLLMs across synthetic and real-world scientific documents]
- Writing Quality: ⭐⭐⭐⭐⭐ [Meticulous formulation, clear probabilistic factorization, and thorough error analysis]
- Value: ⭐⭐⭐⭐⭐ [Crucial diagnostic benchmark for pushing MLLMs from shallow document readers to reliable scientific assistants]