CHARTSTYLE-100K: A Large-Scale Dataset for Structured Visualization Style Transfer¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/yuweiyang-anu/ChartStyle
Area: Image Generation
Keywords: structured visualizations, style transfer, image editing, reverse data generation, diffusion models
TL;DR¶
To overcome data-encoding geometric distortion and reference content leakage in structured visualization stylization, this paper presents ChartForge, a reverse-generation pipeline synthesizing 100K high-fidelity triplets, alongside ChartStyle-Bench and ReChart, a progressively fine-tuned model achieving state-of-the-art element-level style transfer.
Background & Motivation¶
Structured visualizations—spanning statistical charts, infographics, flowcharts, diagrams, and tables—form the bedrock of quantitative communication, translating abstract numeric and relational records into intuitive geometric patterns. Beyond raw data readability, visual style critically transforms an ordinary graphic from a functional plot into a trustworthy, professional artifact by shaping typography, palette harmony, mark aesthetics, and decorative hierarchy. However, manually crafting high-quality stylized visual designs remains labor-intensive and demands deep graphic expertise. While modern text-to-image diffusion backbones and instruction-based editing models show remarkable proficiency across natural scenes, directly deploying them to restyle structured charts leads to catastrophic failures.
These failures stem from three fundamental challenges unique to structured visual domains. First is compositional style: unlike a uniform texture or painterly filter in natural images, chart styling involves a multi-attribute ensemble comprising palettes, font types, glyph shapes, tick marks, and layout density that requires selective, component-level control. Second is zero-tolerance data-encoding geometry: visualizations encode quantitative truth directly into spatial metrics (bar lengths, sector angles, scatter coordinates, edge connections), where minor perceptual distortions corrupt the underlying data. Third is acute style-content entanglement: because both the style exemplar and the content image share an identical visual syntax of axes, bars, and legends, general-purpose image editors frequently suffer from severe content leakage, directly copying elements from the reference into the target. Crucially, addressing this data scarcity via naive forward synthesis (prompting models to restyle an existing content chart toward a reference) inevitably inherits these flaws, yielding heavily contaminated training supervision.
To resolve this dilemma, this work turns the conventional synthesis paradigm upside down: rather than forcing style onto fixed content, it builds paired supervision in reverse—first synthesizing a coherent stylized target from an exemplar, and then deriving a structure-aligned counterpart rendered under a baseline aesthetic. Core idea: develop the reverse-generation pipeline ChartForge to construct structurally disentangled triplets by synthesizing stylized targets before deriving matching content counterparts, curate the 100K-scale ChartStyle-100K dataset, and train ReChart via progressive two-stage fine-tuning to deliver fine-grained, component-aware chart stylization without content leakage.
Method¶
Overall Architecture¶
The structured visualization style transfer task aims to produce a target visual \(I_t = \Phi(I_s, I_c)\) given a style exemplar \(I_s\) and a content visual \(I_c\) endowed with geometry \(G_c\) and textual annotations \(T_c\). The joint optimization objective balances compositional style transfer \(\mathcal{L}_{\text{comp}}(I_t, I_s)\), geometric and text fidelity \(\mathcal{L}_{\text{struct}}(I_t, I_c)\) enforcing \(\|G(I_t) - G_c\| \to 0\) and \(\text{OCR}(I_t) = T_c\), and a leakage penalty \(\mathcal{L}_{\text{entgl}}(I_t, I_s, I_c)\). To provide uncorrupted supervision across these competing criteria, ChartForge implements a four-stage synthetic pipeline that feeds into the two-stage curriculum learning scheme of ReChart.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Style Exemplar Pool<br/>58K real infographic and Canva designs"] --> B["Reference-Driven Target Generation<br/>Synthesize It using independent content parameters"]
B --> C["Restyle-Based Content Generation<br/>Derive Ic under 90 visual families preserving geometry"]
C --> D["Style-Space Resampling<br/>Recycle restyled variants as new exemplars"]
D --> E["Multi-Dimensional Quality Assessment and Filtering<br/>LLM criteria + PaddleOCR token F1 filtering"]
E --> F["Progressive Fine-Tuning Framework<br/>General domain alignment + ChartStyle-100K specialization"]
F --> G["ReChart Image Editing Model<br/>Element-aware transfer with strict data preservation"]
Key Designs¶
1. Reference-driven target generation: neutralizing content leakage at the source
Standard synthetic workflows rely on forward conditioning (\(I_c \xrightarrow{I_s} I_t\)), where generative models invariably entangle semantic elements of \(I_s\) with \(I_c\). To eliminate this failure mode, ChartForge adopts a reverse construction strategy. In Stage I, generation begins solely from a style exemplar \(I_s\), paired with independently sampled content attributes drawn from structured pools spanning 36 chart categories, 26 academic/business domains, multiple panel arrangements, and variable information densities. The generator produces a fresh stylized target \(I_t\) that thoroughly reflects the visual aesthetic, color palette, and mark rendering of \(I_s\). Because \(I_t\) contains newly fabricated data rather than pre-existing content, the model focuses exclusively on style adherence without any competing structural constraints, effectively avoiding content contamination.
2. Restyle-based content generation: structure-preserving transformation into baseline content
With the high-aesthetic target \(I_t\) established, the pipeline must synthesize a structurally identical counterpart \(I_c\) rendered in a neutral or contrasting style. Directly stripping stylistic attributes from \(I_t\) is an ill-posed inverse problem. Instead, Stage II prompts the generative model to restyle \(I_t\) under a newly sampled visual aesthetic chosen from 90 predefined style families (e.g., minimalist, retro, data journalism, 3D-shaded). The prompt strictly enforces the preservation of element-level geometry (bar lengths, pie sector angles, point coordinates) and textual labels, altering only decorative marks, backgrounds, and color palettes. To provide an anchor toward typical real-world use cases, at least one generated variant is constrained to adopt a plain baseline aesthetic such as standard Matplotlib styling or a monochrome sketch. This produces clean \((I_s, I_c, I_t)\) triplets where structure is invariant by construction.
3. Style-space resampling: mitigating real-world exemplar distribution skew
Web-crawled infographic repositories suffer from severe stylistic homogeneity, where over 80% of samples cluster around standard corporate flat-vector graphics, leaving artistic, illustrative, and niche aesthetics critically under-represented. To prevent the downstream editor from collapsing into a narrow palette distribution, Stage III introduces self-augmentation via style-space bootstrapping: high-quality restyled variants produced during Stage II are harvested and reintroduced as novel style references \(\tilde{I}_s\). The full generation and restyling cycles are then re-executed around these bootstrapped exemplars. This expansion broadens coverage across long-tail visual aesthetics without necessitating external manual data collection.
4. Multi-dimensional quality assessment and filtering: hybrid LLM and OCR verification
Raw synthetic triplets can exhibit minor textual distortions, layout drift, or subtle reference leakage. Stage IV applies a stringent multi-dimensional filtration protocol leveraging GPT-4o as a visual judge alongside PaddleOCR: - Visual quality \(R_{cvq}, R_{tvq} \in [1, 5]\): independently evaluates \(I_c\) and \(I_t\) for typography clarity, visual contrast, layout discipline, and generative artifacts; - Content consistency \(R_{cc} \in [1, 5]\): jointly scrutinizes \(I_c\) and \(I_t\) to ensure structural topology, axis titles, data series, and textual annotations remain identical; - Stylistic consistency \(R_{sc} \in [1, 5]\): compares \(I_s\) and \(I_t\) exclusively across visual style dimensions (color palette, typeface features, mark textures); - Character-level OCR overlap \(R_{ocr} \in [0, 1]\): extracts text tokens from \(I_c\) and \(I_t\) to compute token-level \(F_1\) scores. Filtering thresholds are set to \(R_{cvq} \ge 4\), \(R_{tvq} \ge 4\), \(R_{cc} \ge 3\), \(R_{sc} \ge 3\), and \(R_{ocr} \ge 0.8\). This rigorous audit refines an initial pool of 308,409 candidate triplets down to 100,744 pristine instances, forming the ChartStyle-100K dataset.
5. Progressive fine-tuning framework: curriculum transfer from general editing to chart specialization
Directly fine-tuning base diffusion editors (such as Qwen-Image-Edit) entirely on chart data yields suboptimal results because the pre-trained weights lack cross-image feature alignment capabilities. ReChart employs a two-stage training curriculum: the first stage fine-tunes the double-stream MMDiT backbone on broad, open-domain style transfer pairs to establish robust cross-attention routing between disparate visual inputs; the second stage fine-tunes specifically on ChartStyle-100K. Cross-attention analysis confirms that this progressive regimen successfully decouples internal representations: style attention concentrates on appearance-sensitive textures and tones, while content attention focuses tightly on structure-defining coordinates and glyph boundaries.
Loss & Training¶
ReChart is trained using a flow matching objective on rectified flow dynamics. Given image latent \(x_1\) and Gaussian prior noise \(x_0\), the model learns a vector field \(v_\theta(x_t, t, c)\) along the probability trajectory \(x_t = t x_1 + (1 - t) x_0\) conditioned on \(c = \{I_s, I_c\}\):
Training incorporates aspect-ratio-preserving high-resolution image encoding, cosine learning rate scheduling, and gradient clipping to ensure optimization stability.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on ChartStyle-Bench, a curated testbed of 300 diverse, non-overlapping content-style pairs verified by human annotators. Performance is assessed using GPT-4o evaluation (Content Consistency, Style Similarity, Content Leakage indicator, and Overall harmonic score) alongside perceptual and OCR metrics (CLIP Semantic Consistency, CLIP Stylistic Fidelity, and PaddleOCR word-level F1).
| Model | Content↑ | Style↑ | Leakage↓ | Overall↑ | Semantic↑ | Fidelity↑ | OCRScore↑ |
|---|---|---|---|---|---|---|---|
| StyleStudio [24] | 1.00 | 1.67 | 0.06 | 1.17 | 0.58 | 0.47 | 0.01 |
| CSGO [57] | 1.00 | 1.97 | 0.02 | 1.25 | 0.55 | 0.46 | 0.01 |
| OmniGen2 [54] | 1.11 | 2.66 | 0.57 | 1.14 | 0.54 | 0.53 | 0.04 |
| FLUX.2-dev [23] | 1.75 | 3.80 | 0.80 | 1.29 | 0.59 | 0.73 | 0.26 |
| Edit-R1 [27] | 1.85 | 3.25 | 0.63 | 1.35 | 0.65 | 0.66 | 0.28 |
| Seedream 4.0 [44] | 2.39 | 3.70 | 0.51 | 2.00 | 0.74 | 0.56 | 0.50 |
| GPT-Image-1 [38] | 2.31 | 2.96 | 0.05 | 2.30 | 0.77 | 0.46 | 0.54 |
| GPT-Image-1.5 [39] | 3.30 | 3.04 | 0.06 | 2.82 | 0.80 | 0.48 | 0.73 |
| Nano-Banana [3] | 3.57 | 2.66 | 0.20 | 2.44 | 0.87 | 0.50 | 0.77 |
| Nano-Banana-2 [11] | 3.48 | 3.75 | 0.48 | 2.59 | 0.76 | 0.54 | 0.71 |
| Nano-Banana-Pro [10] | 3.79 | 3.58 | 0.32 | 2.84 | 0.77 | 0.54 | 0.72 |
| Qwen-Image-Edit [53] | 2.75 | 2.73 | 0.51 | 1.58 | 0.76 | 0.62 | 0.54 |
| ReChart (Ours) | 3.96 | 2.90 | 0.05 | 3.03 | 0.89 | 0.49 | 0.76 |
Ablation Study¶
1. Progressive two-stage training ablation on ChartStyle-Bench The impact of each training phase is ablated by measuring content retention, style replication, and leakage suppression:
| Config | Content↑ | Style↑ | Leakage↓ | Overall↑ | Semantic↑ | Fidelity↑ | OCRScore↑ |
|---|---|---|---|---|---|---|---|
| Qwen-Image-Edit base | 2.75 | 2.73 | 0.51 | 1.58 | 0.76 | 0.62 | 0.54 |
| + Stage 1 only (General style) | 3.13 | 2.92 | 0.13 | 2.45 | 0.80 | 0.52 | 0.59 |
| + Stage 2 only (Direct ChartStyle) | 2.45 | 3.15 | 0.03 | 2.53 | 0.84 | 0.48 | 0.54 |
| + Stage 1 + Stage 2 (Full ReChart) | 3.96 | 2.90 | 0.05 | 3.03 | 0.89 | 0.49 | 0.76 |
2. Impact of style-space resampling (Stage III)
| Config | Content↑ | Style↑ | Leakage↓ | Overall↑ | Semantic↑ | Fidelity↑ | OCRScore↑ |
|---|---|---|---|---|---|---|---|
| ReChart (Full) | 3.96 | 2.90 | 0.05 | 3.03 | 0.89 | 0.49 | 0.76 |
| w/o style-space resampling | 3.57 | 2.75 | 0.09 | 2.71 | 0.88 | 0.48 | 0.73 |
3. Text fidelity across bounding box scales (PaddleOCR F1)
| Model | <12 px | 12–14 px | 14–18 px | 18–24 px | 24–32 px | ≥32 px |
|---|---|---|---|---|---|---|
| ReChart (Ours) | 0.56 | 0.56 | 0.66 | 0.72 | 0.75 | 0.78 |
| GPT-Image-1.5 [39] | 0.52 | 0.58 | 0.62 | 0.70 | 0.70 | 0.69 |
| Nano-Banana-Pro [10] | 0.50 | 0.60 | 0.68 | 0.74 | 0.71 | 0.72 |
Key Findings¶
- The illusion of inflated style scores: Baseline models such as FLUX.2-dev and Nano-Banana-2 report deceivingly high raw style scores (3.80 and 3.75), yet suffer from severe leakage rates of 0.80 and 0.48. These systems achieve visual similarity primarily by copying elements directly from the reference chart, crashing their effective overall performance.
- Necessity of curriculum staging: Directly training on chart triplets without general style pre-training reduces leakage (0.03) but causes Content Consistency to collapse from 2.75 to 2.45. Only the sequential pairing of general cross-attention grounding followed by domain specialization achieves superior structural fidelity (3.96) while keeping leakage at 0.05.
- Small-scale typography remains an open bottleneck: Scale-stratified OCR analysis demonstrates that while text above 14px is faithfully rendered across all leading models (F1 \(\ge 0.66\)), resolution degrades rapidly below 12px (dropping to 0.50–0.56), highlighting the limits of standard latent patch tokenization for dense, microscopic characters.
Highlights & Insights¶
- Inverted generation logic for zero-leakage supervision: Addressing structural leakage by reversing the synthesis pipeline—generating stylized targets before deriving paired content counterparts—provides an elegant, scalable data creation principle for structured visual domains.
- Mechanistic verification via attention tomography: Analyzing double-stream MMDiT cross-attention layers and element-removal interventions empirically proves that ReChart learns element-aware routing, channeling style influence into textures while binding content influence to layout coordinates.
- Pixel-space editing superiority over code generation: Direct image-to-code-to-render pipelines (e.g., via Gemini-3-Pro) fail to capture rich illustrative textures and fine visual nuances, confirming that end-to-end pixel diffusion offers a substantially higher expressive ceiling for stylized visualization.
Limitations & Future Work¶
- Legibility in highly compressed micro-annotations: Sub-12px annotations and crowded dense data labels can occasionally suffer from pixel blurring or stroke artifacts, pointing to the need for text-aware regional super-resolution or hybrid vector rasterization.
- Handling ultra-complex multi-panel dashboards: For intricate, multi-layered dashboards with non-standard relational topologies, the model occasionally misattributes secondary style accents across distinct sub-charts, motivating future explorations into hierarchical graph-guided conditioning.
Related Work & Insights¶
- vs CSGO [57] & StyleStudio [24]: Existing natural image stylization methods rely on global IP-Adapter embeddings, generating severe structural smearing and unusable illegible artifacts when applied to data visualizations; ReChart enforces strict element-level geometric integrity.
- vs Qwen-Image-Edit [53]: While the base foundation editor exhibits strong open-world capabilities, its lack of structural inductive bias yields over 50% content leakage; ReChart eliminates leakage through targeted progressive fine-tuning on clean reverse-synthesized triplets.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Reversing the data synthesis workflow to decouple style from content provides a brilliant conceptual paradigm.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorously tested across 300 benchmark pairs with multi-LLM scoring, perceptual metrics, human A/B studies, and scale-aware OCR ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Lucid formulation, thorough motivation, and intuitive qualitative and attention analysis figures.
- Value: ⭐⭐⭐⭐⭐ Opens a highly practical, commercially viable frontier in AI-driven graphic design and automated data storytelling.