i-Design: Step-by-Step Graphic Layout Design with Progressive Aesthetic Policy Optimization¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Multimodal VLM / VLM Reasoning
Keywords: Graphic Layout Generation, Progressive Decision Making, Aesthetic Policy Optimization, Multimodal Large Language Model, Reinforcement Learning
TL;DR¶
Addressing the visual imbalance and spatial misalignment inherent to conventional single-shot layout models, i-Design reformulates 2D graphic layout generation as an iterative group-wise trajectory decision process conditioned on partially rendered intermediate canvases, integrating Progressive Imitation Learning (PIL) with online Progressive Aesthetic Policy Optimization (PAPO) driven by ground-truth anchored consensus graph rewards to achieve state-of-the-art geometric alignment and human aesthetic preferences.
Background & Motivation¶
Automated 2D graphic layout generation forms the operational foundation of graphic design, spanning promotional posters, web user interfaces, presentation slides, and digital publications. A compelling graphic composition demands more than geometrically collision-free bounding boxes; it hinges upon nuanced perceptual criteria such as visual hierarchy, reading order, typographical balance, and deliberate negative space. Early approaches relied on rule-based heuristics and energy minimization frameworks, which suffered from rigid constraint modeling and poor stylistic generalization. With the advent of deep generative architectures and multimodal large language models (such as LayoutTransformer, LayoutDiffusion, PosterLLaVA, and LayoutNUWA), data-driven spatial synthesis has advanced considerably. Nevertheless, the vast majority of existing techniques cast layout planning as a single-shot prediction problem, forcing the generator to determine the coordinates of all graphic and textual elements concurrently in a single forward pass.
This single-shot paradigm fundamentally departs from the iterative nature of human design cognition. Professional graphic designers rarely finalize complex compositions in one stroke; rather, they follow a feedback-driven workflow: anchoring salient elements, appraising visual weight and negative space on the working canvas, progressively introducing secondary content, and dynamically fine-tuning alignments. When models predict all elements simultaneously without intermediate visual feedback, spatial collisions, unbalanced margins, and semantic clutter frequently emerge. Furthermore, standard training schemes enforce strict token- or coordinate-level supervised regression, rigidly penalizing valid aesthetic variations, while preference alignment methods applied to graphic layout have remained largely restricted to offline, single-turn ranking without online trajectory exploration.
These challenges pose a foundational research question: can graphic layout generation be reframed as an iterative, trajectory-level decision process steered by visual aesthetic feedback over progressive canvas states? The central angle of attack is to harness the visual perception and spatial grounding of Vision-Language Models (VLMs), enabling the generator to observe the rendered partial canvas at each step, predict the next element group, and optimize long-horizon visual harmony via online reinforcement learning. Core Idea: Formulate graphic layout generation as a sequential, partially rendered canvas-conditioned trajectory process, establishing robust geometric grounding via Progressive Imitation Learning (PIL) and steering compositional balance through online Progressive Aesthetic Policy Optimization (PAPO) powered by a ground-truth anchored consensus graph.
Method¶
Overall Architecture¶
i-Design models the layout generation task as a multi-step trajectory rollout. Given an input set of \(N\) design elements \(E = \{e_i = (p_i, s_i, \tau_i)\}_{i=1}^N\) (comprising center positions, dimensions, and element categories), the system deterministically partitions them into \(G = \lceil N/k \rceil\) ordered placement groups \(A = \{A_1, \dots, A_G\}\) following their intrinsic \(z\)-order hierarchy, where each group contains up to \(k\) elements. At each progressive step \(g\), the Skia graphics engine renders the partial canvas \(C_{g-1} = \mathcal{R}(E^{<g})\) containing all elements placed in preceding turns. A multimodal encoder-decoder architecture based on InternVL-3 8B fuses the rendered canvas, verbalized prompt tokens, and latent element embeddings, predicting the placement coordinates for remaining elements. During inference, only the immediate next group \(A_g\) is committed and rendered to update the canvas, continuing recursively until the full trajectory is executed. The training framework consists of three stages: pretraining, Progressive Imitation Learning (PIL), and Progressive Aesthetic Policy Optimization (PAPO).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Design Elements Set<br/>(Category/Content/Visual Embeddings)"] --> S1["Grouping & Canvas Verbalization<br/>(z-order Grouping/Skia Partial Rendering)"]
S1 --> S2["Progressive Imitation Learning (PIL)<br/>(Autoregressive Supervision on Partial Canvases)"]
S2 --> S3["Progressive Aesthetic Policy Optimization (PAPO)<br/>(Online Sampling of Layout Rollout Trajectories)"]
S3 --> S4["Ground-Truth Anchored Consensus Graph<br/>(DesignSense Pairwise Judging/PageRank Scoring)"]
S4 --> Out["Aesthetically Coherent Layout Generation<br/>(Geometric Precision & Visual Balance)"]
Key Designs¶
1. Grouping & Canvas Verbalization: Grounding layout decisions in progressive visual feedback Prior autoregressive layout models operated at the individual token level without holistic image-space visual grounding, whereas single-shot models entirely lacked opportunities for intermediate correction. i-Design implements a deterministic grouping strategy based on the element stacking order (\(z\)-order), chunking \(N\) elements into \(G\) groups with group size \(k \approx \lceil\sqrt{N}\rceil\). At step \(g\), elements already committed \(e_i \in E^{<g}\) are rendered onto the visual canvas \(C_{g-1} = \mathcal{R}(E^{<g})\) and encoded with explicit discrete bounding-box tokens, while unplaced elements preserve only category and content features. The multimodal prompt concatenates verbalized descriptions of all elements in sequence, allowing the model \(M_{\text{Layout}} = (f_{\text{enc}}, L_\theta)\) to inspect actual visual density, negative space, and salient regions on the rendered canvas while retaining the complete macro design intent for subsequent elements.
2. Progressive Imitation Learning (PIL): Establishing geometric priors via partial canvas conditioning Applying unconstrained reinforcement learning directly to multi-step spatial coordination invites policy collapse and action space divergence. To provide a stable foundation, PIL trains the policy \(\pi_\theta\) via supervised imitation over ground-truth layout progressions. At each training step \(g\), the first \(g-1\) element groups are rendered using ground-truth coordinates into a partial canvas \(C_{g-1}^{GT}\). Conditioned on this visual canvas, the prompt \(\mathcal{P}_g^E\), and element embeddings \(Z_g\), the model is optimized to predict the full future sequence of element groups \(A_{g:G}^{GT}\) by minimizing the multi-group cross-entropy loss: $\(\mathcal{L}_{\text{PIL}}(\theta) = -\sum_{h=g}^G \log \pi_\theta(A_h \mid C_{g-1}^{GT}, \mathcal{P}_g^E, Z_g)\)$ By training the network to forecast all remaining elements from any arbitrary intermediate state, PIL prevents greedy local placements that would otherwise crowd out space required by late-stage components.
3. Progressive Aesthetic Policy Optimization (PAPO): Online policy updates via ground-truth anchored consensus ranking Standard geometric metrics such as bounding-box IoU cannot distinguish between rigid geometric alignment and genuine graphic aesthetics. To optimize for visual appeal, PAPO establishes an online reinforcement learning loop adapted from Grouped Sequence Policy Optimization (GSPO). For each design prompt, the current policy samples \(m\) distinct layout rollouts \(\mathcal{Z} = \{\zeta_1, \dots, \zeta_m\}\), which are rendered into bitmap images \(L_i = \mathcal{R}(\zeta_i)\). Crucially, the ground-truth layout \(L_{GT}\) is integrated as an anchor into the candidate pool \(\mathcal{L} = \{L_{GT}, L_1, \dots, L_m\}\). An open-source dedicated layout judge, DesignSense, executes bidirectional pairwise comparisons across all pairs in \(\mathcal{L}\) to construct a directed comparison matrix \(M_{ij}\). Aesthetic consensus scores \(S(L_i)\) are computed through iterative power propagation: $\(S(L_i) = (1 - \gamma) \sum_{L_j \in \mathcal{N}^-(L_i)} \frac{S(L_j)}{|\mathcal{N}^+(L_j)|} + \frac{\gamma}{|\mathcal{L}|}\)$ Trajectory-level rewards balance absolute visual ranking against ground-truth fidelity: \(r(\zeta_i) = S(L_i) + \alpha \exp\big(\beta(S(L_i) - S(L_{GT}))\big) + r_{\text{fmt}}\). All future element groups \(\{A_g, \dots, A_G\}\) emerging from canvas state \(C_{g-1}^{GT}\) share the same scalar advantage \(A_g^{\text{adv}} = r(\zeta) - b_g\), providing stable long-horizon credit assignment that encourages global compositional harmony.
Loss & Training¶
The complete optimization pipeline comprises three successive phases: 1. Pretraining Initialization: Fine-tuning the InternVL-3 8B backbone across 148k graphic layout samples for 2 epochs, learning basic spatial token discretization (224 coordinate bins) and text-visual grounding; 2. PIL Supervised Training: Training for 5 epochs on a 75% split of the Crello dataset under random multi-stage progressive canvas conditioning; 3. PAPO Online Alignment: Executing online RL for 1 epoch on the remaining 25% split of Crello, with rollout count \(m=4\), damping coefficient \(\gamma=0.85\), scaling hyperparameters \(\alpha=0.1\) and \(\beta=1\), an effective batch size of 128, distributed across 8 NVIDIA A100 (80GB) GPUs.
Key Experimental Results¶
Main Results¶
Evaluation was conducted on two canonical benchmarks: Crello (professional marketing templates and posters) and WebUI (real-world web interface layouts). Metrics include geometric Mean IoU (%) and pairwise aesthetic Win Rate (%) judged by GPT-4o and o3.
| Dataset | Model | Scale | Mean IoU (↑) | GPT-4o Win Rate (%) | o3 Win Rate (%) |
|---|---|---|---|---|---|
| Crello | FlexDM | - | 12.71 | - | - |
| Crello | LayoutDETR | - | 15.04 | 1.71 | 0.62 |
| Crello | LACE | - | 23.18 | 0.85 | 0.09 |
| Crello | PosterLLaVA | - | 25.18 | 2.19 | 1.01 |
| Crello | LayoutNUWA | - | 25.74 | 4.18 | 1.93 |
| Crello | LaDeCo | - | 35.25 | 14.42 | 8.73 |
| Crello | AesthetiQ | 8B | 42.83 | 15.74 | 8.04 |
| Crello | i-Design | 1B | 23.83 | 2.85 | 1.13 |
| Crello | i-Design | 2B | 30.96 | 6.74 | 1.98 |
| Crello | i-Design | 4B | 43.47 | 17.59 | 10.37 |
| Crello | i-Design | 8B | 48.27 | 27.64 | 19.03 |
| WebUI | Desigen | - | 15.36 | 4.81 | - |
| WebUI | LayoutDETR | - | 16.11 | 5.27 | - |
| WebUI | LACE | - | 17.88 | 5.27 | - |
| WebUI | PosterLLaVA | - | 30.19 | 14.73 | - |
| WebUI | LayoutNUWA | - | 32.16 | 15.28 | - |
| WebUI | LaDeCo | - | 41.16 | 19.49 | - |
| WebUI | AesthetiQ | 8B | 48.29 | 24.48 | - |
| WebUI | i-Design | 1B | 39.24 | 18.03 | - |
| WebUI | i-Design | 2B | 43.18 | 21.40 | - |
| WebUI | i-Design | 4B | 47.86 | 24.35 | - |
| WebUI | i-Design | 8B | 53.71 | 31.29 | - |
Ablation Study¶
1. Reinforcement Learning Strategy and Ground-Truth Anchoring (Crello, i-Design-8B)
| Training Strategy Configuration | Mean IoU (↑) | GPT-4o Win Rate (%) | o3 Win Rate (%) | Mechanism & Note |
|---|---|---|---|---|
| i-Design-8B + PIL | 44.39 | 20.41 | 15.28 | Supervised progressive imitation learning only |
| i-Design-8B + PIL + AAPA | 46.62 | 25.60 | 16.79 | Baseline static preference alignment from prior work |
| i-Design-8B + PIL + PAPO (w/o GT anchor) | 45.19 | 23.04 | 15.39 | Consensus graph ranking without ground-truth anchor |
| i-Design-8B + PIL + PAPO (Full Model) | 48.27 | 27.64 | 19.03 | Online trajectory exploration + GT-anchored consensus graph |
2. Group Size \(k\) and Partitioning Strategy Ablations (Crello, i-Design-8B)
| Variable Category | Configuration | Mean IoU (↑) | GPT-4o Win Rate (%) | o3 Win Rate (%) | Key Takeaway |
|---|---|---|---|---|---|
| Group Size \(k\) | \(k = N\) (single-pass) | 40.28 | 17.64 | 13.59 | Absence of progressive visual feedback degrades quality |
| Group Size \(k\) | \(k = 1\) (single-element) | 47.72 | 27.13 | 18.31 | Excessive turns introduce cumulative latency without gain |
| Group Size \(k\) | \(k = \lceil\sqrt{N}\rceil\) (default) | 48.27 | 27.64 | 19.03 | Optimal balance between local feedback and global context |
| Grouping Strategy | Random (\(\sqrt{N}\) elements) | 43.27 | 15.04 | 6.11 | Incoherent intermediate canvases disrupt visual reasoning |
| Grouping Strategy | Semantic (GPT-4o clustering) | 46.21 | 29.23 | 19.91 | Marginal aesthetic gains but requires costly LLM preprocessing |
| Grouping Strategy | \(z\)-order Deterministic (default) | 48.27 | 27.64 | 19.03 | Standalone deterministic ordering yields superior geometric fidelity |
Key Findings¶
- Progressive conditioning overcomes single-pass bottlenecks: Constraining the group size to \(k=N\) (collapsing into a single forward pass) incurs a 7.99% drop in Mean IoU and a 10.00% reduction in GPT-4o win rate, proving that rendered canvas observations are crucial for collision avoidance and proportional spacing.
- Ground-truth anchoring prevents aesthetic drift: Eliminating ground-truth layouts from the comparison graph (w/o GT) drops mIoU by 3.08% and GPT-4o win rate by 4.60%. Ground-truth references act as topological anchors that stabilize graph ranking and prevent the online policy from converging to degenerate local optima.
- Predictable scaling behavior: Scaling model capacity from 1B to 8B drives monotonic improvements across both benchmarks (Crello mIoU expands from 23.83% to 48.27%, win rate from 2.85% to 27.64%), confirming that larger VLM backbones provide superior visual-spatial reasoning.
- Human preference validation: In a 4-way blind user study, i-Design attained a 52.9% preference share, substantially eclipsing prior SOTA AesthetiQ (40.7%), LayoutNUWA (3.6%), and LACE (2.8%).
Highlights & Insights¶
- Human-centric design formulation: Replaces the simplistic "one-shot coordinate generation" assumption with an embodied "observe-place-reevaluate" paradigm, providing a principled formulation for computer vision-assisted creative design.
- Graph-consensus aesthetic policy optimization: Formulates pairwise VLM visual judge evaluations into a directed comparison graph solved via power iteration, overcoming the high variance and inconsistency common to scalar reward models.
- Transferable modular framework: The concept of \(z\)-order chunking coupled with rendered visual feedback (\(k \approx \lceil\sqrt{N}\rceil\)) can be readily extended to automated UI synthesis, multi-slide deck generation, and document parsing agents.
Limitations & Future Work¶
- Reliance on predetermined stacking order (\(z\)-order): While \(z\)-order provides a reliable heuristic for most graphic compositions, complex overlapping layouts with cyclic visual dependencies may benefit from dynamically learned group scheduling.
- Online rendering latency during training: Online sampling within PAPO requires repeated calls to Skia rendering and forward passes through the visual judge, increasing training wall-clock time compared to offline preference tuning.
- Omission of styling attributes: The current framework focuses primarily on spatial geometry (bounding boxes and dimensions), leaving typography selection, font hierarchies, and palette matching for future multimodal extensions.
Related Work & Insights¶
- vs AesthetiQ: AesthetiQ pioneered visual preference alignment for layouts but operated within a single-shot generation framework using offline pairwise tuning; i-Design demonstrates that combining progressive rollouts with online graph-consensus RL yields a +5.44% boost in mIoU and +11.90% increase in GPT-4o win rate on Crello.
- vs LaDeCo: While LaDeCo adopted step-by-step element placement, it was restricted to supervised imitation learning; i-Design integrates online policy gradient optimization grounded in rendered aesthetic rewards.
- vs LayoutNUWA & PosterLLaVA: These models rely heavily on text-tokenized coordinates with limited or static visual conditioning, often producing element overlaps; i-Design's dynamic rendering loop continuously rectifies spatial interference.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates an innovative fusion of progressive trajectory decision-making and graph-consensus aesthetic RL for graphic layouts.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive cross-benchmark evaluations, rigorous scaling analysis (1B to 8B), four detailed ablation studies, and human perceptual trials.
- Writing Quality: ⭐⭐⭐⭐⭐ Clean mathematical formulation, clear design rationale, and intuitive visual narrative.
- Value: ⭐⭐⭐⭐⭐ Offers an inspiring, highly practical blueprint for creative layout generation and multimodal reinforcement learning.