Semantic Line Diffusion: Character-Consistent Line Art from text-annotated Storyboards¶
Conference: ECCV 2026
Paper: ECCV Official Link
Code: https://huggingface.co/SemanticDiffusion-for-Line-Art-code
Area: LLM (Other)
Keywords: webtoon storyboard line art generation, diffusion transformer, cross-panel character consistency, memory-augmented retrieval, storyboard multimodal conditioning
TL;DR¶
To tackle rough geometry and severe cross-panel character identity drift in webtoon storyboard sketches (conti), this paper introduces PanelDiff, a diffusion transformer framework combining multi-reference static pooling, dynamic memory bank retrieval with a temporal recency kernel, and panel-adaptive gating to generate clean, structure-aligned, and identity-consistent line-art sequences.
Background & Motivation¶
In digital comic and webtoon industrial production pipelines, transforming rough storyboard sketches—known as conti—into clean, polished line art remains an exceptionally labor-intensive process requiring extensive artistic expertise. Creators sketch coarse layouts combining minimalist line drafts, rough bounding shapes, panel arrangements, and textual annotations to plan character actions, framing transitions, and narrative rhythm before painstakingly inking final linework. However, because conti sketches provide only sparse geometric outlines and limited semantic clues while imposing strict panelized sequential narrative constraints, standard generative models struggle to reconstruct fine details without losing structural fidelity.
Existing image-to-image translation models, sketch-to-image diffusion frameworks (such as ControlNet, T2I-Adapter, and DiffSketcher), and text-to-image foundation models process individual images in isolation. When applied to multi-panel storyboards, this isolated per-panel generation paradigm cannot track character visual history across sequential panels. Whenever a character undergoes viewpoint shifts, dynamic body poses, or expressive facial variations between cuts, independent generators suffer from severe "identity drift," altering facial geometry, hair structure, and character traits, thereby destroying narrative immersion.
The fundamental challenge is balancing global canonical identity priors defined by reference character model sheets against dynamic appearance shifts accrued during sequential storytelling. The core idea of this paper is to build PanelDiff, a panel-aware diffusion transformer that unifies panel positional indices, conti sketches, textual descriptions, and character identity tokens synthesized via multi-reference static pooling and dynamic memory retrieval, jointly optimized under diffusion denoising, feature contrastive, and temporal consistency losses to generate character-consistent line-art panel sequences.
Method¶
Overall Architecture¶
PanelDiff takes as input a sequence of conti panel images \(x_t \in \mathbb{R}^{H \times W \times 3}\), textual character annotations \(\{l_c\}_{c \in \mathcal{C}_t}\), and a set of artist-provided static canonical reference line drawings \(\mathcal{R}_c = \{r_c^{(k)}\}\) for each character, autoregressively synthesizing clean line-art panels \(y_t\). The architecture coordinates multimodal feature encoding, dual-source memory pooling and retrieval, panel-adaptive gating, and diffusion transformer denoising.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
InConti["Input: Conti sketch x_t<br/>and text descriptions l_c"] --> EncVisual["Multimodal Encoding<br/>Light ViT for conti / language transformer for text"]
InRefs["Input: Static references R_c<br/>and historical dynamic memory"] --> PoolStatic["Multi-Reference Static Pooling<br/>Conti-context attention weighting"]
InRefs --> RetDyn["Dynamic Memory Retrieval<br/>Exponential temporal recency kernel"]
PoolStatic --> GateFusion["Panel-Adaptive Gating Fusion<br/>Adaptive query weighting and projection"]
RetDyn --> GateFusion
EncVisual --> DiTGen["Panel-Aware DiT Generator<br/>Joint attention over panel/conti/text/identity tokens"]
GateFusion --> DiTGen
DiTGen --> GenPanel["Output: Polished line-art panel y_t"]
GenPanel --> UpdMemory["Dynamic Memory FIFO Maintenance<br/>Shared encoder extracts character crops into memory"]
At inference time, a lightweight Vision Transformer encoder \(E_x\) converts the conti sketch into visual tokens \(Z_{x_t}\), while a language transformer extracts textual embeddings \(e_c^{\text{text}}\). A shared CNN-ViT hybrid line-art encoder \(E_r\) extracts feature representations from canonical reference drawings into static memory entries \(\mathcal{M}_0\) and processes detected character crops from previously generated panels into dynamic memory entries. The similarity interpreter computes context-aware attention over static exemplars and temporally discounted attention over dynamic history, gating both streams into a unified identity token \(\tilde{h}_{t,c}\). The diffusion transformer attends jointly to panel index, conti visual tokens, text tokens, and character tokens to predict reverse diffusion noise and decode clean line art.
Key Designs¶
1. Multi-Reference Character Conditioning: Context-Aware Reference Feature Pooling Single canonical portrait references fail to cover extreme camera angles, side profiles, and dynamic poses common in webtoon panels. PanelDiff maintains a collection of static reference drawings \(\mathcal{R}_c\) per character, from which the shared line-art encoder \(E_r\) extracts static features \(\mathcal{F}_{t,c}^{\text{static}} = \{f_{t,c,i}\}\). To weight the most structurally relevant reference against the current sketch geometry, the model forms a panel context query \(u_{t,c} = W_u [g_t \,;\, e_c^{\text{text}}]\) using the mean-pooled conti visual representation \(g_t\) and textual attribute embedding \(e_c^{\text{text}}\). The relevance of each static reference is computed via scaled dot-product attention: $$ \alpha_{t,c,i}^{\text{static}} = \frac{\exp(u_{t,c}^\top W_{\text{st}} f_{t,c,i})}{\sum_{f_{t,c,k} \in \mathcal{F}{t,c}^{\text{static}}} \exp(u $$ The pooled static feature }^\top W_{\text{st}} f_{t,c,k})\(s_{t,c} = \sum_i \alpha_{t,c,i}^{\text{static}} f_{t,c,i}\) selectively emphasizes the reference drawing matching the viewpoint and emotional tone of the current frame, preventing unnatural feature warping when facing non-frontal views.
2. Memory-Augmented Dynamic Bank: Temporal Recency-Discounted Sequential Retrieval As a visual narrative unfolds, subtle stylistic details, temporary costume alterations, and expressive line weights evolve across panels. When each panel \(\hat{y}_t\) is synthesized, character crops \(\Omega_{t,c}\) are encoded into dynamic features \(f_{t,c}^{\text{dyn}} = E_r(\hat{y}_t|_{\Omega_{t,c}})\) and stored alongside metadata \(\text{meta}_{t,c}\) (panel index, bounding box, confidence score, pose/viewpoint, and text). To bound memory growth, a per-character FIFO buffer is capped at \(K_{\text{mem}} = 64\) entries. To reflect the narrative principle that recent frames bear greater appearance relevance than distant cuts, retrieval incorporates an exponential temporal recency kernel \(\kappa(\Delta t_{t,i}) = \exp(-\Delta t_{t,i} / \lambda_{\text{time}})\), where \(\Delta t_{t,i} = t - t_i\). Dynamic attention weights combine semantic similarity and temporal proximity: $$ \beta_{t,c,i}^{\text{dyn}} = \frac{\exp(u_{t,c}^\top W_{\text{dy}} f_{t,c,i}) \cdot \kappa(\Delta t_{t,i})}{\sum_{f_{t,c,k} \in \mathcal{F}{t,c}^{\text{dyn}}} \exp(u $$ Aggregating dynamic entries into }^\top W_{\text{dy}} f_{t,c,k}) \cdot \kappa(\Delta t_{t,k})\(d_{t,c} = \sum_i \beta_{t,c,i}^{\text{dyn}} f_{t,c,i}\) provides the network with temporal context, ensuring smooth stroke continuity across sequential cuts.
3. Panel-Adaptive Gating: Dynamic Arbitration and Identity Token Construction A generator must flexibly balance canonical invariants with episodic appearance variations: during abrupt scene transitions or large camera cuts, over-relying on recent dynamic frames risks propagating transient pose artifacts, necessitating higher weight on static references; conversely, in continuous dialogues, dynamic memory is paramount to eliminate flicker. The similarity interpreter predicts an adaptive gating scalar \(\eta_{t,c} = \sigma(v^\top [g_t \,;\, e_c^{\text{text}}])\) to interpolate between static and dynamic features: $$ m_{t,c} = \eta_{t,c} s_{t,c} + (1 - \eta_{t,c}) d_{t,c} $$ This integrated memory vector \(m_{t,c}\) is concatenated with the canonical static embedding \(\bar{h}_{t,c}\) and mapped through an MLP projection \(\psi(\cdot)\) to form the final character token \(\tilde{h}_{t,c} = \psi([\bar{h}_{t,c} \,;\, m_{t,c}])\) provided to the diffusion transformer.
4. Panel-Aware DiT Generator: Structured Conditioning and Multi-Objective Supervision The generative backbone is built on a Diffusion Transformer. For panel \(t\), the input sequence is assembled by concatenating panel position token \(p_t\), conti visual tokens \(Z_{x_t}\), text tokens \(Z_t^{\text{text}}\), and the character identity tokens \(\{\tilde{h}_{t,c}\}\) into \(Z_{\text{int},s}\). At timestep \(s\), the transformer estimates noise \(\hat{\epsilon}_t^{(s)} = F_\theta(z_t^{(s)}, Z_{\text{int},s}, \gamma(s))\) on noisy latent \(z_t^{(s)}\). In addition to standard diffusion noise prediction loss \(\mathcal{L}_{\text{diff}}\), the architecture is supervised by a triplet feature contrastive loss \(\mathcal{L}_{\text{id}} = \mathbb{E}[\max(0, \|f_{t,c} - f_{t',c}\|_2^2 - \|f_{t,c} - f_{t,c}^{\text{neg}}\|_2^2 + \delta)]\) to enforce identity discriminability, and a temporal smoothness loss penalizing adjacent feature fluctuations: $$ \mathcal{L}{\text{cons}} = \mathbb{E}|_2 \right] $$ The composite objective } \left[ |f_{t,c} - f_{t-1,c\(\mathcal{L} = \mathcal{L}_{\text{diff}} + \lambda_{\text{id}} \mathcal{L}_{\text{id}} + \lambda_{\text{cons}} \mathcal{L}_{\text{cons}}\) drives pixel fidelity, identity clustering, and cross-panel stability simultaneously.
Loss & Training¶
The model is trained at \(512 \times 512\) resolution using a Diffusion Transformer backbone with a Cosine noise schedule. Optimization is performed using AdamW with an initial learning rate of \(10^{-3}\) (decayed to \(10^{-5}\) upon training instability), batch size 64, for 100 epochs. All experiments run on distributed nodes containing eight NVIDIA RTX A6000 GPUs (48GB VRAM each) and 128GB RAM. Static reference embeddings are precomputed to maximize training throughput.
Key Experimental Results¶
Main Results¶
Evaluated on the newly introduced Conti–LineArt dataset comprising 100,000 paired panels crafted by 30 professional webtoon artists across roughly 150 recurring characters, PanelDiff is compared against image translation, sketch control, and story generation baselines (mean over 5 random seeds with standard deviations in parentheses).
| Model | LPIPS ↓ | FID ↓ | MS-SSIM ↑ (Face) | PSNR ↑ | SSIM ↑ | Edge IoU ↑ | Edge F1 ↑ | ID Sim ↑ |
|---|---|---|---|---|---|---|---|---|
| CycleGAN | 0.345 (0.021) | 91.7 (2.40) | 0.812 (0.018) | 18.3 (0.03) | 0.731 (0.020) | 0.54 (0.02) | 0.58 (0.02) | 0.71 (0.02) |
| Pix2PixHD | 0.312 (0.019) | 78.4 (2.10) | 0.835 (0.017) | 19.7 (0.02) | 0.752 (0.019) | 0.58 (0.02) | 0.61 (0.02) | 0.74 (0.02) |
| SPADE | 0.287 (0.018) | 72.1 (1.90) | 0.842 (0.016) | 20.6 (0.02) | 0.768 (0.018) | 0.60 (0.02) | 0.64 (0.02) | 0.76 (0.02) |
| Stable Diffusion | 0.258 (0.016) | 61.3 (1.70) | 0.857 (0.015) | 21.9 (0.02) | 0.781 (0.017) | 0.64 (0.02) | 0.67 (0.02) | 0.79 (0.02) |
| ControlNet | 0.233 (0.015) | 55.1 (1.50) | 0.864 (0.014) | 22.6 (0.02) | 0.792 (0.016) | 0.66 (0.02) | 0.69 (0.02) | 0.81 (0.02) |
| T2I-Adapter | 0.225 (0.014) | 53.2 (1.40) | 0.868 (0.014) | 23.1 (0.02) | 0.799 (0.015) | 0.67 (0.02) | 0.70 (0.02) | 0.82 (0.02) |
| IP-Adapter | 0.216 (0.014) | 51.4 (1.30) | 0.874 (0.013) | 23.6 (0.02) | 0.807 (0.015) | 0.68 (0.02) | 0.72 (0.02) | 0.83 (0.02) |
| DiffSketcher | 0.208 (0.013) | 49.8 (1.20) | 0.881 (0.012) | 24.0 (0.02) | 0.815 (0.014) | 0.69 (0.02) | 0.73 (0.02) | 0.84 (0.02) |
| InstructPix2Pix | 0.241 (0.016) | 64.4 (1.80) | 0.853 (0.015) | 21.2 (0.03) | 0.776 (0.017) | 0.63 (0.02) | 0.66 (0.02) | 0.78 (0.02) |
| GPT-4o | 0.274 (0.017) | 69.2 (2.00) | 0.846 (0.016) | 20.4 (0.03) | 0.763 (0.018) | 0.61 (0.02) | 0.64 (0.02) | 0.76 (0.02) |
| Ours (full) | 0.162 (0.011) | 39.8 (0.90) | 0.903 (0.010) | 26.8 (0.01) | 0.861 (0.011) | 0.74 (0.01) | 0.80 (0.01) | 0.91 (0.01) |
Sequence-level consistency and computational cost evaluated over 50 consecutive panel episodes:
| Model | Identity Drift ↓ | Sequence MS-SSIM ↑ | Layout IoU ↑ | FLOPs (G) |
|---|---|---|---|---|
| CycleGAN | 0.482 ± 0.021 | 0.33 ± 0.02 | 0.41 ± 0.01 | 23 |
| Pix2PixHD | 0.451 ± 0.018 | 0.37 ± 0.02 | 0.45 ± 0.01 | 69 |
| SPADE | 0.429 ± 0.017 | 0.40 ± 0.02 | 0.48 ± 0.01 | 92 |
| Stable Diffusion | 0.377 ± 0.014 | 0.44 ± 0.02 | 0.52 ± 0.01 | 150 |
| ControlNet | 0.355 ± 0.013 | 0.47 ± 0.02 | 0.56 ± 0.01 | 188 |
| T2I-Adapter | 0.346 ± 0.012 | 0.49 ± 0.01 | 0.58 ± 0.01 | 152 |
| IP-Adapter | 0.337 ± 0.012 | 0.51 ± 0.01 | 0.59 ± 0.01 | 149 |
| DiffSketcher | 0.328 ± 0.011 | 0.52 ± 0.01 | 0.60 ± 0.01 | 34 |
| InstructPix2Pix | 0.392 ± 0.016 | 0.43 ± 0.02 | 0.50 ± 0.01 | 160 |
| GPT-4o | 0.371 ± 0.015 | 0.45 ± 0.02 | 0.53 ± 0.01 | — |
| Ours (full) | 0.214 ± 0.009 | 0.68 ± 0.01 | 0.71 ± 0.01 | 41 |
Ablation Study¶
Component contributions evaluated across 50-cut episodes:
| Config | MS-SSIM ↑ | Drift ↓ | Layout IoU ↑ | ID Sim ↑ | Edge IoU ↑ | Note |
|---|---|---|---|---|---|---|
| Ours (full) | 0.903 (0.002) | 0.214 (0.009) | 0.71 (0.01) | 0.91 (0.01) | 0.74 (0.01) | full model |
| w/o Multi-Reference | 0.861 (0.003) | 0.301 (0.012) | 0.65 (0.01) | 0.84 (0.01) | 0.67 (0.01) | removes static reference pooling (+40.7% drift) |
| Single Reference | 0.878 (0.003) | 0.268 (0.011) | 0.67 (0.01) | 0.86 (0.01) | 0.69 (0.01) | single frontal reference image only |
| w/o Memory Bank | 0.842 (0.004) | 0.332 (0.015) | 0.62 (0.01) | 0.82 (0.01) | 0.64 (0.01) | removes dynamic memory entirely (largest drop) |
| Recent-Only Memory | 0.874 (0.003) | 0.259 (0.010) | 0.66 (0.01) | 0.87 (0.01) | 0.69 (0.01) | conditions only on panel \(t-1\) |
| w/o Similarity Interpreter | 0.851 (0.004) | 0.317 (0.013) | 0.63 (0.01) | 0.83 (0.01) | 0.65 (0.01) | removes temporal kernel and gating (uniform avg) |
Under unseen character and unseen artist splits, PanelDiff maintains superior robustness, scoring 0.803 / 0.792 in MS-SSIM and 0.271 / 0.284 in Drift, significantly outperforming DiffSketcher (0.715 / 0.352) and IP-Adapter (0.712 / 0.363).
Key Findings¶
- Dynamic memory is the primary driver of sequence consistency: Removing the memory bank increases drift from 0.214 to 0.332 (+55.1%) and degrades face MS-SSIM to 0.842. Conditioned on only the immediate previous frame (Recent-Only), drift improves to 0.259 but errors compound over 50 cuts, demonstrating the necessity of the FIFO memory buffer.
- Multi-reference exemplars anchor non-frontal geometry: Increasing static references per character from \(K=1\) to \(K=4\) lifts face MS-SSIM from 0.78 to 0.86 and reduces drift from 0.32 to 0.27, proving that varied viewpoints eliminate facial distortion under extreme angles.
- Robustness against rough sketches: Under artificial stroke perturbations (noise level 1.0 with stroke jitter and missing lines), baseline average MS-SSIM collapses to 0.48, whereas PanelDiff sustains 0.79, highlighting its resilience to real-world messy artist contis.
- Exceptional computational efficiency: Operating at only 41 GFLOPs per \(512 \times 512\) panel, PanelDiff is over 70% lighter than ControlNet (188 GFLOPs) and Stable Diffusion (150 GFLOPs), enabling rapid workstation deployment.
Highlights & Insights¶
- Formalizing a key industrial bottleneck: Establishes the first formal benchmark and dataset for webtoon storyboard-to-line-art generation, bridging general image synthesis and professional sequential comic production.
- Decoupled static-dynamic identity coordination: By marrying canonical character reference pooling with temporal-recency dynamic memory retrieval, the model resolves the long-standing trade-off between global identity adherence and localized visual continuity.
- Practical industry-ready efficiency: Delivering high-quality line art at 41 GFLOPs with overwhelming preference in a 30-artist blind evaluation makes it an immediate candidate for comic studio integration.
Limitations & Future Work¶
- Multi-character crowding and occlusion: Cropping character regions \(\Omega_{t,c}\) using bounding box detectors can introduce visual contamination when multiple characters closely overlap or embrace.
- Long-arc episodic attire changes: If a character alters attire and reappears after 60+ panels, the FIFO buffer (\(K_{\text{mem}}=64\)) may have evicted earlier contextual cues, warranting hierarchical or episodic long-term memory structures.
- Future directions: Integrating Conti-to-LineArt with automated flat coloring and cel-shading pipelines to achieve end-to-end storyboard-to-finished-webtoon automation.
Related Work & Insights¶
- vs ControlNet / T2I-Adapter: While adapter-based architectures inject sketch constraints, their independent per-frame inference lacks temporal memory, leading to severe character facial drift across sequential panels; PanelDiff introduces dynamic memory and temporal kernel retrieval to guarantee continuity.
- vs StoryDiffusion / StoryImager: General visual story models focus on loose pictorial coherence in natural images without satisfying the precise, stroke-level contour alignment demanded by professional black-and-white comic inking. PanelDiff employs a dedicated CNN-ViT line encoder and edge IoU alignment losses tailored for digital comics.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates the webtoon conti-to-line-art task with an elegant dual-memory DiT framework.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 100k panels, 50-cut sequence drift, out-of-domain splits, and professional artist user studies.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulations, structured algorithms, and rigorous motivation.
- Value: ⭐⭐⭐⭐⭐ Significant practical value in automating repetitive digital illustration labor for the global webtoon industry.