High-Resolution Artwork Outpainting with Global Blueprint Guidance and Layout Control¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/poohoh/BlueOut
Area: Image Generation
Keywords: Image Outpainting, Artworks, Layout Control, Blueprint Guidance, Diffusion Model
TL;DR¶
This paper presents BlueOut, a two-stage diffusion framework for high-resolution artwork outpainting that plans global composition with bounding-box layout control in Stage 1 and synthesizes high-resolution local patches in parallel in Stage 2 using forward diffusion low-frequency initialization, achieving superior visual fidelity, explicit spatial control, and a 2.4× speedup.
Background & Motivation¶
Image outpainting expands an image beyond its original canvas boundaries, requiring seamless integration that preserves style, color palette, and semantic context rather than merely stretching pixels. In heritage artwork applications—such as large-scale immersive digital exhibitions, projection mapping, and VR/AR displays—legacy paintings must often be adapted to novel canvas sizes or extreme aspect ratios while strictly preserving authentic aesthetic harmony and compositional balance. Unlike image inpainting where missing regions are surrounded by known pixel constraints, outpainting synthesizes unbounded exterior regions, introducing substantially higher uncertainty and requiring a rigorous grasp of global scene harmony.
Existing diffusion-based high-resolution outpainting approaches primarily rely on progressive expansion (e.g., ProOut), iteratively generating and stitching content using sliding local windows. However, this paradigm suffers from three critical bottlenecks: first, the lack of an explicit global planning stage causes inconsistencies in early windows to propagate across successive steps, resulting in severe error accumulation and structural collapse at high resolutions; second, relying strictly on global text prompts precludes fine-grained spatial control, making it impossible to position specific objects (such as crowns, human figures, or architectural elements) at user-designated locations; third, inherently sequential window generation incurs extreme inference latency, creating a major hardware bottleneck where parallel compute cannot be exploited even when memory is abundant.
The core insight of this work is that high-resolution outpainting demands macro-level global planning to coordinate micro-level detail rendering. By decoupling the task into global structural planning and high-resolution local synthesis, a low-resolution global blueprint can establish a reliable compositional anchor. Furthermore, leveraging the low-frequency preservation property of forward diffusion allows this blueprint to initialize structurally coherent noise across all local patches simultaneously. Core idea: decouple high-resolution artwork outpainting into a global blueprint-guided two-stage framework, where Stage 1 plans macro-structure and bounding-box instances via a Layout Adapter at low resolution, and Stage 2 synthesizes local high-resolution patches in parallel across GPUs using forward diffusion low-frequency initialization and global feature guidance.
Method¶
Overall Architecture¶
BlueOut decouples the outpainting workflow into two cooperative stages: Stage 1 performs macro structural planning and layout injection at full canvas scale, while Stage 2 executes parallel, fine-grained high-resolution patch synthesis.
In Stage 1, the system receives the source artwork, the target canvas mask, and user-specified bounding-box layout conditions. After optimizing initial noise via Attention Alignment Measure (AAM), a Layout Adapter and Gated Fuser inject instance tokens into a frozen 9-channel SD Inpainting U-Net backbone. This generates a low-resolution global blueprint \(\hat{x}^g\) and populates a timestep-wise global guidance feature bank \(\mathcal{F}\).
In Stage 2, the high-resolution canvas is partitioned into eight overlapping local patches surrounding the source image. The Stage 1 blueprint is upsampled and cropped into corresponding patches, which are perturbed via forward diffusion using a unified global noise map to construct structurally aligned initial noise states. The patches are then distributed across separate GPUs for parallel denoising. During Stage 2 denoising, each local patch receives both the cached guidance features \(\mathcal{F}\) from Stage 1 and its global canvas position token \(p_i\) through a Gated Fuser. At each step, latent-space overlap averaging ensures seamless boundary transitions, followed by final pixel-space compositing and verbatim source-region re-pasting.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Source Image + Expansion Mask + Bounding Box Layout"] --> S1["Global Blueprint Generation & Instance Layout Injection<br/>AAM noise optimization + Layout Adapter instance tokens"]
S1 --> Blueprint["Low-Resolution Blueprint ˆx_g + Guidance Feature Bank F"]
Blueprint --> InitNoise["Structurally Aligned Initial Noise Construction<br/>Blueprint upsample-crop + Forward diffusion low-frequency retention + Global shared noise"]
InitNoise --> S2["Space-Aware Parallel Local Patch Synthesis<br/>Global feature F and position token p_i injection, multi-GPU denoising"]
S2 --> Blend["Latent Overlap Blending + Pixel Canvas Compositing with Source Pasting"]
Blend --> Out["High-Resolution Seamless Artwork Outpainting Result"]
Key Designs¶
1. Global Blueprint Generation & Instance Layout Injection: Decoupled Planning with Precise Spatial Grounding Progressive outpainting lacks global context, inevitably drifting into distorted textures or repetitive frame-like artifacts. Stage 1 addresses this by operating over the entire canvas at low resolution using a frozen SD Inpainting U-Net. Before denoising, the Attention Alignment Measure (AAM) optimizes the initial noise \(z_T^g\) by ensuring spatial tokens in target regions adequately attend to source tokens in the first denoising step, proactively preventing semantic disconnection. To enable precise spatial control, a Fourier-feature Layout Adapter encodes user-specified bounding boxes \(b_k\) and CLIP text descriptions \(e_k = \mathcal{T}(d_k)\) into instance tokens \(c_k = \text{MLP}([\gamma(b_k); e_k])\), collected into condition matrix \(C\). A zero-initialized Gated Fuser inserted into each transformer block then dynamically fuses \(C\) into intermediate U-Net features \(H_t^g\): $\(F_t^g = H_t^g + \alpha_g \cdot \text{Attn}\Big(Q(H_t^g), K([H_t^g; C]), V([H_t^g; C])\Big)\)$ where \(\alpha_g\) is a learnable gate. Denoising features \(F_t^g\) are cached across all steps into feature bank \(\mathcal{F} = \{F_t^g\}_{t=1}^T\), and clean latent \(z_0^g\) is decoded into global blueprint \(\hat{x}^g = \mathcal{D}(z_0^g)\).
2. Structurally Aligned Initial Noise Construction via Forward Diffusion: Bridging Macro Plan to Micro Init Standard cascaded architectures require separate super-resolution networks, which add training complexity and fail to exploit the high-resolution generative prior of the base diffusion model. BlueOut leverages the low-frequency preservation property of forward diffusion: under typical noise schedules, high-frequency details decay rapidly while macroscopic layout and coarse geometry remain intact even at terminal step \(T\). The decoded blueprint \(\hat{x}^g\) is upsampled to target resolution, cropped into eight local patches, and encoded into latents \(z_{\text{bp}}^{(i)} = \mathcal{E}(\text{Crop}_i(\text{Up}(\hat{x}^g)))\). A single global Gaussian noise map \(\epsilon^{\text{full}} \sim \mathcal{N}(0, I)\) is sampled over the entire canvas, and patch initial noise is constructed via: $\(z_{T, \text{init}}^{(i)} = \sqrt{\bar{\alpha}_T} z_{\text{bp}}^{(i)} + \sqrt{1 - \bar{\alpha}_T} \text{Crop}_i(\epsilon^{\text{full}})\)$ Using crops from a shared global noise map guarantees identical random perturbations along overlapping borders, while the preserved low frequencies ensure all parallel patches begin denoising from a unified, blueprint-aligned geometric layout.
3. Space-Aware Parallel Local Patch Synthesis: Eliminating Sequential Latency Bottlenecks To overcome the sequential latency of progressive sliding windows, Stage 2 synthesizes all local patches simultaneously across multiple GPUs. To prevent parallel branches from drifting semantically during mid-to-late denoising stages, Stage 2 introduces joint spatial coordinate encoding and global feature conditioning. For each patch \(i\), a Position Network encodes its canvas bounding box \(b_i^p\) and a learnable null token \(n_\emptyset\) into position token \(p_i = \text{MLP}([\gamma(b_i^p); n_\emptyset])\). The Stage 2 Gated Fuser integrates local patch features \(H_t^{(i)}\) with Stage 1 guidance features \(F_t^g\) and position token \(p_i\): $\(\tilde{H}_t^{(i)} = H_t^{(i)} + \alpha_p \cdot \text{Attn}\Big(Q(H_t^{(i)}), K([H_t^{(i)}; F_t^g; p_i]), V([H_t^{(i)}; F_t^g; p_i])\Big)\)$ Here, \(p_i\) informs the network of the patch's absolute canvas location, while \(F_t^g\) supplies global scene context. At each step, latent values across overlapping margins are averaged across overlapping sets \(\mathcal{R}(r)\) for target pixels, ensuring seam-free synthesis before final pixel compositing and source re-pasting.
Loss & Training¶
Both Stage 1 and Stage 2 freeze the base SD Inpainting U-Net weights, training only the auxiliary Layout Adapter, Position Network, and Gated Fusers. Standard noise prediction MSE objectives supervise both stages: $\(\mathcal{L}_{\text{S1}} = \mathbb{E}_{z_0^g, \epsilon, t} \left[ \left\| \epsilon - \epsilon_{\theta_1}(z_t^g, t; \tilde{x}^g, m^g, \mathcal{B}, c_G) \right\|_2^2 \right]\)$ $\(\mathcal{L}_{\text{S2}} = \mathbb{E}_{i, z_0^{(i)}, \epsilon, t} \left[ \left\| \epsilon - \epsilon_{\theta_2}(z_t^{(i)}, t; \tilde{x}^{(i)}, m^{(i)}, F_t^g, p_i, c_L) \right\|_2^2 \right]\)$ Stage 1 and Stage 2 are trained independently; Stage 1 remains frozen as a feature bank extractor when optimizing Stage 2. The model is trained on HumanArt, WikiArt, and LAION-5B high-resolution subsets with AdamW (learning rate \(5 \times 10^{-5}\), batch size 512). Inference employs DDIM sampling with 30 steps and CFG = 3.0.
Key Experimental Results¶
Main Results¶
On the IconArt benchmark with default masking ratio \(\rho = 0.333\) (~225% total expansion relative to source), BlueOut was evaluated against state-of-the-art outpainting baselines across visual fidelity, layout accuracy, and runtime.
| Methods | FID ↓ | pFID256 ↓ | pFID512 ↓ | CLIP-S ↑ | CLIP-A ↑ | AP ↑ | IoU ↑ | Time (s) ↓ |
|---|---|---|---|---|---|---|---|---|
| Ground Truth (GT) | – | – | – | – | – | 0.5903 | 0.7716 | – |
| PQDiff (copy) | 25.3032 | 41.1501 | 31.8283 | 0.2031 | 5.0144 | 0.2903 | 0.4304 | – |
| SD Inpainting | 10.7598 | 9.5336 | 7.6219 | 0.2007 | 6.7041 | 0.3526 | 0.5539 | 13.43 |
| PowerPaint | 10.5374 | 9.4403 | 7.4834 | 0.2003 | 6.4423 | 0.3537 | 0.5503 | 13.99 |
| ProOut | 10.3032 | 9.2854 | 7.2515 | 0.2052 | 6.8585 | 0.3595 | 0.5623 | 18.31 |
| BlueOut (Ours) | 9.3064 | 9.0635 | 6.7514 | 0.2033 | 6.7949 | 0.4336 | 0.6382 | 7.50 |
| BlueOut (Ours*) | 9.0729 | 9.0577 | 6.6609 | 0.2055 | 6.8766 | 0.4363 | 0.6361 | 7.50 |
Note: Ours* extracts Stage 1 guidance features after the cross-attention block to better incorporate text prompt conditions. Inference runtime represents average seconds per image across 100 test samples in an 8-GPU parallel setup.
Ablation Study¶
Ablation experiments analyze the contribution of each core component responsible for transferring blueprint guidance into parallel Stage 2 synthesis:
| Config | FID ↓ | pFID256 ↓ | pFID512 ↓ | CLIP-S ↑ | CLIP-A ↑ | Note |
|---|---|---|---|---|---|---|
| SD Inpainting (Parallel) | 37.2675 | 17.6093 | 20.3705 | 0.1932 | 5.8696 | Parallel synthesis with latent blending only, lacking global guidance |
| w/o Guidance Feature | 22.2319 | 12.6436 | 11.7879 | 0.1940 | 5.8561 | Lacking global semantic features causes structural discord across patches |
| w/o Patch Position Token | 17.9578 | 10.7008 | 8.7402 | 0.1990 | 6.2209 | Absence of canvas coordinate priors causes spatial layout drift |
| w/o Forward Diffusion Init | 9.6048 | 9.1144 | 6.9442 | 0.2020 | 6.7598 | Pure Gaussian noise start; cross-attention guidance recovers most structure |
| Full Model | 9.3064 | 9.0635 | 6.7514 | 0.2033 | 6.7949 | Optimal synergy across global features, position tokens, and low-frequency noise |
Key Findings¶
- Guidance features and spatial position tokens are critical for parallel synthesis: Simply applying latent overlap blending (SD Inpainting parallel) fails drastically (FID 37.27). Dropping guidance features causes FID to surge from 9.31 to 22.23, proving that independent local patches cannot maintain contextual consistency through boundary smoothing alone without macro-level feature guidance.
- Robustness under extreme expansion ratios: On landscape art collections from the Cleveland Museum of Art and Art Institute of Chicago, expanding canvas area from 200% to 600% degraded ProOut's FID from 56.22 to 89.84, whereas BlueOut maintained much better stability, shifting from 44.87 to 81.86. Establishing a macro blueprint upfront effectively insulates the synthesis from progressive error runaway.
- 2.4× acceleration via parallel multi-GPU dispatch: In an 8-GPU cluster, distributing the eight patches enables an inference latency of 7.50 s per image versus ProOut's sequential 18.31 s (a 2.44× speedup). Even in a single-GPU serial emulation (15.73 s), BlueOut remains faster than ProOut.
Highlights & Insights¶
- Creative reuse of forward diffusion low-frequency retention: Rather than treating terminal timestep \(t=T\) as pure uninformative noise, BlueOut capitalizes on the FreeInit insight that coarse geometry survives heavy Gaussian corruption. This provides a lightweight, zero-parameter structural alignment across parallel local patches without dedicated super-resolution networks.
- First explicit layout control for artwork outpainting: Prior outpainting frameworks relied strictly on global text prompts, failing to achieve controlled placement of complex objects. Injecting instance bounding boxes into the global planning phase establishes spatial layout precision (AP 0.4363 vs. 0.3595 baseline).
- Decoupled "planning-then-rendering" paradigm: Decoupling macroscopic structural composition from microscopic patch synthesis eliminates the sequential dependencies of progressive sliding windows while aligning seamlessly with modern distributed multi-GPU hardware.
Limitations & Future Work¶
- Static patch partitioning on extreme aspect ratios: The current default layout partitions the canvas into eight symmetrical patches around the center. For extreme aspect ratios, such as traditional Chinese or Japanese handscrolls (emakimono), dynamic adaptive grid partitioning and patch communication topologies will be required.
- Inter-GPU communication synchronization: Performing latent overlap averaging at every denoising step requires frequent inter-device communication, which may become an I/O bottleneck in high-latency or bandwidth-constrained distributed clusters.
- Complex perspective and directional lighting constraints: While bounding boxes effectively constrain object position and scale, highly stylized artworks with strong multi-source illumination or non-linear perspective may still exhibit subtle lighting discrepancies between distant patches.
Related Work & Insights¶
- vs ProOut (Song et al., ICCV 2025): ProOut is a pioneering progressive artwork outpainting framework using sliding windows and intermediate global hints. BlueOut demonstrates that sequential expansion suffers from error accumulation and high latency; by decoupling global blueprint planning from parallel local synthesis, BlueOut achieves superior visual quality, 2.4× faster inference, and layout controllability.
- vs MultiDiffusion (Bar-Tal et al., ICML 2023) / Follow-Your-Canvas (Chen et al., AAAI 2025): MultiDiffusion smooths overlapping boundaries via latent averaging but lacks top-down global semantic control. BlueOut shows that boundary blending alone produces poor high-resolution extensions unless paired with low-frequency noise initialization and blueprint feature guidance.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Elegant two-stage blueprint guidance combined with forward diffusion low-frequency noise initialization for parallel outpainting]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive quantitative benchmarks across resolutions, expansion scales, layout accuracy metrics, and comprehensive ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear structural progression, well-motivated problem formulation, and explicit architectural exposition]
- Value: ⭐⭐⭐⭐⭐ [Addresses the core latency and error accumulation bottlenecks of high-resolution outpainting with high practical value for digital museum preservation and immersive displays]