FlowLess: Controlling Abstract Image Generation¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Image Generation
Keywords: abstract image generation, flow models, visual abstraction set, per-part modulation, minimalist composition
TL;DR¶
FlowLess introduces a self-supervised framework for controllable abstract image generation from a visual abstraction set (disjoint non-semantic primitives and composition details), leveraging geometric augmentations and per-part modulation to learn abstract degrees of freedom, while using a coverage score to suppress generative clutter and synthesize minimalist compositions.
Background & Motivation¶
Modern large-scale text-to-image generative models based on diffusion transformers and rectified flow architectures have demonstrated remarkable capabilities in synthesizing high-fidelity, photorealistic images from descriptive text prompts. However, when faced with non-semantic abstract primitives and the demand for minimalist artistic creation, these models remain tightly constrained by their strong generative priors. As articulated in Leon Battista Alberti's classical architectural treatise in 1485, beauty is that "reasoned harmony of all the parts within a body, so that nothing may be added, taken away, or altered, but for the worse." In stark contrast, modern generative models inherently exhibit a bias toward adding excessive visual clutter, ornamenting compositions and populating the canvas with gratuitous textures and details that users cannot easily suppress.
Existing visual conditioning frameworks predominantly operate on distinct semantic entities or enforce rigid spatial layouts and global style constraints. In abstract art and graphic design—characterized by geometric abstraction, minimalist illustrations, or expressive ink strokes—artists often compose diverse semantic narratives from a collection of disjoint, non-semantic visual primitives (e.g., a few watercolor splashes, loose brush marks, or clean vector blocks). Existing conditioning methods either impose over-restrictive spatial coordinate grids or apply global style transfer that obliterates the discrete geometric identities of the input primitives, entirely lacking a mechanism that allows users to preserve specific shapes while loosely reinterpreting others.
This paper's angle of attack is to rethink the paradigm of conditioning: instead of treating reference inputs as whole-image styles or concrete semantic subjects, decompose them into a non-semantic visual abstraction set and inject them into the flow model's modulation space. Core idea: construct a visual abstraction set of disjoint primitives and relational details, train a flow model via self-supervised geometric augmentations and per-part modulation residuals, and guide the model's expressiveness and minimalist abstraction using a continuous global coverage score.
Method¶
Overall Architecture¶
FlowLess establishes a self-supervised fine-tuning pipeline for rectified flow models. The system begins by decomposing images into visual abstraction sets comprising disjoint visual primitives and localized inter-relation close-up details. During training, the extracted primitives undergo multi-level geometric augmentations alongside random subset dropout to derive a continuous ground-truth coverage ratio. An extended DiT sequence processes text embeddings, noisy latent image patches, and abstraction primitive embeddings jointly. Through modulation layers, learned per-part augmentation residuals and global coverage residuals modulate the intermediate activations, guiding the model's velocity field prediction. At inference time, users can assign distinct generative roles (preserve, free use, or completion) to individual primitives and continuously scale the coverage score to dictate the overall degree of minimalist abstraction.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Sample<br/>SVG Vector / Raster Illustration"] --> B["Self-Supervised Abstraction Set Extraction<br/>Primitive Segmentation + Relation Graph Rendering"]
B --> C["Multi-Level Geometric Augmentation & Dropout<br/>Three-Tier Distraction + Coverage Calculation"]
C --> D["Extended Sequence & Per-Part Modulation<br/>DiT Sequence Concatenation + Residual Injection"]
D --> E["Coverage-Guided Minimalist Control<br/>Global Coverage Modulation + Velocity Prediction"]
E --> F["Output Abstract Image<br/>Semantic Adherence + Pure Primitive Composition"]
Key Designs¶
1. Self-Supervised Abstraction Set Extraction: Automatically Isolating Non-Semantic Primitives and Compositional Details
To circumvent the prohibitive cost of collecting paired datasets of raw abstract primitives and finished artworks, this design implements an automated decomposition and topology extraction pipeline. For vector graphics (SVGs) that inherently encode path and primitive hierarchies, the method directly parses individual shape paths. For rasterized illustrations exhibiting richer artistic textures, the pipeline deploys Segment Anything (SAM) over a dense grid followed by non-maxima suppression, and subsequently repairs occluded regions using relation-aware inpainting. To retain essential compositional information regarding spatial adjacency and material overlay, the method constructs a relation graph connecting primitives that intersect or overlap, rendering circular close-up views centered at their boundaries to serve as visual proxies for inter-element relations.
2. Extended Sequence & Per-Part Modulation: Disentangling Abstraction Degrees of Freedom in Modulation Space
Standard conditioning schemes typically append reference tokens globally, making it difficult to impose heterogeneous constraints on different elements within the same set. FlowLess fine-tunes the Flux.1-dev rectified flow model over extended token sequences where encoded abstraction primitive embeddings are directly concatenated with text and image patch embeddings:
To teach the network how strictly to adhere to the geometry of individual primitives, training applies three tiers of distraction augmentations (fixed position and scale, centered random rotation/scaling, and contour-omitting random crops) alongside relation close-ups. For each primitive, a learned modulation residual \(\Delta_{\text{aug}}\) corresponding to its augmentation category is added to its modulation vector (\(\hat{y}_i = y + \Delta_{\text{aug}}\)). This enables users at test time to assign heterogeneous roles to different primitives—such as strictly locking the geometry of specific shapes (Preserve), flexibly recycling textures and colors into arbitrary object parts (Free-use), or fixing an anchor primitive to synthesize surrounding content (Complete).
3. Coverage-Guided Minimalist Control: Steering Expressiveness and Clutter via a Continuous Variable
Text prompts and reference images fundamentally lack the capacity to specify how much generative ornament a model should introduce. FlowLess resolves this ambiguity through a continuous coverage score mechanism. During training, if an image contains \(A\) total primitives and a random subset of \(n\) elements is sampled, the input covers a fraction \(\tilde{c} = n / A \in [0, 1]\). This continuous scalar is projected into Fourier features and mapped through an MLP to form a global modulation residual \(\Delta_{\text{cov}}\) that modulates the image patch embeddings exclusively. At inference time, setting a low coverage score allows the flow model to draw on its generative priors to populate the canvas, whereas dialing \(\tilde{c}\) toward 1.0 compels the model to construct the concept using strictly the provided primitives, eliminating extraneous visual clutter and producing clean, minimalist compositions.
Loss & Training¶
The framework fine-tunes the pre-trained Flux.1-dev flow model using the standard rectified flow velocity matching objective:
Training employs a cosine schedule with a log-SNR shift of \(2\log(2)\) to allocate higher weight to noisy timesteps. The model is trained on 64 TPU-v3 devices with ZeRO memory partitioning across a batch size of 512 image-abstraction pairs for 4K iterations at a learning rate of \(10^{-4}\). The training dataset integrates 100K curated SVG illustrations (with \(\ge 3\) paths) from MMSVG-Illustration and 50K synthetic raster illustrations generated via Flux.1-dev prompted by Gemini captions.
Key Experimental Results¶
Main Results¶
Evaluation is conducted across 1,000 synthesized images using 100 Gemini-generated conceptual prompts and 10 multi-style abstraction sets (spanning watercolor, ink strokes, marker doodles, and vector primitives). Text alignment is quantified via CLIP score, set consistency is assessed using bidirectional Chamfer similarity over DINOv2 feature embeddings, and perceptual quality is evaluated through a 1-vs-1 two-alternative forced-choice user study with 20 participants (480 total responses).
| Method | Text alignment (CLIP) ↑ | Set alignment (DINO Chamfer) ↑ | User pref. |
|---|---|---|---|
| InstantStyle | 0.185 ± 0.046 | 0.454 ± 0.064 | 78.1% |
| OmniControl | 0.242 ± 0.033 | 0.409 ± 0.061 | 88.3% |
| DreamO | 0.235 ± 0.033 | 0.419 ± 0.062 | 93.5% |
| Qwen-Image-Edit | 0.239 ± 0.029 | 0.419 ± 0.067 | 79.1% |
| Flux.2-dev | 0.239 ± 0.029 | 0.429 ± 0.062 | 93.7% |
| Flux.2-dev + MMPE | 0.207 ± 0.043 | 0.467 ± 0.078 | 80.9% |
| Ours (Free-use) | 0.196 ± 0.045 | 0.475 ± 0.073 | 74.5% |
| Ours (Preserve) | 0.172 ± 0.046 | 0.495 ± 0.065 | — |
Note: User preference indicates the percentage of times human evaluators preferred FlowLess (Free-use) over the respective baseline for faithfully visualizing the concept using the provided visual primitives.
Ablation Study¶
The ablation investigates the trade-offs between coverage score \(\tilde{c}\), conditioning modes, and the synergy of training dataset modalities.
| Training Config / Modality | Coverage Setting \(\tilde{c}\) | Set Alignment (DINO) | Text Alignment (CLIP) | Compositional Characteristics & Failure Analysis |
|---|---|---|---|---|
| SVG-only Dataset | Continuous 0.1 → 1.0 | 0.420 ~ 0.435 | 0.150 ~ 0.170 | Limited artistic styles and textures; yields lower image quality and flat compositions |
| Rasterized-only Dataset | \(\tilde{c} \le 0.4\) | 0.435 ~ 0.455 | 0.220 ~ 0.230 | Decent text adherence, but fails to constrain generation to minimal primitive sets |
| Rasterized-only Dataset | \(\tilde{c} > 0.5\) | Sharp degradation (< 0.44) | 0.160 ~ 0.170 | Lacks training samples with few elements; fails to generalize under sparse primitive budgets |
| Mixed Full Model (Free-use) | \(\tilde{c} = 0.1 \to 1.0\) | 0.445 → 0.475 | 0.228 → 0.196 | Smooth transition from decorated to minimal scenes; balances semantic storytelling with primitive usage |
| Mixed Full Model (Preserve) | \(\tilde{c} = 0.1 \to 1.0\) | 0.465 → 0.495 | 0.210 → 0.172 | Strictly locks primitive contours; attains highest primitive fidelity at peak coverage |
Key Findings¶
- Trade-off Between Expressiveness and Geometric Fidelity: Increasing the coverage score \(\tilde{c}\) forces the model to restrict generation strictly to the provided primitives, raising set alignment up to 0.495 under the Preserve condition. This constraint slightly lowers generic CLIP text scores (from ~0.23 to 0.17–0.19), reflecting the inherent stylistic gap between austere abstract compositions and verbose photorealistic descriptions.
- Dataset Modality Synergy: Training purely on SVG yields clean primitive segmentations but lacks texture richness, while raster-only training is skewed toward high element counts (20–60 primitives), breaking down when tasked with sparse sets (5–8 primitives). Blending both modalities at a 2:1 ratio (100K SVG : 50K raster) is crucial for harmonizing rich artistic texture with sparse primitive abstraction.
- Heterogeneous Per-Part Generative Roles: Per-part modulation successfully enables mixed-constraint synthesis within a single prompt—users can preserve the precise outline of an anchor stroke while allowing the model to freely repurpose circular patches into mechanical wheels or whimsical features.
Highlights & Insights¶
- Operationalizing Classical Aesthetics into Latent Modulations: The paper translates Alberti's architectural philosophy of unadorned structural harmony into a concrete continuous modulation variable, empowering users to explicitly govern the clutter-versus-minimalism slider in generative flow models.
- Self-Supervised Abstraction Decomposition: By combining automated vector extraction, dense grid segmentation, and relation-aware inpainting, the framework bypasses the need for manual abstract art annotations, establishing an effective self-supervised conditioning pipeline.
- Combinatorial Semantic Versatility from Fixed Primitive Kits: Demonstrates that a single set of non-semantic primitives (such as ink strokes or watercolor blobs) can be dynamically re-assembled into vastly different semantic concepts—from animals to musical instruments—providing a fresh paradigm for modular creative design.
Limitations & Future Work¶
- Spatial Reasoning in Highly Complex Compositions: When prompted with complex multi-object scenes or fine-grained spatial actions, the model occasionally struggles to assemble sparse primitives into anatomically coherent or geometrically plausible arrangements.
- Mismatch Under Extreme Primitive Constraints: Forcing an overly simplistic primitive (e.g., a single straight line) to strictly preserve its geometry within an intricate semantic prompt can result in awkward compositional seams or visual incoherence.
- Extension to Temporal Video Dynamics: Future avenues highlighted by the authors include integrating multimodal reasoning models for richer compositional logic and extending abstract primitive conditioning into temporal video generation to control dynamic motion motifs and stylized animations.
Related Work & Insights¶
- vs InstantStyle / OmniControl / DreamO: Existing reference conditioning approaches treat inputs as holistic style exemplars or customized subjects, inadvertently injecting global textural clutter and rigid structural priors. In contrast, FlowLess provides discrete primitive-level control and explicit minimalist restraint.
- vs CLIPDraw / VectorFusion / Word-As-Image: Classical abstract synthesis systems rely on per-instance test-time optimization via differentiable vector graphics, which is computationally sluggish and sensitive to initialization. FlowLess fine-tunes a large-scale rectified flow model to achieve versatile, one-pass generative abstraction.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering framework for controllable abstract image generation from non-semantic primitives with continuous expressiveness regulation.
- Experimental Thoroughness: ⭐⭐⭐⭐☆ Solid empirical evaluation integrating DINO Chamfer similarity, CLIP scores, ablation sweeps, and human forced-choice studies across diverse creative applications.
- Writing Quality: ⭐⭐⭐⭐⭐ Beautifully motivated through architectural and aesthetic theory, with coherent mathematics, clear diagrams, and comprehensive experimental analysis.
- Value: ⭐⭐⭐⭐☆ Highly impactful for graphic design, modular branding, icon design, and the broader study of expressiveness control in generative models.