InstanceControl: Controllable Complex Image Generation without Instance Labeling¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
PDF: ECCV Official PDF
Code: To be released (declared in paper)
Area: Image Generation
Keywords: Controllable Image Generation, Multi-Instance Generation, Vision-Language Model, Adaptive Mask Refinement, Cross-Modal Attention Modulation
TL;DR¶
InstanceControl addresses attribute confusion and the heavy burden of manual instance labeling in complex multi-instance controllable image generation by using a VLM to automatically establish text-visual instance correspondences alongside an adaptive mask refinement strategy that dynamically corrects noisy masks during diffusion denoising.
Background & Motivation¶
Controllable text-to-image generation frameworks like ControlNet introduce spatial visual conditions such as depth maps and edge contours to guide image synthesis, expanding generative models across animation, computer-aided sketching, and image inpainting. However, existing controllable generation models are predominantly reliable only in simple scenes with a limited number of instances. As the count and density of instances grow, these models frequently suffer from attribute confusion and color bleeding across adjacent entities—for instance, when several characters stand side-by-side, the color of one person's coat frequently spills onto a neighboring character's trousers.
To alleviate attribute confusion, recent methods such as FineControlNet, DreamRenderer, and Seg2Any incorporate instance-level spatial annotations, requiring users at test time to supply separate descriptions along with precise bounding boxes or pixel-wise segmentation masks for every object. While effective, this requirement imposes a heavy and labor-intensive annotation burden that limits practicality in real-world creative workflows. The root bottleneck behind the failure of fully automatic controllable models is their inability to accurately associate multi-entity textual descriptions with their corresponding geometric regions in the visual condition. Humans perform this cross-modal association intuitively, and the rich reasoning and grounding capabilities of modern Vision-Language Models (VLMs) present an ideal avenue to bridge this gap automatically.
This paper tackles the challenge by adopting a fine-tuned VLM to parse instance-level noun phrases and predict corresponding condition masks without any human intervention, followed by a diffusion-level dynamic calibration. Core idea: develop a two-stage framework comprising instance-level text-visual condition association and instance-aware controllable generation, where a fine-tuned VLM automatically resolves instance correspondences from raw prompts and conditions, and an adaptive mask refinement module dynamically rectifies noisy masks using diffusion cross-attention to enforce isolated token interactions.
Method¶
Overall Architecture¶
The InstanceControl pipeline consists of two primary stages: instance-level text-visual condition association (Stage 1) and instance-aware controllable generation (Stage 2). In the first stage, the system takes the global prompt \(p\) and visual condition \(c\) (e.g., Canny edges, depth map, or HED edges), employing a Sa2VA-based vision-language backbone to autoregressively parse instance phrases while predicting initial masks via a shared SEG token aggregation mechanism, yielding correspondence tuples \(\mathcal{C} = \{(\mathbf{t}_i^{pred}, \mathbf{m}_i^{pred}, s_i^{pred})\}_{i=1}^N\). In the second stage, because initial VLM masks may contain boundary inaccuracies or localization jitter, a lightweight Mask Refinement Module (MRM) dynamically extracts diffusion cross-attention maps during FLUX denoising to produce calibrated masks \(\mathbf{m}_i^{rfn}\). Finally, a correspondence mask modulates the image-text attention layers to isolate cross-instance interactions and prevent attribute bleeding.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Text Prompt p + Visual Condition c"] --> B["Stage 1: Text-Visual Condition Association<br/>VLM parses phrases and predicts instance masks"]
B --> C["Shared SEG Token Aggregation<br/>Consolidates representations across mentions"]
C --> D["Stage 2: Instance-Aware Controllable Generation<br/>FLUX backbone denoising & feature extraction"]
D --> E["Mask Refinement Module<br/>Fuses confidence, attention maps & latent features"]
E --> F["Correspondence Mask Construction<br/>Constrains image-text cross-attention"]
F --> G["High-Fidelity Multi-Instance Output"]
Key Designs¶
1. Instance-Level Text-Visual Association: Cross-Modal Phrase Grounding with Shared SEG Tokens
In complex multi-instance prompts, descriptions of entities are often interleaved with prepositions, verbs, and relational clauses, and the same entity is frequently mentioned at multiple points in a paragraph. Standard referring segmentation models designed for RGB imagery struggle when applied to sparse, textureless condition maps like edges or depth maps. InstanceControl employs the Sa2VA framework and LoRA fine-tuning to output a structured prediction sequence \(\mathbf{y}_{txt}^{pred}\) (e.g., <p>a man in a black jacket</p><seg><1>). When multiple <SEG> tokens correspond to the same instance ID \(i\), a Shared SEG Token (SST) strategy groups and concatenates their final-layer MLP-projected query representations:
$\(\mathbf{h}_i = \text{Concat}(\mathbf{h}_i^1, \mathbf{h}_i^2, \dots, \mathbf{h}_i^M)\)$
The consolidated query vector \(\mathbf{h}_i\) and dense visual condition features \(\mathbf{f}\) extracted by the SAM image encoder are fed into the SAM decoder to predict dense instance masks \(\mathbf{m}_i^{pred}\) alongside scalar confidence scores \(s_i^{pred} \in [0, 1]\), ensuring consistent grounding across long, multi-clause prompts.
2. Mask Refinement Module: Confidence-Guided Dynamic Spatial Rectification
Initial masks \(\mathbf{m}_i^{pred}\) generated by the VLM inevitably exhibit occasional noise, under-segmentation, or localization offsets. Enforcing these imperfect masks as rigid hard constraints can restrict generation and introduce visible spatial artifacts. Observing that the diffusion model's internal cross-attention maps \(\mathbf{m}_i^{att}\) inherently reflect its generative spatial intent and complement the VLM masks, the authors introduce a lightweight U-Net Mask Refinement Module (MRM). At each denoising timestep \(t\), MRM receives the predicted mask \(\mathbf{m}_i^{pred}\), attention-based mask \(\mathbf{m}_i^{att}\), confidence score \(s_i^{pred}\), and image latent features \(\mathbf{H}_{img}\): $\(\mathbf{m}_i^{rfn} = f(\mathbf{m}_i^{pred}, \mathbf{m}_i^{att}, s_i^{pred}, \mathbf{H}_{img})\)$ Equipped with intra- and inter-instance attention blocks, MRM trusts \(\mathbf{m}_i^{pred}\) when confidence \(s_i^{pred}\) is high, but adaptively shifts weight toward the generative attention cues \(\mathbf{m}_i^{att}\) when confidence is low, achieving smooth, noise-resilient spatial guidance.
3. Correspondence Mask Construction: Soft-Gated Cross-Attention Modulation
To seamlessly inject the refined correspondences \(\{(\mathbf{t}_i^{pred}, \mathbf{m}_i^{rfn})\}_{i=1}^N\) into the Diffusion Transformer (FLUX) without cross-instance attribute interference, a correspondence mask \(\mathbf{M} \in \mathbb{R}^{L_{img} \times L_{txt}}\) is constructed. Letting \(\mathcal{T}_i\) denote the set of text tokens for instance \(i\), \(\mathcal{I}_i\) denote image tokens belonging to \(\mathbf{m}_i^{rfn}\), and \(\mathcal{I}_{bg}\) denote background tokens: $\(\mathbf{M}[q, k] = \begin{cases} \mathbf{m}_i^{rfn}, & \text{if } q \in \mathcal{I}_i \text{ and } k \in \mathcal{T}_i \\ 1, & \text{if } q \in \mathcal{I}_{bg} \\ 0, & \text{otherwise} \end{cases}\)$ The mask enters the attention computation in logarithmic form: $\(\text{Attention}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \text{Softmax}\left(\frac{\mathbf{Q}\mathbf{K}^T}{\sqrt{d}} + \log(\mathbf{M})\right)\mathbf{V}\)$ When \(\mathbf{M}[q, k] \to 1\), \(\log(\mathbf{M}) \to 0\), recovering standard uninhibited cross-attention; when \(\mathbf{M}[q, k] \to 0\), \(\log(\mathbf{M}) \to -\infty\), completely suppressing attention weights from irrelevant entities. This hard-isolated yet soft-weighted mechanism eliminates cross-entity attribute contamination.
Loss & Training¶
The framework is optimized in two disjoint phases: 1. Stage 1 (VLM Grounding Optimization): LoRA adapters (rank 256) are inserted into Sa2VA's VLM, image encoder, and SAM encoder/decoder. The overall loss combines autoregressive text cross-entropy with mask segmentation losses: $\(\mathcal{L}_{Stage1} = \mathcal{L}_{txt} + \sum_{i=0}^N \left(\lambda_{bce} \text{BCE}(\mathbf{m}_i^{pred}, \mathbf{m}_i^{gt}) + \lambda_{dice} \text{DICE}(\mathbf{m}_i^{pred}, \mathbf{m}_i^{gt})\right)\)$ with weights \(\lambda_{bce}=2.0, \lambda_{dice}=0.5\), trained for 30k steps with batch size 64 using AdamW (initial learning rate \(4 \times 10^{-5}\)). 2. Stage 2 (Generation and Refinement Optimization): Built on FLUX.1 (Canny, Depth, HED ControlNet). First, LoRA is applied across all linear layers of FLUX DiT blocks on ground-truth masks for 80k steps (batch size 4, cosine schedule starting at \(1 \times 10^{-4}\)). Subsequently, MRM is trained on predicted masks for 10k steps under flow-matching loss \(\mathcal{L}_{flux}\) paired with mask loss \(\mathcal{L}_{mask}(\mathbf{m}_i^{rfn}, \mathbf{m}_i^{gt})\).
Key Experimental Results¶
Main Results¶
On the multi-instance benchmark MIG-Eval (comprising 5,400 complex test images averaging 11.64 instances and 183.15 prompt tokens), InstanceControl is evaluated against both instance-labeled methods (requiring manual box/mask inputs) and label-free FLUX ControlNet. Evaluation encompasses spatial segmentation alignment (MIoU), region-wise attribute fidelity (Local CLIP, Qwen2-VL-72B VQA accuracy for spatial, color, shape, and texture), and global visual quality (ImageReward, FID).
| Methods | Instance Labeling | Condition | MIoU ↑ | Local CLIP ↑ | Spatial ↑ | Color ↑ | Shape ↑ | Texture ↑ | FID ↓ |
|---|---|---|---|---|---|---|---|---|---|
| EliGen | Required (Box) | None | 0.6104 | 17.49 | 91.97% | 84.98% | 85.59% | 87.93% | 27.82 |
| CreatiLayout | Required (Box) | None | 0.5247 | 16.84 | 88.16% | 80.87% | 83.00% | 85.91% | 21.40 |
| Seg2Any | Required (Mask) | None | 0.8316 | 18.62 | 91.71% | 87.25% | 88.39% | 90.14% | 21.23 |
| DreamRenderer | Required (Manual) | Canny | 0.6497 | 16.25 | 80.74% | 67.55% | 69.20% | 74.68% | 14.51 |
| DreamRenderer | Required (Manual) | Depth | 0.7738 | 17.88 | 89.52% | 80.91% | 81.45% | 86.69% | 15.36 |
| DreamRenderer | Required (Manual) | HED | 0.7060 | 15.08 | 78.93% | 60.56% | 63.18% | 68.79% | 22.67 |
| FLUX ControlNet | Label-Free | Canny | 0.6526 | 16.48 | 84.67% | 73.30% | 73.93% | 79.24% | 14.04 |
| InstanceControl (Ours) | Label-Free | Canny | 0.8250 | 18.51 | 93.54% | 87.78% | 88.19% | 90.88% | 10.03 |
| FLUX ControlNet | Label-Free | Depth | 0.7782 | 17.73 | 90.14% | 75.87% | 78.20% | 82.33% | 14.11 |
| InstanceControl (Ours) | Label-Free | Depth | 0.8116 | 18.18 | 92.87% | 86.11% | 86.49% | 89.91% | 11.92 |
| FLUX ControlNet | Label-Free | HED | 0.6817 | 15.67 | 83.09% | 70.15% | 70.51% | 77.08% | 20.08 |
| InstanceControl (Ours) | Label-Free | HED | 0.8472 | 18.74 | 94.64% | 88.75% | 89.55% | 91.18% | 10.14 |
Ablation Study¶
An ablation study on MIG-Eval (Canny condition) examines the impact of different mask guidance variants (raw predicted mask \(\mathbf{m}^{pred}\), heuristic score-weighted linear fusion \(\mathbf{m}^{fuse}\), and the full MRM \(\mathbf{m}^{rfn}\)) along with the Shared SEG Token (SST) mechanism.
| Config | Mask Formulation | SST | MIoU ↑ | Local CLIP ↑ | Color ↑ | Texture ↑ | FID ↓ | Note |
|---|---|---|---|---|---|---|---|---|
| Config A | Raw predicted \(\mathbf{m}^{pred}\) | Yes | 0.8155 | 18.07 | 84.88% | 89.10% | 10.93 | Unrefined mask introduces segmentation noise |
| Config B | Linear fusion \(\mathbf{m}^{fuse}\) | Yes | 0.8210 | 18.39 | 85.91% | 89.99% | 10.06 | Naive score-weighted interpolation |
| Config C (Full MRM) | Refined mask \(\mathbf{m}^{rfn}\) | Yes | 0.8250 | 18.51 | 87.78% | 90.88% | 10.03 | Dynamic fusion with attention and latents |
| Config D (w/o SST) | Refined mask \(\mathbf{m}^{rfn}\) | No | 0.8214 | 18.31 | 85.89% | 89.83% | 10.30 | Unconsolidated multi-mention tokens degrade consistency |
Key Findings¶
- Without requiring manual instance labeling, InstanceControl significantly outperforms the standard FLUX ControlNet baseline across all metrics under Canny conditions, improving Color accuracy by 14.48%, Texture accuracy by 11.64%, and reducing FID from 14.04 to 10.03.
- InstanceControl matches or surpasses methods relying on human-annotated bounding boxes or masks (such as CreatiLayout, EliGen, and DreamRenderer) across nearly all regional fidelity metrics.
- The Mask Refinement Module (MRM) is critical for overcoming segmentation noise: directly utilizing \(\mathbf{m}^{pred}\) leads to a drop of nearly 3% in color accuracy, whereas MRM dynamically leverages diffusion cross-attention to stabilize boundaries.
- The Shared SEG Token (SST) prevents degradation in long prompts where entities appear across multiple clauses, lifting Local CLIP from 18.31 to 18.51.
Highlights & Insights¶
- Perception-Guided Generation without Annotation: Leverages the open-vocabulary referring segmentation ability of a VLM as an automated parsing frontend, liberating multi-instance controllable generation from cumbersome manual bounding box and mask labeling.
- Bi-Directional Dynamic Calibration via MRM: Rather than treating VLM predictions as rigid ground truth, the Mask Refinement Module reconciles prediction confidence with internal diffusion cross-attention maps, self-correcting perception flaws during the generation process.
- Interactive In-the-Loop Refinement: Offers an optional user-in-the-loop mechanism where users can provide simple point or box prompts to override challenging failure cases, boosting average VQA accuracy from 90.10% to 93.19%.
Limitations & Future Work¶
- Perception Bottleneck in Heavily Occluded Scenes: In extremely dense scenes with overlapping boundaries and ambiguous depth cues, the Stage 1 VLM can miss objects or predict offset bounding regions; when an entity is missed entirely in Stage 1, Stage 2 cannot isolate it.
- Cascaded Inference Latency and VRAM Overhead: Running a full VLM + SAM grounding pass followed by U-Net mask rectification and Diffusion Transformer denoising increases memory consumption and runtime latency compared to a single monolithic model. A unified, end-to-end multimodal architecture represents a promising future research direction.
Related Work & Insights¶
- vs FLUX ControlNet: FLUX ControlNet applies global condition conditioning without mapping specific phrases to spatial locations, causing severe attribute leakage in multi-entity scenes. InstanceControl resolves this by building explicit instance-level bindings and attention isolation.
- vs DreamRenderer / Seg2Any: DreamRenderer and Seg2Any rely on user-annotated spatial boxes or pixel masks during inference. InstanceControl achieves comparable or superior fidelity fully automatically.
- vs Layout-to-Image Models (EliGen / CreatiLayout): Layout-based approaches use coarse bounding boxes without geometric contours, frequently producing distorted shapes in overlapping regions. InstanceControl pairs detailed structural conditions (Canny/Depth) with fine-grained semantic grounding.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Pioneering approach that bridges VLM referring segmentation with diffusion cross-attention modulation to remove the labeling barrier in controllable multi-instance generation.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluation across Canny, Depth, and HED conditions on both MIG-Eval and COCO-POS benchmarks with extensive VQA-based attribute metrics and ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Well-structured, focused motivation, clear formulations, and high-quality qualitative illustrations.
- Value: ⭐⭐⭐⭐⭐ Substantially advances the usability of controllable generative pipelines in complex, real-world multi-entity applications.