Skip to content

The Path to Reconciling Quality and Safety Alignment in Text-to-Image Generation

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Image Generation
Keywords: text-to-image generation, safety alignment, preference optimization, diffusion models, synergistic alignment

TL;DR

Addressing the debilitating trade-off in text-to-image safety alignment where safety constraints degrade generation aesthetics and instruction fidelity, this paper proposes a unified framework comprising the dual-annotated dataset LibraAlign-100K, the Synergistic Preference Optimization algorithm (T2I-SPO), and the balanced Unified Alignment Score (UAScore), effectively eliminating diverse NSFW concepts while fully preserving generative quality.

Background & Motivation

The swift advances of diffusion models in text-to-image (T2I) generation enable the synthesis of photorealistic images from open-ended natural language prompts, yet simultaneously heighten severe ethical and societal risks regarding malicious misuse, such as generating explicit, violent, or hate-promoting content. Existing safety mechanisms primarily follow three paradigms: post-hoc safety checkers are inherently reactive and brittle against adversarial attacks; parameter-editing concept erasure techniques frequently suffer from catastrophic concept forgetting and semantic confusion (e.g., erasing the concept of "nudity" unintentionally degrades normal representations of skin textures or muscle definition); and reinforcement learning or preference fine-tuning methods rely on myopic reward signals focused solely on safety penalties.

This myopic single-objective formulation forces a debilitating zero-sum trade-off between safety and visual appeal. The authors term this the "Sleeping Venus Dilemma" — when tasked with generating classical art or stylized depictions involving sensitive keywords, an overly conservative aligned model either rejects the prompt outright or collapses into a low-fidelity, instruction-violating output stripped of artistic value. Crucially, this friction is not fundamental to diffusion models, but stems from systemic deficiencies across three dimensions: flawed data construction relying on rigid prompt editing that causes severe semantic shift; single-objective optimization treating safety and quality as opposing forces; and quality-agnostic binary classification benchmarks incapable of measuring continuous safety cost or aesthetic degradation.

To overcome this false dichotomy, the paper's core angle is shifting safety alignment from punitive single-objective filtering to holistic, dual-objective synergistic preference modeling. Core idea: develop a benchmark data pipeline providing dual supervision of continuous safety cost and aesthetic quality, derive a synergistic preference optimization algorithm (T2I-SPO) that couples quality reward with safety cost in a unified policy objective, and actively boost hard-to-learn examples via a Dynamic Focusing Mechanism.

Method

Overall Architecture

The proposed framework reconciles safety and quality across three interconnected pillars: data curation (LibraAlign-100K with dual annotations), algorithmic optimization (T2I-SPO with a dynamic focusing mechanism), and holistic evaluation (UAScore). First, a continuous Safety Cost Model (SCM) is trained on 90K coarse-grained paired comparisons across 7 NSFW categories and 634 concepts. Next, high-fidelity image pairs are synthesized via safety-aware inpainting to preserve background composition and style, with each image annotated with both an aesthetic quality score and a safety cost. Finally, the diffusion denoising network is optimized using a composite preference objective that jointly weights quality gain against safety penalty, reinforced by active gradient-maximizing augmentations on lagging hard pairs.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Text Prompt & Diffusion Latents"] --> B["Safety-Aware Inpainting & Dual Annotation<br/>Preserve layout + label quality and safety cost"]
    B --> C["Continuous Safety Cost Modeling<br/>CLIP backbone + contrastive ranking & cost anchoring"]
    C --> D["Composite Preference Optimization T2I-SPO<br/>Jointly optimize quality reward and safety penalty"]
    D --> E["Dynamic Focusing Mechanism DFM<br/>Loss velocity tracking + gradient-maximizing augmentations"]
    E --> F["Aligned Diffusion Model Generating Safe & High-Fidelity Images"]

Key Designs

1. Safety-Aware Inpainting and Dual Annotation: Decoupling Safety from Quality Degradation Prior preference datasets rely heavily on crude text editing (e.g., changing "nude woman" to "woman"), which introduces drastic scene and composition changes, falsely training models to associate safety with minimal visual complexity. The authors introduce a safety-aware inpainting pipeline: SDXL first synthesizes a harmful image from an unsafe prompt, GPT-4 generates a semantically aligned safe prompt, and an image-to-image inpainting pass guided by this prompt replaces only sensitive regions while retaining original lighting, background, and artistic nuances. This produces 8,265 high-fidelity harmful-safe pairs. Combined with 2,495 safe-safe pairs from Pick-a-Pic to anchor benign generative capability, the resulting 10,760-pair LibraAlign-HF dataset is annotated with quality scores from a pretrained reward model and continuous safety costs from the SCM, providing unbiased dual-objective supervision.

2. Continuous Safety Cost Modeling: Fine-Grained Risk Quantification and Baseline Anchoring Standard binary classifiers fail to differentiate mild suggestiveness from severe illicit imagery, providing discontinuous gradient signals. The authors train an SCM \(C(\cdot)\) utilizing an Open-CLIP ViT-H/14 backbone with a lightweight adapter on 90K GPT-4o-annotated image pairs ranked with a 4-level severity rubric. To enforce relative order while signing positive costs for harmful images (\(S=+1\)) and negative costs for safe ones (\(S=-1\)), the model minimizes a contrastive ranking loss: $\(\mathcal{L}_{\text{CTRS}} = \mathbb{E}_{(I_w, I_l, S_w, S_l)} \left[ -\log\sigma(C(I_w) - C(I_l)) - \eta \cdot \big(\log\sigma(S_w \cdot C(I_w)) + \log\sigma(S_l \cdot C(I_l))\big) \right]\)$ To prevent ranking losses from producing excessive variance among benign samples — which would cause subsequent alignment to penalize innocuous visual patterns — a Cost Anchoring Loss regularizes safe samples: $\(\mathcal{L}_{\text{Anchor}} = \mathbb{E}_{(I_w, I_l), S_w=S_l=-1} \left[ (C(I_w) - \mu)^2 + (C(I_l) - \mu)^2 \right]\)$ pulling all benign costs toward the empirical mean \(\mu\) of verified safe images.

3. Composite Preference Optimization T2I-SPO: Integrating Quality Reward with Safety Constraints To prevent optimization from collapsing into overly conservative, aesthetically degraded regimes, T2I-SPO formulates a composite reward function \(R_\lambda(T, I) = R(T, I) - \lambda \cdot C(I)\), where \(R(T, I)\) evaluates prompt alignment and visual quality via PickScore, \(C(I)\) is the continuous safety cost, and \(\lambda > 0\) balances the trade-off. Under the Bradley-Terry preference model, sample pairs are dynamically re-indexed into winner \(I^+\) and loser \(I^-\). Adapting the Diffusion-DPO analytical derivation, policy alignment translates directly into an implicit denoising regression loss: $\(\mathcal{L}_{\text{T2I-SPO}}(\epsilon_\theta, \mathcal{D}_\lambda) = -\mathbb{E}_{(T, I^+, I^-) \sim \mathcal{D}_\lambda} \left[ \log\sigma \left( K \cdot \left[ \|\epsilon_t - \epsilon_\theta(I_t^+, T)\|^2 - \|\epsilon_t - \epsilon_{\text{ref}}(I_t^+, T)\|^2 - \left( \|\epsilon_t - \epsilon_\theta(I_t^-, T)\|^2 - \|\epsilon_t - \epsilon_{\text{ref}}(I_t^-, T)\|^2 \right) \right] \right) \right]\)$ This guarantees that policy updates concurrently optimize aesthetic fidelity while penalizing hazardous representations.

4. Dynamic Focusing Mechanism (DFM): Active Gradient Reinforcement on Hard Concepts Across diverse concepts, optimization progress varies widely; certain pairs suffer learning lag due to conflicting gradient updates. DFM maintains a sliding queue of the past \(m\) training steps' loss for each pair and computes its local descent rate \(v_i^{(t)}\). When a sample's rate is persistently below \(\eta\) times the batch average, it is marked as a hard sample. The system computes its current gradient \(g_k\) and evaluates an augmentation pool \(\mathcal{A}\) (image sharpening, color jittering, latent noise perturbation, and frequency band energy compensation) to select the transformation \(a^*\) maximizing gradient divergence: $\(a^* = \arg\max_{a \in \mathcal{A}} \|g_k - \nabla_\theta \mathcal{L}(\epsilon_\theta, T_k, a(I_k^+), a(I_k^-))\|\)$ Re-injecting the augmented pair into the current batch dynamically amplifies corrective signals for stubborn, boundary-blurring concepts without destabilizing overall convergence.

Key Experimental Results

Main Results

On standard benchmarks covering sexual content (I2P-Sexual, NSFW-56K-Sexual) and the complete I2P suite across 7 NSFW categories, T2I-SPO was evaluated against unaligned SD-v1.5, concept erasure baselines, and RL/DPO-based alignment methods across Inappropriate Probability (IP), Safety Cost (SC), PickScore (PS), and the Unified Alignment Score (UAScore).

Benchmark Metric SD-v1.5 (Baseline) AlignGuard (Prev. SOTA) T2I-SPO (Ours) Relative Gain / Reduction
I2P-Sexual IP (↓, %) 31.90 10.53 4.40 -27.50% (-86.2% relative)
SC (↓) -0.41 -5.08 -8.14 -7.73 lower risk
PickScore (↑) 19.33 18.82 19.38 +0.05 (exceeds base model)
UAScore (↑) 0.302 0.651 0.757 +0.455 (+150.7%)
NSFW-56K-Sexual IP (↓, %) 57.05 19.80 13.65 -43.40% (-76.1% relative)
SC (↓) 0.50 -7.08 -7.47 -7.97 lower risk
PickScore (↑) 19.77 18.86 19.14 +0.28 vs AlignGuard
UAScore (↑) 0.328 0.689 0.725 +0.397 (+121.0%)
Full I2P (7 Categories) IP (↓, %) 35.49 13.18 17.07 Best fine-grained SC
SC (↓) -4.21 -6.14 -7.73 -3.52 (best overall)
PickScore (↑) 19.56 18.94 19.61 +0.05 (superior aesthetics)
UAScore (↑) 0.696 0.690 0.784 +0.088 (+12.6%)

In fine-grained NudeNet evaluations on I2P-Sexual, T2I-SPO reduced the total count of exposed sensitive body parts to 43 (Breast: 3, Genitalia: 0, Buttocks: 1), markedly lower than unaligned SD-v1.5 (485) and AlignGuard (137), completely eliminating genitalia exposures.

Ablation Study

The ablations investigate the safety penalty weight \(\lambda\), the dynamic focusing sensitivity \(\eta\), and the composition of the augmentation pool.

Module / Hyperparameter Target Metric Value / Trend Insight
\(\lambda = 0.01\) I2P IP / COCO CLIPScore IP: 26.3% / CLIPScore: 26.54 Safety penalty too weak; insufficient defensive suppression
\(\lambda = 0.15\) (Default) I2P IP / COCO CLIPScore IP: 21.8% / CLIPScore: 26.46 Optimal Pareto frontier balancing safety and text alignment
\(\lambda = 0.50\) I2P IP / COCO CLIPScore IP: 20.7% / CLIPScore: 25.64 Safety overemphasized; collapses into standard DPO, degrading quality
DFM: Disabled (\(\eta = 0\)) Final Loss@Step 20K Plateaued around ~0.375 Hard concept pairs stall and compromise convergence
DFM: \(\eta = 0.2\) (Default) Final Loss@Step 20K Reaches optimal 0.359 Actively breaks stagnation via targeted gradient shifts
Augmentation Components: 1 Converged Loss 0.3737 Limited perturbation diversity constrains defensive robustness
Augmentation Components: 4 (All) Converged Loss 0.3593 Joint sharpening, jitter, noise, and energy compensation yields lowest loss

Key Findings

  • Zero Quality Penalty under Benign Prompts: When evaluated on purely safe prompts from Pick-a-Pic, HPD, and DrawBench, T2I-SPO maintained PickScore (20.4, 20.7, 21.0) and ImageReward (0.282, 0.233, 0.032) matching or outperforming the base unaligned SD-v1.5 (ImageReward: 0.161, 0.130, -0.014), proving that synergistic alignment prevents the aesthetic degradation seen in conventional methods.
  • Superior Adversarial Jailbreak Defense: Under adversarial jailbreak prompts from SneakyPrompt and MMA, T2I-SPO maintained the lowest inappropriate generation rates of 5.5% and 30.2%, respectively, substantially outperforming concept erasure baselines like ESD-u (9.0% / 33.8%) and UCE (12.0% / 36.8%).

Highlights & Insights

  • Compositional Inpainting for Unbiased Preference Pairs: Rather than relying on textual prompt substitution that causes massive global scene shifts, using image-to-image inpainting strictly localizes safety modifications while preserving background context, eliminating spurious correlation between safety and poor visual fidelity.
  • Closed-Form Direct Alignment with Composite Rewards: Mapping continuous safety costs alongside quality rewards into a unified Bradley-Terry likelihood avoids fragile adversarial policy training, enabling scalable full-parameter diffusion alignment.
  • Gradient-Maximizing Targeted Augmentation: Monitoring loss velocity in real time to trigger adaptive, gradient-divergent augmentations specifically for lagging concepts provides a principled paradigm for addressing hard-sample plateaus in generative modeling.

Limitations & Future Work

  • Author-Acknowledged Limitations: Full-parameter fine-tuning of diffusion backbones requires high GPU memory and compute budgets, limiting lightweight on-device deployment; nuanced concepts involving subtle metaphor or cultural ambiguity can still yield minor scoring ambiguities in the SCM.
  • Identified Limitations: The scalar weight \(\lambda\) remains static across all prompt types rather than adaptively modulating based on prompt risk level.
  • Future Directions: Extending T2I-SPO to parameter-efficient tuning regimes (such as LoRA modules) and expanding the dual-objective synergistic paradigm to text-to-video (T2V) diffusion models to safeguard temporal coherence and motion safety.
  • vs AlignGuard / SafetyDPO: While AlignGuard and SafetyDPO leverage preference optimization for safety, their training pairs suffer from large semantic divergence and their reward focuses purely on safety, degrading artistic details; T2I-SPO preserves scene context via inpainting and introduces dual-dimension rewards to resolve the quality trade-off.
  • vs Concept Erasure (ESD, UCE, RECE): Parameter-editing techniques modify cross-attention matrices via closed-form solutions, causing concept confusion and collateral damage to benign attributes like human textures; T2I-SPO guides diffusion trajectories through distribution alignment, preserving benign visual semantics.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates the "Sleeping Venus Dilemma" and resolves the long-standing safety-quality trade-off with a complete dataset-method-metric framework.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive coverage of 7 NSFW categories, SD-v1.5 and SDXL backbones, adversarial jailbreak benchmarks, and fine-grained anatomical part audits.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear progression, rigorous formulations, and insightful conceptual metaphors.
  • Value: ⭐⭐⭐⭐⭐ Establishes a highly practical and foundational paradigm for responsible deployment of frontier text-to-image foundation models.