SCALE: Semantic-Calibrated Guidance Enhancement for Prompt-Faithful Diffusion¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/SudongCAI/SCALE
Area: Image Generation
Keywords: Diffusion Model, Text-to-Image Generation, Semantic Alignment, Classifier-Free Guidance, Decoupled Control
TL;DR¶
Addressing the quality–alignment tension caused by coupled linear extrapolation in Classifier-Free Guidance (CFG), SCALE introduces a training-free, orthogonal-invariant enhancement mechanism that selectively scales semantic progress along the prompt axis while strictly preserving orthogonal corrective components, significantly improving prompt adherence with negligible inference latency.
Background & Motivation¶
Text-to-image diffusion models have achieved unprecedented visual fidelity, yet photorealism does not guarantee prompt faithfulness. In complex text-guided visual synthesis, generative models frequently suffer from concept bleeding, attribute mismatch, numerical miscounting, and spatial relationship inversions. Classifier-Free Guidance (CFG) serves as the ubiquitous cornerstone for conditional steering by performing linear extrapolation between unconditional and conditional score predictions. However, this scalar extrapolation intrinsically induces a severe quality–alignment tension: aggressively scaling up the guidance factor \(\omega\) amplifies both the semantic steering signal and structural perturbations, driving the denoising trajectory off the underlying data manifold and causing structural distortions, oversaturation, and noise-like artifacts.
Prior efforts to navigate this tension either suppress artifacts under high guidance scales (such as APG) or operate within conservative guidance ranges while reshaping the sampling path via computationally expensive test-time interventions (such as Z-Sampling or DAS). To rigorously inspect the theoretical limits of greedy alignment maximization, this work introduces Semantic Alignment Projection (SAP) as a boundary probe, constraining sampling updates strictly to the 1D semantic axis. Counterintuitively, SAP causes catastrophic visual collapse into noise. Theoretical analysis reveals that stable denoising critically relies on high-dimensional orthogonal degrees of freedom to rectify accumulated trajectory drift; eliminating the orthogonal complement strips the diffusion dynamics of its corrective capability.
This diagnosis exposes a fundamental architectural incompatibility between semantic progress and manifold stability. Core idea: establish the "Decoupled Satisfaction" paradigm through SCALE, applying a minimum-perturbation amplification strictly to the semantic projection parallel to the text guidance vector while enforcing invariant preservation of the orthogonal corrective channel, decoupling prompt adherence from structural stability.
Method¶
Overall Architecture¶
SCALE is a lightweight, training-free, plug-and-play step-level guidance enhancement mechanism. It operates directly on the sampler update vectors during the reverse diffusion trajectory without requiring model retraining, fine-tuning, or internal attention map modifications.
At each reverse timestep \(t\), given the current latent state \(x_t\) and prompt condition \(c\), the model computes the conditional noise prediction \(\epsilon_\theta(x_t, c)\) and unconditional prediction \(\epsilon_\theta(x_t, \emptyset)\). These form the conditional guidance vector \(g_t = \epsilon_\theta(x_t, c) - \epsilon_\theta(x_t, \emptyset)\) and its sampler update-induced reference direction \(\hat{g}_t = -g_t\). For the baseline sampler's tentative trial update \(\tilde{d}_t = \tilde{x}_{t-1} - x_t\), SCALE performs an orthogonal decomposition into a semantic projection component \(d_t^\parallel\) parallel to \(\hat{g}_t\) and a corrective component \(d_t^\perp\) residing in the orthogonal complement. An adaptive hybrid gain scheduler then computes the cosine alignment degree between the update and the semantic axis, dynamically determining a semantic scaling factor \(\lambda_t\). By scaling \(d_t^\parallel\) with \(\lambda_t\) while leaving \(d_t^\perp\) completely undistorted, SCALE outputs the calibrated update \(z_t\) and stably steps to the next state \(\hat{x}_{t-1}\).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input prompt c & latent state x_t"] --> B["Orthogonal Projection Decomposition<br/>Decompose into semantic d∥ & orthogonal d⊥"]
B --> C["Boundary Extremum Probe<br/>Diagnose SAP failure from lost orthogonal freedom"]
C --> D["Orthogonal-Invariant Semantic Enhancement<br/>Preserve d⊥ and apply minimal-perturbation scaling"]
D --> E["Adaptive Hybrid Gain Scheduling<br/>Gate scaling factor λt via cosine alignment"]
E --> F["Calibrated Update Output z_t<br/>Advance reverse step to x_t-1"]
Key Designs¶
1. Orthogonal Projection Decomposition: Disentangling Semantic Steering from Structural Rectification
Standard CFG directly adds an extrapolated noise delta to the unconditional prediction, conflating class-posterior semantic driving forces with structural prior restoration in a single coupled update vector. This design explicitly establishes a geometric reference frame in update space. Formally, define the guidance vector \(g_t \triangleq \epsilon_\theta(x_t, c) - \epsilon_\theta(x_t, \emptyset)\), and the update-aligned direction \(\hat{g}_t \triangleq -g_t\). For any baseline update \(d_t\), the orthogonal projector \(\mathcal{P}_{\hat{g}_t}(\cdot)\) projects \(d_t\) onto \(\text{span}(\hat{g}_t)\): $\(d_t^\parallel = \mathcal{P}_{\hat{g}_t}(d_t) = \frac{\langle d_t, \hat{g}_t \rangle}{\|\hat{g}_t\|^2} \hat{g}_t, \quad d_t^\perp = d_t - d_t^\parallel\)$ This decomposition establishes a clear functional separation: the one-dimensional component \(d_t^\parallel\) carries the prompt-conditional likelihood gradient driving text adherence, while the ambient orthogonal complement \(d_t^\perp\) acts as a restorative force that absorbs discretization errors and pulls the trajectory toward the high-density data manifold.
2. Boundary Extremum Probe: Diagnosing Orthogonal Correction Collapse in SAP
To explore the theoretical limit of alignment progress, the authors formulate Semantic Alignment Projection (SAP) as a steepest-ascent greedy probe. Under a fixed step budget, maximizing the instantaneous semantic gain proxy \(\mathcal{J}_t(z_t) \triangleq \langle z_t, \hat{g}_t \rangle\) saturates Cauchy–Schwarz equality only when \(z_t\) is strictly collinear with \(\hat{g}_t\). Consequently, SAP imposes \(z_t = \mathcal{P}_{\hat{g}_t}(d_t)\) and discards \(d_t^\perp\) entirely. By Lemma 1, while orthogonal projection provides the unique minimum-perturbation approximation within the subspace \(\text{span}(\hat{g}_t)\), it systematically eliminates \(d_t^\perp\). Under the manifold hypothesis, natural images concentrate near a low-dimensional manifold with an anisotropic prior potential: normal curvatures deviating from the manifold are substantially steeper than tangential curvatures. Removing \(d_t^\perp\) completely suppresses the restoring forces along the normal directions, causing trajectory drift to compound across timesteps and triggering catastrophic generative collapse into noise. This failure rigorously underscores the necessity of orthogonal corrective degrees of freedom.
3. Orthogonal-Invariant Semantic Enhancement: Decoupled Satisfaction and Minimal-Perturbation Optimality
Motivated by the SAP diagnosis, SCALE formalizes the Decoupled Satisfaction paradigm: selectively amplify the semantic axis without altering the orthogonal channel. This is framed as a constrained convex optimization problem seeking the minimal Euclidean perturbation to the baseline update \(d_t\) under a target semantic gain lower bound \(\langle z, \hat{g}_t \rangle \ge \lambda \langle d_t, \hat{g}_t \rangle\) with \(\lambda \ge 1\): $\(\min_{z \in \mathbb{R}^n} \|z - d_t\|^2 \quad \text{s.t.} \quad \langle z, \hat{g}_t \rangle \ge \lambda \langle d_t, \hat{g}_t \rangle\)$ Theorem 2 proves that this convex program admits the unique closed-form minimizer: $\(z^* = \lambda d_t^\parallel + d_t^\perp\)$ From an operator perspective, SCALE applies an axis-aligned linear operator \(\mathcal{A}_\lambda \triangleq \mathcal{P}_{\hat{g}_t}^\perp + \lambda \mathcal{P}_{\hat{g}_t} = I + (\lambda - 1)\mathcal{P}_{\hat{g}_t}\) to the update vector. The operator exhibits eigenvalue \(\lambda\) on \(\text{span}(\hat{g}_t)\) and eigenvalue \(1\) on \(\text{span}(\hat{g}_t)^\perp\). By Corollary 3, \(\mathcal{A}_\lambda\) acts as an exact identity mapping on the orthogonal complement, guaranteeing zero distortion on the orthogonal channel and decoupling semantic enhancement from structural stability.
4. Adaptive Hybrid Gain Scheduling: Cosine-Based Directional Gating and Dynamic Scaling
In practical diffusion sampling, the semantic alignment of baseline updates varies dynamically across timesteps. Aggressively boosting updates during initial coarse-layout stages or when updates deviate from the prompt direction (\(\langle \tilde{d}_t, \hat{g}_t \rangle \le 0\)) risks generating severe visual artifacts. SCALE introduces an adaptive hybrid gain schedule. It first computes the cosine alignment score in the update space: $\(\cos(\Theta_t) \triangleq \frac{\langle \tilde{d}_t, \hat{g}_t \rangle}{\|\tilde{d}_t\| \|\hat{g}_t\|}\)$ The semantic scaling factor is dynamically modulated as: $\(\lambda_t \triangleq \lambda_{\text{base}} + \cos(\Theta_t)\)$ Combined with a step-gating schedule, SCALE is activated only during mid-to-late denoising steps when \(\cos(\Theta_t) > 0\), bounding \(\lambda_t \in [\lambda_{\text{base}}, \lambda_{\text{base}} + 1]\). When \(\cos(\Theta_t) \le 0\) or during inactive steps, \(\lambda_t\) defaults to \(1\), gracefully degenerating to the unperturbed baseline update. This gating ensures that semantic amplification is applied assertively when alignment direction is confident, and bypassed when uncertain.
A Worked Example¶
Consider a generation pass with prompt "A red car and a white sheep" at reverse step \(t=15\):
1. The denoiser predicts \(\epsilon_\theta(x_{15}, c)\) and \(\epsilon_\theta(x_{15}, \emptyset)\), forming the guidance vector \(g_{15}\) and reference direction \(\hat{g}_{15}\).
2. The sampler computes trial update \(\tilde{d}_{15}\), yielding cosine alignment \(\cos(\Theta_{15}) = 0.65 > 0\).
3. With active gating and \(\lambda_{\text{base}} = 1.0\), the adaptive scaling factor evaluates to \(\lambda_{15} = 1.0 + 0.65 = 1.65\).
4. \(\tilde{d}_{15}\) is projected into parallel component \(d_{15}^\parallel\) and orthogonal remainder \(d_{15}^\perp\).
5. The calibrated update is formed as \(z_{15} = 1.65 d_{15}^\parallel + d_{15}^\perp\), cleanly binding "red" to the car body and "white" to the sheep without distorting background structures or ground texture.
Key Experimental Results¶
Main Results¶
SCALE is comprehensively benchmarked on SDXL (\(1024 \times 1024\) resolution) across DrawBench, GenEval, and T2I-CompBench using five evaluation metrics: CLIPScore, BLIPScore, VQAScore, DSGScore, and ImageReward.
| Dataset | Method | CLIPScore ↑ | BLIPScore ↑ | VQAScore ↑ | DSGScore ↑ | ImageReward ↑ |
|---|---|---|---|---|---|---|
| DrawBench | SDXL (Base) | 0.2775 | 0.5087 | 0.7417 | 0.7068 | 0.6402 |
| PAG (ECCV'24) | 0.2696 | 0.4896 | 0.7239 | 0.6727 | 0.5509 | |
| SEG (NeurIPS'24) | 0.2723 | 0.4998 | 0.7155 | 0.6662 | 0.6255 | |
| AYS (ICML'24) | 0.2703 | 0.4983 | 0.7239 | 0.7003 | 0.3235 | |
| Z-Sampling (ICLR'25) | 0.2811 | 0.5129 | 0.7438 | 0.7064 | 0.7983 | |
| CFG++ (ICLR'25) | 0.2804 | 0.5067 | 0.7487 | 0.7133 | 0.6022 | |
| APG (ICLR'25) | 0.2863 | 0.5176 | 0.7800 | 0.6912 | 0.7857 | |
| Golden Noise (ICCV'25) | 0.2814 | 0.5043 | 0.7439 | 0.7084 | 0.6694 | |
| GA-Eval (ICLR'26) | 0.2889 | 0.5198 | 0.7867 | 0.7034 | 0.8522 | |
| TAG (ICML'26) | 0.2717 | 0.4957 | 0.7259 | 0.6697 | 0.4028 | |
| SCALE (Ours) | 0.2885 | 0.5244 | 0.7889 | 0.7305 | 0.8305 | |
| GenEval | SDXL (Base) | 0.2802 | 0.5120 | 0.7998 | 0.8439 | 0.5746 |
| Z-Sampling (ICLR'25) | 0.2860 | 0.5180 | 0.8165 | 0.8455 | 0.7624 | |
| CFG++ (ICLR'25) | 0.2814 | 0.5128 | 0.8170 | 0.8331 | 0.6537 | |
| APG (ICLR'25) | 0.2873 | 0.5173 | 0.8176 | 0.8295 | 0.7599 | |
| SCALE (Ours) | 0.2887 | 0.5190 | 0.8345 | 0.8556 | 0.8168 | |
| T2I-CompBench | SDXL (Base) | 0.2771 | 0.5194 | 0.7192 | 0.7306 | 0.4852 |
| Z-Sampling (ICLR'25) | 0.2811 | 0.5261 | 0.7347 | 0.7490 | 0.7125 | |
| APG (ICLR'25) | 0.2815 | 0.5282 | 0.7382 | 0.7462 | 0.7424 | |
| SCALE (Ours) | 0.2823 | 0.5282 | 0.7593 | 0.7593 | 0.7725 |
In the fine-grained GenEval benchmark, SCALE demonstrates marked gains in challenging multi-object compositional skills:
| Method | Single object | Two object ↑ | Counting ↑ | Colors ↑ | Position ↑ | Color attribution ↑ | Overall ↑ |
|---|---|---|---|---|---|---|---|
| SDXL | 100.00% | 68.69% | 40.00% | 81.91% | 9.00% | 18.00% | 53.44% |
| Z-Sampling | 100.00% | 76.77% | 38.75% | 86.17% | 12.00% | 20.00% | 55.61% |
| SCALE (Ours) | 100.00% | 71.72% | 42.50% | 88.30% | 17.00% | 19.00% | 56.41% |
Ablation Study¶
To isolate the functional roles of parallel and orthogonal components w.r.t. \(\hat{g}_t\), an ablation study on DrawBench compares component pruning and scaling strategies:
| Config / Variant | CLIPScore ↑ | BLIPScore ↑ | VQAScore ↑ | DSGScore ↑ | ImageReward ↑ | Note |
|---|---|---|---|---|---|---|
| SDXL (Baseline) | 0.2771 | 0.5188 | 0.7184 | 0.6945 | 0.5921 | Standard CFG sampling |
| Simultaneous Enhancement | 0.2624 | 0.4784 | 0.6519 | 0.5881 | -0.1666 | Scaled both \(d^\parallel\) and \(d^\perp\) equally (manifold departure) |
| Parallel-Only (SAP) | 0.1418 | 0.2950 | 0.1446 | 0.0092 | -2.2422 | Discarded \(d^\perp\) completely (collapse of orthogonal freedom) |
| Perp-Only | 0.1202 | 0.2864 | 0.2900 | 0.1097 | -2.0469 | Discarded \(d^\parallel\) completely (loss of prompt steering) |
| SCALE (Full Model) | 0.2885 | 0.5244 | 0.7889 | 0.7305 | 0.8305 | Amplified \(d^\parallel\) while preserving \(d^\perp\) exactly |
In addition, inference latency benchmarking highlights SCALE's efficiency: standard SDXL sampling takes 5.3486s per image, while SCALE takes 5.3687s (relative latency ratio of 1.0038×). In contrast, trajectory-optimization baselines incur substantial latency penalties (Z-Sampling at 2.82×, DAS at 55.60×). When applied on top of Diffusion-DPO preference tuning, SCALE boosts ImageReward from 0.9856 to 1.0278 on GenEval, confirming strong orthogonality and complementarity to post-training alignment.
Key Findings¶
- Crucial Role of Orthogonal Correction: Parallel-Only (SAP) triggers a complete collapse across all alignment metrics (DSGScore drops precipitously from 0.6945 to 0.0092, ImageReward plunges to -2.2422), confirming that eliminating orthogonal corrective degrees of freedom critically breaks denoising trajectory stability.
- Superiority of Decoupled Scaling: Uniformly scaling both \(d_t^\parallel\) and \(d_t^\perp\) (Simultaneous Enhancement) drives ImageReward negative (-0.1666), proving that unconstrained scalar extrapolation amplifies structural noise and confirming that selective semantic amplification is the correct path.
- Enhanced Compositional Control: On GenEval, SCALE raises object counting accuracy to 42.50% (+2.50% over SDXL) and nearly doubles spatial position accuracy to 17.00% (from 9.00%), demonstrating robust attribute binding in complex scenes.
Highlights & Insights¶
- Theoretical Insight from Boundary Probing: Rather than proposing heuristics ad-hoc, the authors first probe alignment limits with SAP, rigorously tracing its catastrophic failure to the loss of manifold-rectifying orthogonal degrees of freedom via differential geometry and Bayesian score decomposition.
- Elegant Closed-Form Minimal-Perturbation Operator: Formalizing the decoupled requirement as a constrained convex program yields the axis-aligned operator \(\mathcal{A}_\lambda = I + (\lambda - 1)\mathcal{P}_{\hat{g}_t}\), which strictly acts as an identity mapping on the orthogonal complement while provably delivering minimal Euclidean perturbation.
- Negligible Overhead with Universal Compatibility: SCALE requires no gradient backpropagation, attention patching, or iterative trajectory loops. Executed solely via basic vector dot products at inference time, it operates with a 1.0038× latency ratio and orthogonally boosts preference-tuned models like Diffusion-DPO.
Limitations & Future Work¶
- Dependency on Base Model Semantic Priors: SCALE amplifies guidance directions predicted by the base model. If the pre-trained denoiser completely lacks knowledge of rare concepts, resulting in degenerate guidance vectors, geometric amplification cannot synthesize unknown concepts ex nihilo.
- Step Gating Hyperparameters: The onset of semantic amplification relies on empirical step gating. Identifying optimal activation schedules across diverse sampler solvers (e.g., DDIM, DPM-Solver++, Euler a) currently requires heuristic tuning.
- Future Work: Extending orthogonal-invariant guidance operators to spatio-temporal video diffusion dynamics and multi-condition generation scenarios (e.g., joint ControlNet structural conditioning and text guidance).
Related Work & Insights¶
- vs CFG [Ho & Salimans, 2021]: CFG applies uniform scalar extrapolation that entangles semantic driving forces with structural noise; SCALE decouples update space via orthogonal projection, achieving superior alignment without over-guidance artifacts.
- vs APG [Sadat et al., ICLR 2025]: APG suppresses artifacts under extreme guidance scales using momentum and norm clipping; SCALE directly operates on the geometric structure of updates, inherently avoiding artifacts via minimal-perturbation orthogonal invariance.
- vs Z-Sampling [LiChen et al., ICLR 2025] & DAS [Kim et al., ICLR 2025]: Z-Sampling and DAS rely on reflection loops or reward gradients that incur massive latency overheads (2.8× to 55.6×); SCALE provides closed-form, gradient-free updates at 1.0038× speed.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates the quality–alignment tension through geometric projection and manifold anisotropy, proposing an elegant orthogonal-invariant operator.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorously evaluated across three major benchmarks, fine-grained compositional categories, diverse metrics, and comprehensive ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Mathematically rigorous, clearly motivated, and structured with clean geometric visualizations.
- Value: ⭐⭐⭐⭐⭐ A training-free, practically zero-overhead plug-and-play technique with immediate utility across existing diffusion pipelines.