BiSLW: Bi-Spectral Latent Watermarking for Generative Diffusion Models¶
Conference: ECCV 2026
Paper: CVF / ECCV 2026
Code: https://github.com/OVER-CODER/BiSLW
Area: Image Generation
Keywords: diffusion model watermarking, bi-spectral latent embedding, discrete cosine transform, cross-band consistency, regeneration attack defense
TL;DR¶
Addressing copyright attribution and regeneration-stripping vulnerabilities in generative diffusion models, BiSLW losslessly decomposes latent tensors into complementary low-frequency semantic and high-frequency textural bands via channel-wise 2D-DCT, injecting aligned watermark identities prior to VAE decoding with cross-band consistency constraints to achieve 37.40 dB PSNR and superior regeneration robustness with just 1 ms overhead.
Background & Motivation¶
With denoising diffusion probabilistic models (DDPM) and latent diffusion models (LDM) establishing dominance in photorealistic image synthesis, machine-generated visual content has reached perceptual parity with real-world imagery, precipitating urgent challenges in copyright provenance, unauthorized dissemination, and deepfake accountability. Invisible watermarking offers a primary path for verifiable provenance attribution. However, conventional pixel-domain post-processing watermarking techniques incur substantial computational latency and remain notoriously fragile against "regeneration attacks"—adversarial pipelines that inject forward noise into the watermarked image and denoise it through the reverse diffusion trajectory, thoroughly stripping fragile pixel modifications.
While latent-space watermarking substantially amortizes inference overhead, prevailing frameworks treat intermediate latent tensors strictly as flat spatial feature maps, disregarding the rich hierarchical frequency structure intrinsic to the latent manifold. Meanwhile, frequency-domain pioneers such as Tree-Ring embed static circular Fourier patterns into the initial Gaussian noise and rely on inversion-based recovery, failing to adaptively learn dual-band representations end-to-end. Anchoring identity marks to a single representational mode—spatial or fixed-frequency—leaves them fundamentally vulnerable to structural perturbations, compression artifacts, or regeneration filtering. Crucially, the low-frequency components representing global semantic layout and the high-frequency components encoding fine-grained textures exhibit natural, complementary resilience profiles under distinct physical and algorithmic distortions.
To bridge this representational vulnerability, this paper attacks the problem by establishing a dual-spectral collaborative embedding pipeline directly on the final clean latent representation before VAE decoding. Core idea: decompose diffusion latent representations via channel-wise discrete cosine transform (DCT) into complementary low-frequency semantic and high-frequency textural spectral bands, independently injecting aligned watermark identities via learned residual encoders with cross-band consistency constraints to construct bi-spectral redundancy against regeneration and image distortions.
Method¶
Overall Architecture¶
The BiSLW pipeline operates precisely at the interface between the clean latent tensor \(z_0 \in \mathbb{R}^{C \times H \times W}\) produced by the reverse diffusion trajectory and the frozen pre-trained VAE decoder \(D\). The latent tensor is transformed into the frequency domain via channel-wise two-dimensional discrete cosine transform (2D-DCT) and split via a radial binary mask into a low-frequency semantic band and a high-frequency textural band. Target binary message bits \(w \in \{-1, +1\}^{d_w}\) modulate lightweight residual perturbation networks via Feature-wise Linear Modulation (FiLM) to adaptively inject identity perturbations into both bands. The modified spectrum is losslessly recomposed and mapped back to the spatial latent domain via inverse DCT before being rendered by the VAE decoder. During extraction, a distorted image is re-encoded into latent space, spectrally decomposed, and decoded by dual spectral extraction heads under a cross-band consistency criterion.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Clean Latent z₀ from Diffusion<br/>+ Watermark Message w"] --> S1["Latent Spectral Decomposition<br/>Channel-wise 2D-DCT + Radial Mask"]
S1 --> S2["Bi-Spectral Watermark Embedding<br/>FiLM-Conditioned Residual Injection"]
S2 --> S3["Spectral Inversion & Reconstruction<br/>Inverse 2D-DCT & VAE Decoding"]
S3 --> Out["Watermarked Image x̃"]
Out -.->|Distortion / Attack Channel| Ex["Image VAE Re-encoding E(x̃')"]
Ex --> S4["Dual-Band Spectral Extraction & Voting<br/>DL/DH Decoders + Sign Voting"]
S4 --> Rec["Recovered Watermark ŵ"]
Key Designs¶
1. Latent Spectral Decomposition: Radial Partitioning of Complementary Bands Spatial latent perturbations easily wash out when subjected to reverse diffusion denoising or Gaussian filtering. To resolve this fragility, BiSLW capitalizes on the frequency hierarchy of latent diffusion tensors: low-frequency DCT coefficients govern coarse spatial layout and composition, while high-frequency coefficients dictate fine local texture. The clean latent tensor \(z_0\) undergoes channel-wise 2D-DCT: $\(Z_c^{\mathrm{freq}} = \mathrm{DCT}_{2D}(\mathbf{z}_{0,c}), \quad c=1,\ldots,C\)$ Because the DC coefficient dominates AC coefficients by orders of magnitude—which would destabilize neural training dynamics—BiSLW performs channel-wise z-score normalization on \(Z^{\mathrm{freq}}\) prior to decomposition. A radial binary mask centered at the low-frequency origin partitions the spectrum: $\(\mathcal{M}[i, j] = \begin{cases} 1, & \text{if } \sqrt{(i/H)^2 + (j/W)^2} \le r \\ 0, & \text{otherwise} \end{cases}\)$ yielding \(Z^{\mathrm{low}} = Z^{\mathrm{freq}} \odot \mathcal{M}\) and \(Z^{\mathrm{high}} = Z^{\mathrm{freq}} \odot (1 - \mathcal{M})\). This decomposition preserves exact lossless additivity \(Z^{\mathrm{low}} + Z^{\mathrm{high}} = Z^{\mathrm{freq}}\), isolating complementary spectral substrates that exhibit divergent survival characteristics under distinct attacks.
2. Bi-Spectral Watermark Embedding: FiLM-Conditioned Adaptive Residual Modulation Unlike rigid, hand-crafted Fourier patterns that lack adaptability to diverse visual content, BiSLW deploys two lightweight perturbation networks, \(\Delta_L\) and \(\Delta_H\), built from convolutional encoder-decoder blocks with residual skip connections. The binary watermark vector \(w\) is projected through Feature-wise Linear Modulation (FiLM) layers to produce per-channel affine scaling and shift factors that condition intermediate latent features. The spectral residuals are injected via: $\(\tilde{Z}^{\mathrm{low}} = Z^{\mathrm{low}} + \alpha_L \Delta_L(Z^{\mathrm{low}}, \mathbf{w}), \quad \tilde{Z}^{\mathrm{high}} = Z^{\mathrm{high}} + \alpha_H \Delta_H(Z^{\mathrm{high}}, \mathbf{w})\)$ where \(\alpha_L\) and \(\alpha_H\) denote embedding strength scalars. Crucially, the authors configure \(\alpha_L > \alpha_H\) (empirically \(\alpha_L=0.8, \alpha_H=0.3\)); low-frequency latent components can absorb higher-energy modifications while surviving severe low-pass filtering and geometric distortions, whereas high-frequency injection operates with lower energy to protect visual textures from high-frequency artifacts.
3. Spectral Inversion & Reconstruction: Pre-Decoding Intrinsic Integration Rather than altering pixels post-hoc, BiSLW sums the watermarked spectral bands and inverts the z-score normalization to obtain \(\tilde{Z}^{\mathrm{freq}} = \tilde{Z}^{\mathrm{low}} + \tilde{Z}^{\mathrm{high}}\). The spatial latent tensor is recovered via channel-wise inverse 2D-DCT: $\(\tilde{\mathbf{z}}_0 = \mathcal{T}^{-1}(\tilde{Z}^{\mathrm{freq}})\)$ The final watermarked image is synthesized through the frozen VAE decoder \(\tilde{x} = D(\tilde{\mathbf{z}}_0)\). Because perturbations are structured inside the latent frequency domain before decoding, they pass through the non-linear VAE generative manifold, aligning naturally with image reconstruction statistics and eliminating the color banding or boundary halos that plague direct pixel modifications in smooth image regions.
4. Dual-Band Spectral Extraction & Voting: Cross-Band Alignment & Redundant Ensembling During extraction from a distorted image \(\tilde{x}'\), the image is mapped to latent space \(z' = E(\tilde{x}')\) via the frozen VAE encoder, spectrally transformed, and partitioned into \(Z'^{\mathrm{low}}\) and \(Z'^{\mathrm{high}}\). Two dedicated spectral decoders extract bit predictions: $\(\hat{\mathbf{w}}_L = D_L(Z'^{\mathrm{low}}), \quad \hat{\mathbf{w}}_H = D_H(Z'^{\mathrm{high}})\)$ To prevent identity drift between the two pathways during training, BiSLW penalizes spectral discrepancy via a cross-band consistency loss \(\mathcal{L}_{\mathrm{cons}} = \|\hat{\mathbf{w}}_L - \hat{\mathbf{w}}_H\|_2^2\). At inference, the final bit estimate combines both predictions via mean ensembling \(\hat{\mathbf{w}} = \frac{1}{2}(\hat{\mathbf{w}}_L + \hat{\mathbf{w}}_H)\) followed by a binary sign function. When low-pass filtering degrades high-frequency coefficients, the semantic low-frequency branch preserves decoding fidelity; conversely, when spatial crops or localized edits disrupt low frequencies, the high-frequency textural branch compensates.
Loss & Training¶
The pre-trained diffusion backbone and VAE encoder/decoder are frozen throughout training; only \(\{\Delta_L, \Delta_H, D_L, D_H\}\) are optimized over 100,000 MIRFLICKR-1M images. The joint objective balances watermark recovery, cross-band alignment, latent preservation, and distortion robustness: $\(\mathcal{L} = \lambda_w \mathcal{L}_w + \lambda_{\mathrm{cons}} \mathcal{L}_{\mathrm{cons}} + \lambda_z \mathcal{L}_z + \lambda_{\mathrm{rob}} \mathcal{L}_{\mathrm{rob}}\)$ Watermark recovery enforces \(\|\hat{\mathbf{w}}_L - \mathbf{w}\|_2^2 + \|\hat{\mathbf{w}}_H - \mathbf{w}\|_2^2\), latent fidelity penalizes drift via \(\|\tilde{\mathbf{z}}_0 - \mathbf{z}_0\|_2^2\), and the robustness loss computes extraction error under differentiable attacks sampled from \(\mathcal{P}_A\) (JPEG compression, Gaussian blur, additive noise, random cropping, and diffusion regeneration). Loss weights are balanced at \(\lambda_w=1.0, \lambda_{\mathrm{cons}}=0.5, \lambda_{\mathrm{rob}}=1.0, \lambda_z=2.0\) with split radius \(r=0.25\).
Key Experimental Results¶
Main Results¶
On AI-generated images (32-bit watermark) and the CLIC dataset (48-bit watermark), BiSLW was benchmarked against eight classical and deep watermarking baselines across nine individual distortions and a composite attack (Comb.):
| Method (bit#) | PSNR/SSIM ↑ | Emb. time (ms) | None Acc ↑ | C.Crop Acc ↑ | Rot. 15° Acc ↑ | Blur Acc ↑ | JPEG-70 Acc ↑ | Comb. Acc ↑ | Avg. Acc ↑ |
|---|---|---|---|---|---|---|---|---|---|
| DCT-DWT (32) | 39.47 / 0.97 | 139 | 0.91 | 0.51 | 0.52 | 0.51 | 0.51 | 0.51 | 0.55 |
| HiDDeN (30) | 32.59 / 0.95 | 17 | 0.91 | 0.91 | 0.79 | 0.76 | 0.53 | 0.59 | 0.77 |
| SSL (32) | 33.23 / 0.89 | 870 | 1.00 | 0.74 | 0.99 | 1.00 | 0.99 | 0.85 | 0.92 |
| RoSteALS (32) | 29.31 / 0.93 | 144 | 1.00 | 0.50 | 0.47 | 1.00 | 1.00 | 0.50 | 0.77 |
| RivaGAN (32) | 40.53 / 0.98 | 57 | 0.99 | 0.98 | 0.91 | 0.99 | 0.98 | 0.93 | 0.92 |
| Tree-Ring (32) | 33.64 / 0.90 | 46 | 0.88 | 0.81 | 0.72 | 0.88 | 0.83 | 0.81 | 0.83 |
| LaWa* (32) | 34.25 / 0.89 | 1 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.97 | 0.99 |
| BiSLW* (32, Ours) | 37.40 / 0.91 | 1 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.98 | 0.99 |
| BiSLW-post-gen (32) | 37.42 / 0.91 | 33 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.98 | 0.99 |
Generative quality, manifold drift, and regeneration resilience (Table 2):
| Method | FID ↓ | CLIP Sim. ↑ | PSNR (dB) ↑ | KL Shift ↓ | Latent Shift ↓ | Regen. Acc (\(t^*=250/500\)) ↑ |
|---|---|---|---|---|---|---|
| Vanilla SD | 8.7 | 0.312 | — | — | — | — |
| LaWa* | 9.8 | 0.307 | 34.25 | 0.021 | 0.014 | 0.94 / 0.89 |
| Tree-Ring | 9.9 | 0.302 | 33.64 | 0.024 | 0.018 | 0.88 / 0.81 |
| BiSLW* (Ours) | 9.0 | 0.311 | 37.40 | 0.018 | 0.011 | 0.96 / 0.92 |
Ablation Study¶
To isolate the contribution of dual-band redundancy, BiSLW was ablated against single-band baselines on AI-generated images with 48-bit watermark payload (Table 4):
| Embedding Strategy | PSNR (dB) ↑ | SSIM ↑ | Bit Acc. (Comb.) ↑ | Note |
|---|---|---|---|---|
| Low-frequency only | 38.96 | 0.92 | 0.88 | Highest visual fidelity, but vulnerable under severe composite attacks |
| High-frequency only | 37.12 | 0.89 | 0.90 | Higher resilience against distortions, but compromises image quality |
| BiSLW (Low + High) | 37.71 | 0.89 | 0.93 | Combines both bands for optimal robustness and balanced fidelity |
Ablation on the latent fidelity regularization weight \(\lambda_z\) (48-bit payload, Table 5):
| Regularization \(\lambda_z\) | PSNR (dB) ↑ | SSIM ↑ | Bit Acc. (None) ↑ | Bit Acc. (Comb.) ↑ |
|---|---|---|---|---|
| 0.5 | 36.85 | 0.86 | 1.00 | 0.96 |
| 2.0 (Selected default) | 37.71 | 0.89 | 0.99 | 0.93 |
| 5.0 | 38.53 | 0.91 | 0.98 | 0.91 |
| 10.0 | 39.14 | 0.92 | 0.98 | 0.89 |
| 20.0 | 39.59 | 0.92 | 0.97 | 0.88 |
Key Findings¶
- Dual-Band Redundancy Outperforms Single-Band Variants: Low-frequency-only embedding yields the highest PSNR (38.96 dB) but drops to 0.88 accuracy under composite attacks. High-frequency-only attains 0.90 accuracy with lower quality. The full bi-spectral framework synergistically elevates composite accuracy to 0.93, verifying that cross-band complementarity successfully covers disparate attack vulnerabilities.
- Superior Regeneration Attack Resilience: When subject to diffusion regeneration attacks at \(t^*=500\) (where substantial noise replaces structured latent features), BiSLW retains 0.92 bit accuracy, significantly outperforming LaWa (0.89) and Tree-Ring (0.81).
- Minimal Generative Manifold Degradation: BiSLW achieves an FID of 9.0, closely matching unwatermarked Vanilla SD (8.7) and outperforming LaWa (9.8). Its KL shift (0.018) and latent drift (0.011) are the lowest among all evaluated methods, showing that bi-spectral residuals preserve the latent manifold distribution.
Highlights & Insights¶
- From Flat Latent Maps to Dual-Spectral Hierarchy: Transcends conventional spatial latent watermarking by recognizing the intrinsic frequency hierarchy in diffusion latents, exploiting channel-wise 2D-DCT to build orthogonal semantic and textural channels.
- Adaptive Learned Injection over Hardcoded Patterns: Rather than injecting static circular rings into initial Gaussian noise, BiSLW dynamically modulates the final latent tensor via FiLM conditioning, combining learned non-linear adaptability with spectral robustness.
- Cross-Band Alignment with Negligible Latency: The explicit cross-band consistency constraint \(\mathcal{L}_{\mathrm{cons}}\) guarantees zero identity drift, while lightweight latent convolutions execute in just 1 ms during sampling, introducing zero practical bottleneck to generation throughput.
Limitations & Future Work¶
- Rotation Degradation under High Payloads: When scaling to 96 or 128 bits under 15° rotation, bit accuracy degrades (0.77 at 96 bits) because spatial grid rotation disperses 2D-DCT frequency energy. Incorporating polar-coordinate transforms or Spatial Transformer Networks (STN) could mitigate rotational phase shift.
- Static Radial Frequency Cutoff: The fixed radial threshold \(r=0.25\) partitions frequency bands statically. An adaptive frequency allocation network conditioned on latent spectral entropy could further optimize trade-offs between texture-heavy and flat scenes.
Related Work & Insights¶
- vs Tree-Ring Watermarks: Tree-Ring hardcodes circular Fourier rings into initial noise and requires deterministic DDIM inversion for extraction, precluding end-to-end training and large payload scaling. BiSLW operates on the clean latent before decoding via learnable networks, supporting direct multi-bit payload decoding.
- vs Stable Signature / LaWa: Stable Signature fine-tunes the VAE decoder weights, whereas LaWa injects spatial latent perturbations without frequency awareness. BiSLW freezes both the diffusion backbone and VAE, achieving a 3.15 dB PSNR improvement over LaWa while sustaining matching or superior robustness.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Pioneering frequency-domain decomposition of diffusion latent tensors for learned dual-band watermarking]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across 9 individual attacks, composite distortions, regeneration curves, payload scaling, and manifold drift]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically rigorous formulation, and coherent experimental presentation]
- Value: ⭐⭐⭐⭐⭐ [Provides a practical, high-fidelity, 1 ms watermarking solution critical for generative content attribution and security]