Skip to content

DiffUE: Enhancing Utility-Unlearnability Trade-off of Unlearnable Examples via Diffusion Autoencoders

Conference: ECCV 2026
Paper: ECCV Official
Code: To be confirmed
Area: AI Safety / Human Understanding
Keywords: Unlearnable Examples, Diffusion Autoencoders, Semantic Perturbations, Data Privacy, Adversarial Relearning

TL;DR

DiffUE leverages a pre-trained diffusion autoencoder (DiffAE) to map defensive perturbations of unlearnable examples from conventional pixel space into a disentangled semantic latent space, utilizing bi-level minimization optimization alongside directional attribute manipulation to preserve natural image manifold structures and visual fidelity while achieving superior robustness against adversarial training, filtering, relearning attacks, and lossy compression.

Background & Motivation

The rapid advancement of deep learning and foundational models has been largely fueled by vast datasets scraped from social media and public platforms without explicit user consent. This rampant unauthorized data harvesting poses severe privacy threats, such as non-consensual facial recognition and exploitative targeted profiling. Although regulatory frameworks like GDPR and CCPA have been enacted, they primarily focus on pre-collection governance; once personal data enters commercial AI training pipelines, users lack viable technical mechanisms to prevent continued unauthorized exploitation. To counteract this, researchers proposed unlearnable examples (UEs), which defensively inject imperceptible perturbations into images to induce near-zero training loss, creating the illusion that the data has already been learned and effectively suppressing gradient signals required for meaningful feature extraction.

However, existing unlearnable example techniques face an acute dilemma between defensive fragility and practical usability. On one hand, current methods rely on pixel-space error-minimizing noise that is easily circumvented by adaptive relearning strategies—such as adversarial training, grayscale transformation, adversarial data augmentations (e.g., UEraser), and diffusion purification pipelines (e.g., LUE, Avatar). On the other hand, attempts to enhance robustness (e.g., REM, SEM) by increasing perturbation budgets or incorporating adversarial/random components introduce conspicuous high-frequency noise, color banding, and texture corruption directly in pixel space. For photographers, digital artists, and everyday social media users, such degraded visuals are entirely unacceptable for public sharing and professional portfolios.

To resolve this fundamental tension, this paper shifts the defensive battlefield away from pixel space to high-level image semantics on the natural data manifold. Core idea: leverage a pre-trained diffusion autoencoder (DiffAE) to disentangle input images into semantic and stochastic latent codes, inject defensive error-minimizing noise into the continuous semantic latent space, and reconstruct unlearnable examples via conditional DDIM decoding, thereby eliminating high-frequency pixel artifacts while enabling controlled, aesthetically pleasing image edits that provide robust defense against unauthorized model training.

Method

Overall Architecture

The core architecture of DiffUE is built on a pre-trained diffusion autoencoder (DiffAE). Given an input image \(x_0\), a semantic encoder \(\mathcal{E}\) and a conditional DDIM encoder extract two disentangled representations: a 512-dimensional continuous latent vector \(z_{sem} = \mathcal{E}(x_0) \in \mathbb{R}^{512}\) capturing global, non-spatial semantics, and a stochastic latent code \(x_T\) preserving low-level spatial geometry and textural details. The defensive objective is to search for a semantic perturbation \(z_\delta\) bounded in norm, forming a perturbed defensive semantic representation \(z_{def} = z_{sem} + z_\delta\). Subsequently, a conditional DDIM decoder \(\mathcal{D}\) takes \(z_{def}\) as the conditioning vector and performs deterministic reverse denoising starting from \(x_T\), synthesizing the protected unlearnable image \(x_u = \mathcal{D}(z_{sem} + z_\delta, x_T)\). When fed into a surrogate classification model \(g_\theta\), \(x_u\) induces a near-zero cross-entropy loss through bi-level optimization, suppressing the model's ability to extract genuine discriminative features.

The end-to-end framework comprises dual-latent disentangled encoding, semantic defensive perturbation optimization, optional controlled attribute guidance, and conditional DDIM deterministic decoding:

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Clean Image x0"] --> B["Dual-Latent Disentangled Encoding<br/>Semantic encoder E yields zsem + stochastic code xT"]
    B --> C["Semantic Latent Defense Optimization<br/>Projected Gradient Descent PGD iterates zδ"]
    C --> D["Directional Semantic Edit Guidance<br/>Cosine similarity aligns with target attribute direction zM (optional)"]
    D --> E["Conditional DDIM Deterministic Decoding<br/>Reverse sampling conditioned on zdef and xT yields xu"]
    E --> F["Surrogate Model Bi-Level Game<br/>Induces near-zero loss to suppress learning gradients"]

Key Designs

1. Semantic Latent Defense Optimization: Shifting from Pixel Corruption to Manifold Semantic Shifts Conventional unlearnable example methods directly superimpose norm-bounded additive noise \(\delta\) on the RGB pixel grid. Such high-frequency perturbations deviate sharply from the natural image distribution, rendering them vulnerable to low-pass filters, spatial smoothing, and compression codecs. DiffUE operates within the compact, continuous latent space of DiffAE, constraining the defensive perturbation strictly to the high-level semantic vector \(z_{sem}\). During optimization, Projected Gradient Descent (PGD) iteratively updates the defensive noise \(z_\delta\) within an \(L_p\)-norm ball of radius \(\rho_u\) in the direction that minimizes surrogate classification loss: $\(z_{\delta}^{(t+1)} = \Pi_{\rho_u} \left( z_{\delta}^{(t)} - \eta \cdot \mathrm{sign}\left(\nabla_{z_\delta} \mathcal{L}_{ule}\right) \right)\)$ where \(\Pi_{\rho_u}\) denotes the projection clipping operator, \(\eta\) is the step size, and \(\mathcal{L}_{ule}\) is the surrogate objective function. Because perturbations are confined to the learned image manifold, decoding them back to image space translates into coherent, natural visual variations—such as mild illumination shifts or subtle tonal adjustments—without foreign pixel-level artifacts. Furthermore, because neural networks exhibit acute sensitivity to latent semantic features, minor semantic perturbations drastically alter model representations, enabling DiffUE to achieve potent unlearnability with a much smaller perturbation budget while inherently resisting pixel-domain augmentations.

2. Surrogate Model Bi-Level Game: Aligning with Empirical Model Learning Dynamics Unlike adversarial examples that seek to induce classification failures during inference, unlearnable examples must manufacture artificial shortcuts that mislead the model into perceiving the training samples as already mastered, effectively starving back-propagation gradients. Consequently, the defensive noise must align with the parameter trajectory of neural networks during initial training. DiffUE models the synthesis of unlearnable examples as a min-min bi-level optimization problem: $\(\min_{\theta} \; \mathbb{E}_{(x_0, y)} \left[ \min_{\|z_\delta\|_p \le \rho_u} \mathcal{L}\left( g_\theta\left( \mathcal{D}(\mathcal{E}(x_0) + z_\delta, x_T) \right), y \right) \right]\)$ The inner minimization freezes model parameters \(\theta\) and updates the semantic noise \(z_\delta\) to minimize classification loss, while the outer minimization optimizes \(\theta\) to further depress the training error. To ensure realistic gradient trajectories, the surrogate network \(\theta\) is pre-trained for \(Q\) warm-up steps before updating \(z_\delta\). In each outer step, 10 inner PGD iterations are performed. This alternating optimization continues until the surrogate model achieves a training accuracy exceeding a threshold \(\lambda\) (e.g., 99%), ensuring that the injected semantic shortcut fully obscures genuine sample representations.

3. Directional Semantic Edit Guidance: Transforming Defensive Noise into Purposeful Enhancements To overcome the compromise where privacy protection degrades visual aesthetics, DiffUE repurposes unconstrained defensive noise into purposeful, visually appealing photo edits. Benefiting from the linear controllability of the DiffAE latent space, the authors train lightweight linear classifiers on semantic codes (\(z_{sem}\)) to identify target visual attributes (e.g., smiling, hairstyle, gender) or stylistic color filters. The normalized weight vector of the attribute classifier defines a target semantic direction \(z_M = w_{cls} / \|w_{cls}\|\). The defensive loss is then augmented with a directional alignment penalty: $\(\mathcal{L}_{ule} = \mathcal{L}\left( g_\theta\left( \mathcal{D}(\mathcal{E}(x_0) + z_\delta, x_T) \right), y \right) + \gamma \cdot \left( 1 - \cos(z_\delta, z_M) \right)\)$ where \(\cos(\cdot)\) denotes cosine similarity and \(\gamma\) is a balancing hyperparameter governing the edit intensity. By aligning the perturbation vector with \(z_M\), the defensive noise steers image attributes toward intended enhancements (such as adding a natural smile or professional portrait grading) while simultaneously suppressing model learnability. This transforms privacy defense from a destructive nuisance into a constructive, user-embraced enhancement.

Loss & Training

DiffUE employs ResNet-18 as the default surrogate model \(g_\theta\). Because latent space Euclidean distances do not directly correspond to pixel shifts due to nonlinear decoding, the semantic perturbation radius \(\rho_u\) was determined empirically: by perturbing 50 out of 512 latent dimensions across 2,000 samples per dataset, the authors set \(\rho_u = 0.20\) for CIFAR-10/100 and \(\rho_u = 0.35\) for ImageNet. These correspond to a maximum pixel shift of only 4/255—half the 8/255 budget allocated to pixel-space baselines. For controlled attribute editing on CelebA-HQ, \(\rho_u\) is set to 0.35 to generate visible modifications. Bi-level optimization runs with 10 inner PGD steps per surrogate step until surrogate training accuracy exceeds 99%.

Key Experimental Results

Main Results

To evaluate the resilience of DiffUE against adversarial training (AT) relearning countermeasures, experiments were conducted on CIFAR-10, CIFAR-100, and an ImageNet Subset (first 100 classes) using ResNet-18. In this adversarial relearning setting, attackers train models with an adversarial perturbation radius \(\rho_t\) to disrupt unlearnable protections. Lower clean test accuracy reflects superior defensive protection.

Dataset Adv. Train. Radius \(\rho_t\) Clean Baseline EM [9] REM [4] (\(\rho_a=4/255\)) SEM [20] (\(\rho_r=4/255\)) Ours (DiffUE)
CIFAR-10 0 95.64% 11.98% 31.27% 25.14% 10.79%
2/255 89.24% 20.96% 37.75% 26.65% 12.59%
4/255 86.92% 76.95% 64.64% 46.85% 17.11%
CIFAR-100 0 85.32% 1.76% 19.30% 18.61% 2.17%
2/255 76.76% 67.29% 18.35% 12.56% 3.78%
4/255 72.19% 64.59% 32.22% 26.19% 4.92%
ImageNet Subset 0 82.54% 1.89% 17.74% 10.11% 2.11%
2/255 73.22% 70.67% 31.40% 36.22% 12.45%
4/255 70.23% 67.86% 45.49% 38.57% 17.20%

Ablation Study & Robustness Under Relearning

Table below summarizes classification test accuracy on CIFAR-10 under spatial filtering and advanced relearning strategies (lower accuracy indicates greater robustness), along with image perceptual quality metrics on CelebA-HQ and ImageNet (lower FID is better; higher SSIM and PSNR are better).

Evaluation Domain Metric / Attack Config EM [9] REM [4] SEM [20] Ours (DiffUE) Performance Advantage Note
CIFAR-10 Spatial Filtering Mean Filter 37.87% 29.92% 28.55% 13.32% Unlearnability fully preserved against spatial smoothing
Gaussian Filter 32.71% 28.53% 24.22% 12.43% Outperforms runner-up by >11.7% margin
GrayScale Transform 83.23% 63.87% 34.23% 13.99% Complete immunity against color-channel stripping
CIFAR-10 Relearning Attacks UEraser (Aug. + AT) 90.23% 87.39% 81.56% 14.19% Baselines breached (>81%), DiffUE remains solidly defensive
LUE (Diffusion Purification) 90.33% 89.57% 89.29% 33.97% Massive >55% protection margin against diffusion removal
Lossy Image Compression CIFAR-10: JPEG 45.22% 38.12% 32.92% 17.22% Mitigates accuracy resurgence by 28.0% vs EM
CIFAR-10: WebP 55.57% 41.78% 35.96% 18.05% Robust protection under advanced modern compression
Perceptual Quality (CelebA-HQ) FID (↓) 11.612 10.971 10.145 5.355 47.2% reduction in FID, aligning with natural image distribution
SSIM (↑) / PSNR (↑) 0.704 / 30.11 0.561 / 29.41 0.611 / 29.87 0.814 / 32.44 Leading in structural fidelity and peak signal-to-noise ratio

Key Findings

  • Decisive Resilience Against Relearning Attacks: When evaluated against UEraser (which combines adversarial data augmentations with adversarial training to defeat UEs), all pixel-based baselines collapse completely (EM 90.23%, REM 87.39%, SEM 81.56%), whereas DiffUE holds model accuracy down to 14.19%. Against diffusion purification (LUE), which eliminates defensive noise through joint-conditional diffusion, DiffUE limits model accuracy to 33.97%, outperforming baselines by over 55 percentage points.
  • Robustness Against Lossy Compression: Real-world image sharing typically involves JPEG, WebP, or HEIC compression. Pixel-level high-frequency noise is severely compromised by compression algorithms, causing model test accuracy to rebound significantly (e.g., EM rebounding to 45.22% on CIFAR-10 under JPEG). In contrast, DiffUE modifies low-frequency semantic attributes on the image manifold, retaining low test accuracy (17.22% on CIFAR-10 and 4.12% on ImageNet).
  • Strong Cross-Architecture Transferability: When evaluating unlearnable examples crafted on ResNet-18 across VGG-16, ResNet-50, and DenseNet-121, Grad-CAM visualizations reveal that all architectures divert their attention entirely away from key discriminative features (e.g., a dog's face). On CIFAR-100, transferring to DenseNet-121 suppresses accuracy to 10.98% (compared to 50%–66% for baselines).
  • Overwhelming User Study Preference: In a double-blind user study with 56 professional photographers, digital artists, and content creators, DiffUE achieved a naturalness score of 4.65 ± 0.12 (near the clean image benchmark of 4.90 ± 0.10) and an acceptability rating of 98.5% (compared to only 19.2% for REM), ranking first in overall preference with an average rank of 1.22.

Highlights & Insights

  • Shifting Defensive Paradigms from Pixel to Latent Manifold: By moving beyond vulnerable pixel-space additive noise to the generative semantic manifold of diffusion autoencoders, DiffUE achieves an elegant separation between spatial structural details and global semantic cues, rendering the defense impervious to spatial filtering and augmentations.
  • Harmonizing Data Privacy with Aesthetic Utility: Aligning error-minimization noise with semantic attribute directions (such as adding a smile or subtle tone enhancement) creatively transforms privacy preservation into an enjoyable user feature, reconciling the historical trade-off between protection strength and image usability.
  • Cross-Domain Reusability: The concept of injecting unlearnable constraints into diffusion latent representations can naturally extend to other generative modalities, including text-to-image personalization protection, 3D Gaussian Splatting asset defense, and audio diffusion watermarking.

Limitations & Future Work

  • Computational Overhead of Diffusion Sampling: Generating UEs requires multiple DDIM denoising steps during each inner optimization loop along with back-propagation through the diffusion network, resulting in significantly higher generation time compared to simple first-order pixel methods. Future work could incorporate latent consistency models or flow matching for faster synthesis.
  • Dependence on Pre-trained Autoencoder Distribution: The fidelity and unlearnability of DiffUE rely on the semantic representation capacity of the underlying pre-trained DiffAE. For out-of-distribution inputs (such as specialized medical or aerial imagery), latent reconstruction artifacts may arise.
  • Future Directions: The authors suggest exploring adaptive, sample-specific semantic noise budget selection and extending semantic unlearnability frameworks to multimodal large language models (MLLMs) and temporal video tracking defenses.
  • vs EM (Huang et al., ICLR 2021): EM pioneered unlearnable examples via pixel-space error-minimizing noise but suffers from acute fragility under adversarial training and spatial filters; DiffUE transposes the defense into semantic space, achieving superior robustness and visual fidelity.
  • vs REM (Fu et al., ICLR 2022) & SEM (Liu et al., AAAI 2024): REM and SEM incorporate adversarial or random noise in pixel space to resist relearning, but this severely degrades visual quality (FID > 10.0, user acceptability down to 19.2%); DiffUE achieves stronger defense with half the effective pixel shift while maintaining near-clean visual fidelity.
  • vs LUE (Jiang et al., ACM MM 2023) & Avatar (Dolatabadi et al., SaTML 2024): These works deploy diffusion models as purification countermeasures to strip defensive noise from UEs; DiffUE demonstrates that defensive noise embedded directly in diffusion semantic latent representations effectively withstands diffusion-based purification.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering shift of unlearnable examples into diffusion latent semantic space with controlled attribute editing]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across 4 benchmarks, 5 relearning attacks, multiple compression standards, cross-architecture transfer, and a 56-participant user study]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically rigorous formulation, comprehensive empirical discussions, and clean visual presentation]
  • Value: ⭐⭐⭐⭐⭐ [Provides a practical, robust, and user-friendly defense paradigm for safeguarding personal image privacy in the generative AI era]