EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/HiDream-ai/EraseSAE
Area: Video Generation / AI Safety
Keywords: Concept Erasure, Sparse Autoencoders, Text-to-Video Diffusion Models, Monosemantic Features, Spatiotemporal Masks
TL;DR¶
Addressing concept entanglement and collateral background degradation in text-to-video diffusion models, EraseSAE introduces a decompose-attribute-erase paradigm leveraging a dual-branch Partitioned Convolutional Sparse Autoencoder (PConvSAE) to extract monosemantic features and dynamic spatiotemporal masks for spatially-modulated classifier-free guidance, achieving surgical erasure with minimal side effects.
Background & Motivation¶
Diffusion Transformer (DiT)-based text-to-video (T2V) models have demonstrated remarkable generative capabilities, synthesizing high-fidelity dynamic visual scenes from arbitrary textual prompts. However, because these models rely heavily on web-scale, loosely curated training datasets, they inevitably memorize and generate copyright-infringing content, explicit material, and deepfakes of public figures. While retraining on sanitized data offers a straightforward solution, the immense computational burden makes it practically infeasible. Concept erasure has consequently emerged as a principal post-training remedy to selectively eliminate sensitive semantics while maintaining the model's remaining generative capacity. Prior approaches broadly divide into training-free inference steering (such as negative prompting and orthogonal text embedding projection) and training-based parameter editing (such as cross-attention fine-tuning, closed-form weight manipulation, and neuron pruning). Nevertheless, both paradigms suffer from severe granularity mismatch.
The core tension lies in the fundamental discrepancy between the coarse intervention level of existing methods and the fine-grained, polysemantic nature of neural concept representations. In deep networks, individual neurons simultaneously encode multiple distinct semantic concepts. Modifying global weights induces catastrophic forgetting of benign content, localized cross-attention editing leaves latent traces vulnerable to adversarial probing, and neuron pruning inflicts irreversible collateral damage on co-encoded concepts. Compounding this challenge, the 3D full-attention mechanism in DiT models couples concepts across both spatial coordinates and temporal frames. Existing methods lack spatiotemporal locality: applying unconstrained global suppression across the entire video canvas irreversibly corrupts unrelated background textures, scene illumination, and temporal coherence even when the target concept occupies only a localized spatio-temporal region.
Overcoming this bottleneck requires shifting the intervention granularity to monosemantic units where each latent feature maps strictly to a single interpretable concept. Core idea: leverage a Partitioned Convolutional Sparse Autoencoder (PConvSAE) that preserves native spatiotemporal topology while decomposing dense representations into decoupled monosemantic concept features and scene context, isolate pure concept kernels via contrastive log-ratio attribution against hard negatives, and execute surgical concept removal through timestep-resolved spatiotemporal masks paired with spatially-modulated classifier-free guidance.
Method¶
Overall Architecture¶
EraseSAE establishes a principled three-stage decompose-attribute-erase pipeline tailored for DiT-based T2V diffusion architectures. In the decompose stage, an identified transformer intervention layer is equipped with a dual-branch Partitioned Convolutional Sparse Autoencoder (PConvSAE). By folding the temporal dimension into the batch axis, PConvSAE directly applies 2D convolutions over native feature topologies, disentangling dense representations into a Context Branch (capturing concept-agnostic scene canvas) and a Concept Branch (isolating localized concept semantics). In the attribute stage, a one-time offline contrastive procedure compares activation distributions across paired positive and negative prompts using a log-ratio scoring mechanism against hard-negative baselines, locking a compact set of high-purity concept kernels. In the erase stage, the locked kernels dynamically track the evolving concept geometry across denoising timesteps to synthesize spatiotemporal masks, guiding a spatially-modulated classifier-free guidance (SM-CFG) mechanism that confines negative guidance strictly to active target regions while leaving global background content untouched.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Dense Spatiotemporal Activation Tensor<br/>X ∈ R^(B×T)×D×H×W"] --> B["Dual-Branch Disentangled PConvSAE<br/>Topology Preservation + Spatial-Aware Local Activation"]
B --> C["Contrastive Log-Ratio Attribution<br/>One-vs-Max Hard-Negative Exclusion + Stability Filtering"]
C --> D["Dynamic Spatiotemporal Mask SM-CFG<br/>Timestep-Resolved Mask Intersection + Localized CFG"]
D --> E["Output: Surgically Sanitized Video<br/>Target Semantics Excised & Background Fully Intact"]
Key Designs¶
1. Dual-Branch Disentangled PConvSAE: Topology Preservation and Semantic Separation Conventional MLP-based sparse autoencoders flatten high-dimensional hidden activations into one-dimensional feature vectors, completely destroying spatial neighborhood topologies and temporal coherence, which leads to blurred spatial localization and temporal flickering. PConvSAE preserves native tensor geometry by folding the temporal dimension into the batch dimension and employing 2D spatial convolutional encoders and decoders. Crucially, the architecture partitions the latent space into two structurally decoupled, non-parameter-sharing branches: a Context Branch (\(f_{\text{ctx}}\)) dedicated to global structural priors, background textures, and illumination, and a Concept Branch (\(f_{\text{cpt}}\)) segmented into \(C\) mutually exclusive channel subspaces corresponding to target concepts. To enforce spatial coherence, a Spatial-Aware Local Activation mechanism identifies the top-\(K\) responsive channels based on peak spatial activations \(s_c = \max_{h,w} z_{c,h,w}\), activating selected channels across their spatial footprint. Guided by cross-attention target masks \(M_{\text{tgt}}\) and background complements \(M_{\text{bg}}\), PConvSAE is trained under a joint five-component loss: $\(\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{ctx}} + \mathcal{L}_{\text{cpt}} + \mathcal{L}_{\text{leak}} + \mathcal{L}_{\text{temp}} + \mathcal{L}_{\text{aux}}\)$ where \(\mathcal{L}_{\text{ctx}}\) forces the Context Branch to reconstruct solely background regions while being zeroed inside target boundaries, \(\mathcal{L}_{\text{cpt}}\) ensures additive recovery of target semantics, \(\mathcal{L}_{\text{leak}}\) penalizes concept kernel activations outside target regions to eliminate semantic leakage, \(\mathcal{L}_{\text{temp}}\) penalizes inter-frame fluctuations of context features to stabilize temporal dynamics, and \(\mathcal{L}_{\text{aux}}\) activates dormant latents via reconstruction residuals.
2. Contrastive Log-Ratio Attribution: One-vs-Max Hard-Negative Exclusion Dynamically searching for concept-specific features at inference time incurs prohibitive computational overhead and risks confounding target features with benign co-occurring semantics. EraseSAE introduces an offline contrastive attribution mechanism executed once before deployment. For a target concept \(c_i\), mean activations \(\mu_{c_i}\) are computed within target regions across positive samples. To eliminate spurious background correlations and semantically adjacent concept activations, a competitive hard-negative baseline is constructed: $\(\mu_{\text{base}} = \max\left(\mu_{\text{bg}},\, \max_{j \neq i} \mu_{c_j}\right)\)$ This establishes a rigorous one-versus-max exclusion principle: a kernel qualifies as concept-specific only if its response to \(c_i\) exceeds both general background activation and the maximum response triggered by any competing concept. Concept specificity is then quantified via contrastive log-ratio scoring: $\(S_{\text{attr}} = \mu_{c_i} \odot \max\left(0,\, \log\frac{\mu_{c_i} + \epsilon}{\mu_{\text{base}} + \epsilon}\right) \odot M_{\text{stable}}\)$ where multiplication by \(\mu_{c_i}\) prioritizes high-magnitude dominant features, and binary mask \(M_{\text{stable}}\) filters out transient features whose firing frequency across positive samples falls below threshold \(\tau\). Intersecting the top-ranked kernels with the pre-allocated concept subspace yields the definitively locked kernel set \(K_{\text{tgt}}\) with verified monosemantic purity.
3. Dynamic Spatiotemporal Mask SM-CFG: Timestep-Resolved Localized Suppression Because diffusion features progressively evolve from coarse global layouts to fine-grained textures across the reverse trajectory, applying a static spatial mask causes boundary tearing and semantic leakage as concept boundaries shift. EraseSAE adaptively generates a dynamic spatiotemporal mask at each denoising timestep \(t\). Feeding intermediate states into the frozen PConvSAE, the spatial complement of the Context Branch heatmap yields a dynamic foreground prior \(M_{\text{fg}}\), while the activation map of the locked concept kernels \(K_{\text{tgt}}\) provides the target concept region \(M_{\text{cpt}}\). The intervention mask is formed via normalized element-wise intersection: $\(M_t = \text{Norm}(M_{\text{fg}}) \odot \text{Norm}(M_{\text{cpt}})\)$ followed by dynamic threshold binarization. This mask constraints erasure strictly to the intersection of foreground entity and active concept pixels. Integrated into classifier-free guidance, the spatially-modulated CFG is formulated as: $\(\hat{\epsilon}_\theta(x_t) = \epsilon_\theta(x_t, c_{\text{pos}}) - s \cdot \left(\epsilon_\theta(x_t, c_{\text{pos}}) - \epsilon_\theta(x_t, c_{\text{neg}})\right) \odot M_t\)$ Within active concept boundaries (\(M_t \approx 1\)), strong negative guidance pushes the trajectory away from forbidden semantics; outside (\(M_t \approx 0\)), the original conditioned prediction remains entirely untouched, ensuring lossless background preservation.
Loss & Training¶
PConvSAE is trained on activations extracted from the probed intervention layer using the Adam optimizer with an initial learning rate of \(1 \times 10^{-4}\) for 30 epochs across 8 NVIDIA A100 GPUs. The diffusion backbone remains completely frozen throughout training. Since concept attribution is performed offline, inference requires only lightweight forward feature evaluation without backward propagation or dynamic search, preserving real-time generation speed.
Key Experimental Results¶
Main Results¶
EraseSAE was systematically evaluated on two primary DiT-based text-to-video architectures, CogVideoX-5B and HunyuanVideo, across nudity erasure and celebrity identity erasure tasks, as well as on the Flux.1 [dev] image foundation model.
The following table presents quantitative results for nudity erasure across standard prompts (Gen) and the adversarial Ring-A-Bell benchmark across difficulty tiers (K16, K38, K77):
| Method | Backbone | Nudity Rate (Gen) ↓ | Object Class (↑) | Subject Consistency (↑) | SSIM (↑) | Ring-A-Bell K16 ↓ | Ring-A-Bell K38 ↓ | Ring-A-Bell K77 ↓ | Latency (s/frame) |
|---|---|---|---|---|---|---|---|---|---|
| Original | CogVideoX | 28.25 | 79.91 | 95.12 | - | 14.47 | 20.26 | 29.34 | 3.61 |
| Neg Prompt | CogVideoX | 27.75 | 75.68 | 95.87 | 47.02 | 10.79 | 34.34 | 38.03 | 3.66 |
| SAFREE | CogVideoX | 5.75 | 45.19 | 94.49 | 42.39 | 15.13 | 16.58 | 15.66 | 3.72 |
| VideoEraser | CogVideoX | 18.88 | 73.86 | 96.47 | 44.36 | 7.11 | 18.03 | 23.68 | 4.03 |
| T2VUnlearning | CogVideoX | 2.88 | 51.98 | 93.55 | 42.39 | 5.30 | 7.60 | 8.65 | 3.67 |
| Ours (EraseSAE) | CogVideoX | 2.62 | 77.80 | 94.30 | 74.99 | 2.63 | 5.26 | 4.74 | 3.66 |
| Original | HunyuanVideo | 68.13 | 83.88 | 96.12 | - | 21.97 | 30.53 | 23.55 | 3.55 |
| Neg Prompt | HunyuanVideo | 57.25 | 87.59 | 95.48 | 47.06 | 22.37 | 29.34 | 31.97 | 6.95 |
| SAFREE | HunyuanVideo | 46.38 | 71.47 | 95.33 | 41.78 | 36.05 | 31.71 | 19.74 | 3.59 |
| T2VUnlearning | HunyuanVideo | 10.89 | 77.33 | 95.96 | 60.67 | 5.13 | 7.82 | 9.95 | 3.60 |
| Ours (EraseSAE) | HunyuanVideo | 7.13 | 80.61 | 96.04 | 65.69 | 4.47 | 6.84 | 5.13 | 6.97 |
In celebrity erasure across five public figures (Donald Trump, Barack Obama, Elon Musk, Queen Elizabeth, Taylor Swift), EraseSAE reduced the average target identification rate on CogVideoX to 8.00 (from 63.40 original) while preserving non-target identities at 58.60 and achieving a high background SSIM of 63.38. On HunyuanVideo, target detection dropped to 19.20 (from 67.10 original) with 61.50 non-target retention and an SSIM of 77.47, substantially outperforming baselines.
Ablation Study¶
Ablations on HunyuanVideo for the nudity erasure task validate the necessity of each component.
-
Cumulative Loss Function Ablation: | Config | \(\mathcal{L}_{\text{ctx}}\) | \(\mathcal{L}_{\text{cpt}}\) | \(\mathcal{L}_{\text{leak}}\) | \(\mathcal{L}_{\text{temp}}\) | Nudity Rate (Gen) ↓ | SSIM (↑) | Note | |---|---|---|---|---|---|---|---| | Context Recon Only | ✓ | | | | 24.50 | 40.12 | Incomplete concept-context isolation | | + Concept Recon Loss | ✓ | ✓ | | | 10.21 | 43.33 | Dual-branch functional specialization improves erasure | | + Leakage Penalty | ✓ | ✓ | ✓ | | 8.09 | 47.37 | Suppresses false activations, boosting kernel purity | | Full Model (+ Temporal Consistency) | ✓ | ✓ | ✓ | ✓ | 7.13 | 65.69 | Eliminates inter-frame jitter, surging SSIM by +18.32 |
-
Architectural Variants & Inference Strategies: | Dimension | Variant Method | Nudity Rate (Gen) ↓ | SSIM (↑) | Note | |---|---|---|---|---| | SAE Architecture | Linear SAE (Flattened) | 12.76 | 56.97 | Destroyed spatiotemporal topology causes feature blending | | | PConvSAE (Single-Branch) | 9.06 | 61.37 | Retains topology but lacks structural semantic decoupling | | | PConvSAE (Dual-Branch) | 7.13 | 65.69 | Optimal semantic disentanglement and structural preservation | | Inference Intervention | Feature Subtraction (SAE-Sub) | 51.30 | 62.31 | Fails to redirect reverse trajectory away from concept | | | Spatial Zeroing (SAE-Mask) | 6.56 | 53.51 | Strong suppression but severely corrupts scene context | | | Spatially-Modulated CFG (SM-CFG) | 7.13 | 65.69 | Localized negative guidance achieves surgical erasure |
Key Findings¶
- Temporal consistency is indispensable for video autoencoders: Incorporating \(\mathcal{L}_{\text{temp}}\) dramatically improves video structural similarity, raising SSIM from 47.37 to 65.69 (+18.32 points), demonstrating that smoothing context features across consecutive frames is crucial for preventing temporal flicker.
- Superior monosemantic purity: In feature purity evaluations, PConvSAE concept kernels achieve an average purity of \(\bar{P} = 0.76\), more than double that of a Linear SAE (0.37). The cross-identity activation matrix demonstrates sharp diagonal dominance, validating channel partitioning and one-vs-max attribution.
- Robust defense against adversarial attacks: While surface-level filtering baselines fail on Ring-A-Bell adversarial prompts (nudity rates exceeding 30%), EraseSAE reliably suppresses unsafe generations below 5-7%, proving that intervening on monosemantic latent units neutralizes deep semantic representations regardless of surface phrasing.
Highlights & Insights¶
- Repurposing mechanistic interpretability for actionable generative control: While SAEs are predominantly deployed for analytical probing in LLMs, EraseSAE pioneers their application as actionable surgical scalpels in complex video diffusion transformers, bridging interpretability and safe content generation.
- Principled one-versus-max contrastive attribution: The contrastive log-ratio formulation \(\mu_{\text{base}} = \max(\mu_{\text{bg}}, \max_{j \neq i} \mu_{c_j})\) mathematically eliminates spurious correlations, endowing the learned feature dictionary with native mutual exclusivity.
- Seamless integration of dynamic masks with spatially-modulated CFG: By dynamically updating spatiotemporal intervention boundaries at each timestep, the method avoids edge tearing while introducing negligible computational latency.
Limitations & Future Work¶
- Author-admitted limitations: In highly cluttered scenes with severe physical overlap or intimate interaction between the target entity and retained foreground subjects, decomposing foreground priors and concept heatmaps becomes challenging, occasionally leading to slight boundary artifacts.
- Future directions: The current implementation operates at an empirically identified single transformer layer; future extensions could explore multi-scale hierarchical SAE intervention across multiple network depths, as well as fine-grained editing of complex physical interactions and video motion dynamics.
Related Work & Insights¶
- vs SAFREE / VideoEraser: These training-free approaches rely on input prompt modification or text embedding projection, leaving internal latent representations vulnerable to adversarial circumvention; EraseSAE intervenes directly on intermediate monosemantic features, achieving vastly superior adversarial robustness.
- vs T2VUnlearning: T2VUnlearning fine-tunes backbone parameters globally, leading to catastrophic degradation of entangled benign concepts; EraseSAE leaves backbone weights intact and employs localized spatial CFG guidance, boosting SSIM by 5 to 28 points while achieving lower residual concept exposure.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ (First framework to introduce sparse autoencoders for surgical concept erasure in text-to-video diffusion models with an elegant decompose-attribute-erase pipeline)
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ (Extensive evaluation across CogVideoX-5B, HunyuanVideo, and Flux.1, spanning nudity, celebrities, adversarial robustness, and in-depth ablations)
- Writing Quality: ⭐⭐⭐⭐⭐ (Clear mathematical formulation, compelling motivation, and rigorous empirical validation)
- Value: ⭐⭐⭐⭐⭐ (Provides a foundational methodology for safety alignment, copyright protection, and mechanistic interpretability in generative video AI)