Skip to content

Open-Vocabulary Domain Unlearning

Conference: NeurIPS2026 (acceptance information supplied in the task metadata)
arXiv: 2609.31356
Area: Knowledge Editing / Multimodal VLM
Keywords: open-vocabulary domain unlearning, Fisher mask, targeted manifold scattering, few-shot learning, zero-shot retention

TL;DR

The paper extends visual-domain unlearning from failure on training classes to failure on classes absent from unlearning training, combining Fisher-masked visual-parameter updates with Targeted Manifold Scattering (TMS) to improve the forgetting–retention trade-off in few-shot CLIP classification, without establishing certified data deletion or guaranteed semantic preservation.

Background & Motivation

Vision-language models (VLMs) support zero-shot recognition using textual class descriptions, but object semantics and image style are not inherently separated. Domain unlearning seeks to reduce recognition of one visual style while preserving recognition of the same objects in other styles. Unlearning cartoons, for example, does not mean deleting the class “car”: the intended outcome is incorrect classification of cartoon cars while real cars remain correctly classified. Autonomous driving and medical diagrams motivate the paper, but the actual evaluation uses general image-classification datasets and does not establish clinical or driving-safety benefits.

Prior Approximate Domain Unlearning (ADU) uses the same classes for unlearning training and evaluation. This can damage specific class–style combinations without producing class-independent domain forgetting. The paper holds out some object classes and asks whether they remain recognizable in retain domains while also suffering recognition degradation in the forget domain. The difficulty is that individual visual parameters can support both style and semantics: increasing forget-domain classification loss can damage retention, while prompt tuning with Maximum Mean Discrepancy (MMD) may produce a global feature displacement biased toward seen classes.

The paper separates which weights to update from where forget representations should move. Differences in Fisher sensitivity between domains address the former; a geometric objective with explicit destinations addresses the latter, rather than merely increasing classification errors. Core idea: within a restricted set of visual parameters, make retain samples prefer same-class retain samples and move forget samples toward cross-class retain samples, seeking localized domain forgetting that transfers across classes.

Method

Overall Architecture

The inputs are pretrained OpenCLIP, a small set of images with class and domain labels, and designated forget and retain domains. Before training, domain-specific parameter sensitivity produces a fixed Fisher mask. During training, masked updates jointly optimize retain classification, an inverted forget-classification objective, and the two geometric TMS constraints. The output remains a visual–text classification model, with no additional inference-time domain detector.

The open-vocabulary protocol partitions classes into mutually exclusive seen and unseen sets. The seen-class fraction is \(\gamma\in\{0.25,0.50,0.75,1.0\}\), and unlearning training accesses only seen classes. Evaluation covers the union of both sets—all classes—not just unseen classes. At \(\gamma=1.0\), no classes are held out, making this a closed-vocabulary control. Because the main table aggregates seen and unseen classes, higher overall HM cannot substitute for a separate unseen-class forgetting result.

The Fisher mask, Retention Preference, and Targeted Confusion are the three contribution components. The latter two act together within one training objective; their vertical ordering in the diagram is explanatory, not sequential training of two networks. Solid edges indicate training constraints, and dashed edges indicate inference data flow after updating the model.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Seen-class images<br/>Retain + forget domains"] --> B["Fisher Mask"]
    A --> C["Retention Preference"]
    A --> D["Targeted Confusion"]
    B -->|Parameter-update gate| E["Joint objective<br/>Visual-weight updates only"]
    C -->|TMS retention term| E
    D -->|TMS confusion term| E
    E --> F["Updated OpenCLIP"]
    G["Test images + class text"] -.-> F
    F -.-> H["Classification over all classes"]

Key Designs

1. Fisher Mask: concentrate updates on visual weights more sensitive to the forget domain

The paper first establishes a NegGrad+Fisher baseline. Ordinary NegGrad minimizes retain-domain cross-entropy and maximizes forget-domain cross-entropy, giving the gradient difference in source equation (1):

\[ g_{i}^{\mathrm{NegGrad}}=\nabla_{\theta_{i}}\mathcal{L}_{\mathrm{CE}}(x_{r},y_{r};\theta)-\lambda\nabla_{\theta_{i}}\mathcal{L}_{\mathrm{CE}}(x_{f},y_{f};\theta) \]

Reversing a gradient does not distinguish semantic weights from stylistic weights. The authors therefore estimate diagonal empirical Fisher information separately for the two domain groups, averaging squared per-sample loss gradients rather than taking their maximum, to reduce the influence of outliers on importance estimates. Source equation (2) is:

\[ F_{\mathcal{D}}(i)=\frac{1}{N}\sum_{j=1}^{N}\left(\nabla_{\theta_{i}}\ell_{j}(\theta)\right)^{2} \]

They then compute the normalized sensitivity difference in source equation (3), rather than treating parameters important to both domains as domain-specific:

\[ \rho_{i}=\frac{F_{\mathcal{D}_{f}}(i)-F_{\mathcal{D}_{r}}(i)}{F_{\mathcal{D}_{f}}(i)+F_{\mathcal{D}_{r}}(i)+\epsilon} \]

Only parameters with \(\rho_i>\tau\) may update; other gradients are zeroed. The main text uses \(\tau=0.3\), so this is sensitivity-threshold selection, not selection of a fixed percentage of weights. The mask is computed once before unlearning and held fixed, while the text encoder remains frozen. Given the experimental setup, Baseline in the main table should be understood as NegGrad+Fisher, not the unedited base model.

The rationale is that the two domains share class semantics, so differential sensitivity may locate updates more selectively than single-domain gradient magnitude. However, activation in both domains does not imply that semantic parameters must satisfy \(\rho_i\leq0\), and style and semantics need not be coordinate-separable. The source equates masking with projection into a semantic nullspace and claims mathematically guaranteed zero-shot preservation. Its diagonal gradient statistics and binary selection do not establish that claim. This note treats masking as an empirical interference constraint, not an orthogonal projection or preservation theorem.

2. Retention Preference: favor same-class retain samples over same-class forget samples

Making forget samples misclassify does not ensure that retain clusters remain well structured. The first TMS term anchors on a retain embedding, takes another same-class retain embedding as the positive, and treats a same-class forget embedding as the less desirable neighbor. Comparing their cosine similarities encourages the model to maintain semantic proximity within retain domains.

Source equation (5) is shown below. Visual embeddings are L2-normalized, \(s\) denotes cosine similarity, \(\sigma\) is sigmoid, \(\beta\) is a scaling coefficient, and \(m_{\mathrm{retain}}\) is the preference margin:

\[ \mathcal{L}_{\mathrm{retain}}=\mathbb{E}_{(z_{r},z_{f})\mid y_{r}=y_{f}}\left[-\log\sigma\left(\beta\left(s(z_{r},z_{\mathrm{r\_pos}})-s(z_{r},z_{f})-m_{\mathrm{retain}}\right)\right)\right] \]

This constrains relative preferences rather than freezing retain embeddings. It supports same-class retain cohesion while reducing intrusion by same-class forget samples. The loss operates in the visual encoder's pre-projection CLS-token space, whereas classification cross-entropy continues to constrain final visual–text classification. The geometric and classification terms are therefore not identical constraints.

3. Targeted Confusion: give forget embeddings a local cross-class destination

With sparse editable weights and few seen classes, NegGrad may merely move samples across decision boundaries while leaving stylistic structure largely intact. TMS instead selects the currently most similar cross-class retain embedding for each forget embedding and requires their similarity to exceed a confusion margin. Here, hard mining takes the maximum similarity among cross-class candidates—the easiest local confusion target—not the most distant negative.

Source equation (6) is:

\[ \mathcal{L}_{\mathrm{confuse}}=\mathbb{E}_{z_{f}}\left[-\log\sigma\left(\beta\left(\max_{z_{r}\mid y_{r}\neq y_{f}}s(z_{f},z_{r})-m_{\mathrm{confuse}}\right)\right)\right] \]

Different forget samples can approach different cross-class regions, motivating dispersed local geometric changes rather than an MMD-style global translation. Log-sigmoid provides a continuous soft penalty on the preference gap, unlike a hard-margin hinge loss. However, switches in the maximizing candidate can remain nondifferentiable, so the full hard-mining objective is not necessarily everywhere smooth. Increasing similarity to an incorrect class is also a measurable representation perturbation, not automatic proof that stylistic information has disappeared from the model.

Loss & Training

Retention Preference and Targeted Confusion are combined into TMS using source equation (7):

\[ \mathcal{L}_{\mathrm{TMS}}=w_{\mathrm{retain}}\mathcal{L}_{\mathrm{retain}}+w_{\mathrm{confuse}}\mathcal{L}_{\mathrm{confuse}} \]

Retain cross-entropy and negative forget cross-entropy are then added, as in source equation (8):

\[ \mathcal{L}_{\mathrm{total}}=(1-\lambda_{\mathrm{ce}})\mathcal{L}_{\mathrm{retain\_CE}}-\lambda_{\mathrm{ce}}\mathcal{L}_{\mathrm{forget\_CE}}+\lambda_{\mathrm{pref}}\mathcal{L}_{\mathrm{TMS}} \]

The negative sign increases forget-domain classification loss during optimization, rather than training correct classification in that domain. The implementation description additionally states that the two cross-entropy terms are dynamically scaled by retain and forget sample proportions. It does not fully explain their relationship to \(\lambda_{\mathrm{ce}}\) in equation (8), so this note does not combine them into an equation absent from the source.

The main experiments use OpenCLIP ViT-B/16, 8-shot training, AdamW, a learning rate of \(1\times10^{-5}\), and 10 epochs. The appendix adds batch size 64, weight decay \(1\times10^{-2}\), and an A100 40GB GPU. Each configuration uses a fresh optimizer and the same pretrained initialization. Fisher information is estimated using 10 minibatches; measured runtime is not reported in the main text, and the appendix's efficiency discussion is primarily qualitative.

The appendix sets \(\beta=0.1\), \(m_{\mathrm{retain}}=0\), and \(w_{\mathrm{retain}}=w_{\mathrm{confuse}}=1\). Grid-search ranges are \(\lambda_{pref}\in\{0.5,1.0,2.0,5.0\}\) and \(m_{confuse}\in\{0.3,0.5,0.75\}\), without the final selected values for each setting. It also describes a validation-based rank-matched mask for multi-domain unlearning: average the per-domain mask sizes, recompute sensitivity on mixed forget data, and select top-k parameters. This should not be conflated with the fixed-threshold configuration in the main text.

Key Experimental Results

Main Results

Each classification benchmark contains four domains: OfficeHome has 65 classes, Mini-DomainNet has 126, and PACS has 7. The main table abbreviates Mini-DomainNet as DomainNet. One domain is designated for forgetting and the others for retention. The table below preserves the source's F / R / HM numerical order; values are percentages, and evaluation covers all classes.

F is forget-domain classification accuracy, where lower is better; R is retain-domain classification accuracy, where higher is better. HM is the harmonic mean of R and the erasure rate, not of R and F:

\[ \mathcal{E}=100-\mathcal{A}_{f},\qquad \mathcal{H}=\frac{2\cdot\mathcal{A}_{r}\cdot\mathcal{E}}{\mathcal{A}_{r}+\mathcal{E}} \]
Dataset Method 0.25: F / R / HM 0.50: F / R / HM 0.75: F / R / HM 1.0: F / R / HM
OfficeHome Baseline 59.54 / 65.65 / 50.07 51.62 / 67.92 / 56.51 36.97 / 71.90 / 67.17 31.31 / 77.90 / 73.01
OfficeHome ADU 21.97 / 27.61 / 40.79 28.26 / 49.22 / 58.38 33.63 / 67.96 / 67.16 36.37 / 80.36 / 71.02
OfficeHome TMS 53.33 / 63.24 / 53.71 39.34 / 68.33 / 64.27 24.90 / 68.68 / 71.75 24.61 / 75.45 / 75.42
DomainNet Baseline 55.58 / 77.72 / 56.53 27.98 / 77.08 / 74.46 20.04 / 77.24 / 78.58 18.02 / 77.77 / 79.82
DomainNet ADU 14.64 / 24.45 / 38.01 19.68 / 49.25 / 61.06 23.41 / 67.27 / 71.63 27.86 / 81.48 / 76.53
DomainNet TMS 34.04 / 74.97 / 70.18 20.06 / 75.39 / 77.60 14.60 / 75.97 / 80.41 14.44 / 76.84 / 80.97
PACS Baseline 78.64 / 90.02 / 34.52 57.75 / 93.76 / 58.25 59.40 / 93.42 / 56.60 41.01 / 95.60 / 72.96
PACS ADU 64.97 / 28.65 / 31.52 64.07 / 59.53 / 44.81 68.36 / 74.40 / 44.40 72.48 / 93.16 / 42.49
PACS TMS 76.085 / 90.39 / 37.82 47.06 / 93.49 / 67.60 47.07 / 93.24 / 67.53 38.35 / 95.64 / 74.97

These values come from source Table 1; the original precision of 76.085 is retained. The paper states that results average multiple independent runs but does not specify the run count or confidence intervals. Aggregated HM need not equal HM recomputed from aggregated F and R, so the reported values are preserved rather than “corrected.”

Ablation Study

The following subset of Table 2 includes overall averages and local counterexamples, distinguishing best average performance from best performance in every domain. The configurations keep only TMS confusion, only TMS retention, or both; they should not be interpreted as removing the base cross-entropy terms or Fisher mask.

Forget domain / aggregate Class fraction Confuse Only: F / R / HM Retain Only: F / R / HM Both: F / R / HM
Clipart 1.00 16.03 / 78.10 / 80.93 18.73 / 77.78 / 79.49 17.30 / 77.08 / 79.79
Real 0.50 31.27 / 71.48 / 70.08 38.57 / 72.86 / 66.66 32.38 / 72.70 / 70.06
Sketch 0.50 14.44 / 75.08 / 79.98 13.49 / 76.98 / 81.47 12.53 / 75.23 / 80.88
DomainNet (Avg) 0.50 21.67 / 75.49 / 76.81 24.68 / 77.06 / 75.97 20.06 / 75.39 / 77.60
DomainNet (Avg) 1.00 17.30 / 77.53 / 79.97 17.90 / 77.59 / 79.69 14.44 / 76.84 / 80.97

Combining both terms gives the best overall average HM, but a single term performs better for Clipart at 1.00, Real at 0.50, and Sketch at 0.50. Thus, claiming both components are strictly necessary in every setting exceeds the evidence in Table 2.

To examine semantic retention, Table 4 additionally reports external zero-shot classification after unlearning individual DomainNet domains. Its caption does not specify the class fraction, so it should not be assigned an assumed open-vocabulary configuration.

Forget domain Method ImageNet-val (%) CIFAR-100 (%)
No unlearning Base Model 63.38 61.99
Clipart Baseline 54.11 57.40
Clipart ADU 36.73 49.37
Clipart TMS 57.33 57.50
Painting Baseline 51.76 50.45
Painting ADU 28.45 32.54
Painting TMS 49.30 52.26
Sketch Baseline 46.87 52.49
Sketch ADU 39.42 57.01
Sketch TMS 48.82 56.21

Key Findings

  • In Table 1, TMS achieves the highest HM for all three datasets and four fractions, but not always the lowest F or highest R. For DomainNet at 0.25, ADU forgets more strongly but reduces R to 24.45; TMS retains R of 74.97 and reaches HM of 70.18.
  • Table 3 forgets three domains. TMS has HM of 73.05, 75.34, and 78.02 at the first three fractions, outperforming the alternatives. At 1.0, ADU's 80.02 exceeds TMS's 78.86, so TMS is not best in every multi-domain setting.
  • The text accompanying Figure 3 reports 4-shot HM of 80.80 at 1.0, versus 8-shot Baseline 79.76 and ADU 76.53; at 0.5 it reports 76.79, versus 74.09 and 61.06. The two Baseline values differ from Table 1's 79.82 and 74.46 without explanation. Both sources are retained rather than merged.
  • Table 4 supports improved retention relative to ADU, but all TMS external accuracies remain below the base model. Painting ImageNet accuracy is 49.30, below Baseline 51.76; Sketch CIFAR-100 accuracy is 56.21, below ADU 57.01. The text's claim of best performance across all domains is not supported.
  • Appendix Tables 5–7 add ViT-L-14 results. For Real at 8-shot and 0.5 in Table 5, Baseline F / R / HM is 0.8651 / 0.8386 / 0.2324, versus TMS 0.2937 / 0.8069 / 0.7533. These are 0–1 fractions, not the percentages used in the main tables.

Highlights & Insights

  • Separating evaluation classes from unlearning-training classes is stricter than evaluating only seen classes. It exposes the distinction between damaging training classes and transferring stylistic suppression, although separate unseen-class results are still needed.
  • Differential sensitivity selection and geometric objectives address different problems: the former restricts where interference can occur, while the latter specifies a local perturbation direction. Their combination offers a clearer account of structural change under few-shot training than increasing classification loss alone.
  • Retention and confusion are not mirror-image objectives. One anchors on retain samples to preserve same-class relations; the other anchors on forget samples to seek cross-class destinations, allowing localized forgetting without requiring global domain translation.

Limitations & Future Work

  • The authors acknowledge that diagonal Fisher ignores parameter dependencies and suggest richer sensitivity structures. The current binary mask does not prove semantic-nullspace projection or zero-shot preservation.
  • Freezing the text encoder limits the intervention, and the authors note potential cross-modal residual information. Lower classification accuracy establishes impaired recognition under the evaluation, not certified removal of training-data influence.
  • The main tables aggregate all classes without seen/unseen breakdowns, run counts, error bars, or an explicit validation-class split. Future evaluation should separately measure forgetting transfer and retention damage while excluding unseen classes from hyperparameter tuning.
  • There is no mask-removal ablation, threshold sweep, or direct geometric measurement sufficient to isolate Fisher's claimed semantic-protection effect. Matched-budget comparisons with feature-subspace constraints would be informative, without assuming these mechanisms are equivalent.
  • Claims about 4-shot advantages, best external-test performance, and universally optimal multi-domain results should be narrowed to the table evidence. Generic style-classification results also cannot replace deployment-specific risk assessment.
  • vs ADU: ADU uses prompt tuning and MMD for domain-distribution separation; this paper uses restricted visual-weight updates and sample-level geometric objectives, emphasizing cross-class evaluation. Local geometric changes are more targeted than global shifts, but these comparisons do not establish that all MMD or prompt-tuning methods must fail.
  • vs NegGrad+Fisher: Both use the same type of parameter gating; TMS adds retention preferences and cross-class confusion, with improved average trade-offs in the main table. The gain should not be reduced to “first using Fisher.”
  • vs DPO / NPO: The paper borrows a log-sigmoid preference form but optimizes visual-similarity differences rather than language-generation sequence probabilities. It does not use the original DPO reference-policy probability ratio, so the mechanisms are not identical.
  • vs SEMU: The appendix argues that highly overlapping domain subspaces hinder direct SVD-style isolation. Figures 4–5 provide qualitative supporting material, but the cache contains no numerical SEMU table suitable for citation. A research direction is to use independent unseen-class evaluation to test whether sensitivity selection actually reduces semantic interference rather than merely increasing aggregated HM.

Rating

  • Novelty: 4/5 — The combination of cross-class domain-unlearning evaluation and localized geometric objectives has clear value.
  • Experimental Thoroughness: 3/5 — Three benchmarks, ablations, multiple-domain experiments, and a larger backbone are included, but unseen-class breakdowns and statistical reporting are insufficient.
  • Writing Quality: 3/5 — The mechanism is fairly clear, but mathematical guarantees and several optimality claims exceed the evidence.
  • Value: 4/5 — A stricter perspective on domain-unlearning evaluation, requiring independent validation and retention audits before practical use.