Contrastive Representation Shaping for LLM Unlearning¶
Conference: NeurIPS2026
arXiv: 2601.22028
Code: https://github.com/HaoranTang/CLReg
Area: LLM Safety
Keywords: machine unlearning, contrastive learning, representation shaping, forget–retain entanglement, privacy evaluation
TL;DR¶
CLReg uses augmented views of the same forget example as positives and retain examples as negatives to shape hidden representations alongside a base unlearning objective, improving aggregate scores in multiple settings without equating representation separation with knowledge removal or providing privacy guarantees.
Background & Motivation¶
Machine unlearning for large language models (LLMs) aims to revoke the influence of selected training data while preserving other knowledge. Retraining only on retained data is the ideal reference but is usually expensive; directly increasing loss on forget data can damage general capabilities. NPO, SimNPO, and UnDIAL therefore mainly constrain output probabilities or logits, reducing target-content generation while avoiding the instability of gradient ascent. Yet failing to answer is not equivalent to eliminating the underlying knowledge. If forget and retain knowledge still share representations, subsequent learning may reverse output suppression, whereas strengthening suppression can damage retained capabilities.
The specific bottleneck examined here is forget–retain representation entanglement. Earlier contrastive unlearning methods mostly target small classifiers with class labels that naturally define positive and negative pairs. Generative LLM unlearning requests often concern an example or a semantic concept without reliable supervised classes. Rather than forcing all hidden representations back toward a retrained model or sending forget activations in a random direction, the authors use augmented views of each forget input to identify representations that should remain close, and retain examples to identify directions to repel. The representation space can then change while the base training objective protects retained behavior.
This change creates conditions that make selective modification easier for the base algorithm; it does not independently perform deletion. Core Idea: align different views of each forget example, separate them from retain representations, and add this directional representation shaping as a regularizer to existing unlearning losses.
Method¶
Overall Architecture¶
The inputs are a fine-tuned model, a forget set, and a retain set; the output is the same model after unlearning updates. Training proceeds through “Same-Example Views,” “Directional Contrastive Separation,” and “Joint Unlearning Update,” combining representation regularization with base forgetting and retaining objectives. CLReg adds no trainable module, and inference requires neither contrastive retrieval nor a negative-example branch.
Token hidden states from a selected Transformer layer are mean-pooled using the attention mask and normalized to unit length, yielding sentence or passage representations for cosine comparison. The paper calls the subspace spanned by these representations of forget inputs the forget concept. This is an operational representation-level definition, not evidence that all relevant parameters have been identified or that an independently removable subnetwork has been found.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
F["Forget inputs"] --> V["Same-Example Views"]
V --> C["Directional<br/>Contrastive Separation"]
R["Retain inputs"] -->|Negative embeddings| C
C -.->|Training regularizer| J["Joint Unlearning<br/>Update"]
F -->|Base forget loss| J
R -->|Base retain loss| J
J --> M["Updated model"]
M -->|Model only at inference| O["Ordinary text output"]
The two views and negative examples in the diagram provide training signals only. The base forget loss still carries the removal objective; the representation-repulsion branch is not an inference-time knowledge filter.
Key Designs¶
1. Same-Example Views: define positives from the input itself without class labels
Each forget example is an anchor, and its positive comes from a paraphrase of that same input or independent dropout. Before TOFU training, Llama-3.1-8B-Instruct generates five paraphrases per forget question, modifying only the question and keeping the answer unchanged. One paraphrase is sampled uniformly at each step, and independent dropout is applied to anchor and positive hidden states. Paraphrases must preserve names, numbers, and facts: the aim is to change wording, not the forgetting target. MUSE provides no paraphrases, so the same passage under independent dropout masks supplies the two views.
Dropout probabilities are sampled from a normal distribution with mean 0.1 and standard deviation 0.05, then clamped to 0–0.2. The contrastive loss compares pooled, normalized representations rather than raw token probabilities; retain negatives do not receive this additional dropout. Positive relationships therefore require no manual labels and do not assume that different forget examples have identical semantics.
This distinction matters: CLReg aligns each example with its own augmented view, rather than explicitly collapsing all forget inputs into one point. The paper describes the overall object as a forget-concept subspace, but training directly constrains local positive-pair relationships and cross forget–retain relationships. Augmentation quality consequently determines whether positives preserve the target semantics; very small forget sets or semantically distorted paraphrases can weaken this signal.
2. Directional Contrastive Separation: repel retain representations instead of randomly perturbing forget activations
Each forget anchor forms negative pairs with retain-input representations. By default, an anchor sees 32 negatives gathered across eight data-parallel workers with four examples per worker. The effective batch size is 128 for TOFU and 256 for MUSE, but gradient accumulation does not enlarge the negative pool of an individual contrastive computation. Effective batch size must therefore not be mistaken for the number of negatives.
In the paper’s DPO-style version, positive-pair similarity should exceed each negative-pair similarity:
Here \(B_f\) and \(B_r\) are the numbers of forget anchors and retain negatives, \(s\) is cosine similarity between normalized representations, \(\sigma\) is the sigmoid, and \(\tau\) is temperature. When an anchor is too similar to a retain negative, the positive–negative similarity margin is small and the term produces a stronger update. Already-separated negatives contribute less. The loss thus automatically gives harder negatives more weight without necessarily requiring a separate mining strategy.
“DPO-style” refers to logistic pairwise ranking of similarity differences, not standard DPO on preferred responses and policy probabilities. This regularizer introduces neither a reward model nor a reference policy. The authors also test InfoNCE, treating similarities to one positive and all negatives as classification logits with the positive as the correct item. Both forms can exchange forget and retain anchor roles to create symmetric variants. Different base methods and models favor different variants; no combination is universally best.
A single-anchor analysis explains the repulsion direction, but its scope must remain explicit. In embedding space with fixed positive and negative vectors, one direct anchor-gradient update moves toward the positive and away from the negative. With multiple negatives, the repulsion target becomes a difficulty-weighted negative centroid. Decreasing similarity to that centroid requires an additional geometric condition, observed in 94% of logged anchor updates rather than guaranteed automatically.
Actual training updates shared network parameters, changing anchors, positives, and retain representations together, with normalization also involved afterward. The single-step derivation therefore does not establish sustained full-network separation, undistorted retain representations, or certified knowledge removal. Representation distances, entanglement measurements, and intervention experiments—not this local derivation—provide the main evidence for practical separation.
3. Joint Unlearning Update: representation shaping must support the base forgetting objective
CLReg integrates with GradDiff, NPO, SimNPO, UnDIAL, or PDU by adding a loss term, not by replacing their respective forgetting mechanisms:
The base forget loss changes behavior on forget data, the retain loss protects retained tasks, and CLReg changes the relative geometry of the two representation distributions. Moving forget representations elsewhere can leave knowledge intact while merely making it more distinguishable. Representation separation therefore cannot substitute for behavioral, privacy, and recovery evaluations. The conclusion likewise describes CLReg as a regularizer that strengthens the base removal process, not an independent removal mechanism.
The authors test this division of labor directly: first train for five epochs using the retain objective and CLReg without a base forget loss, then remove CLReg and train the original base method for ten epochs. Better subsequent unlearning would not be attributable solely to the extra signal from optimizing both losses simultaneously. Against a fifteen-epoch base-method control with matched total epochs, the representation-shaped initialization still improves subsequent scores. This control matches epochs, however, not strictly wall-clock time or FLOPs, because CLReg can itself add forward computation.
The current experiments mainly target high-level semantic knowledge and apply CLReg to the final feature layer by default. Earlier representations typically contain more shared information, making repulsion more likely to affect retained capabilities. “Final layer” identifies where the regularizer reads hidden representations; it does not mean that only this layer’s parameters are updated, nor is it a universal localization rule for every forget concept.
Loss & Training¶
Main experiments use ten epochs, a learning rate of \(10^{-5}\), and H100 hardware. The authors tune base-algorithm parameters first, then search contrastive temperature, symmetry, and DPO-style versus InfoNCE objectives. Temperature candidates are 0.1, 0.3, 0.5, 0.7, and 0.9. Main tables use \(\lambda=1\) by default, but retaining weights are not identical for every method: the UnDIAL baseline has no retain loss and receives a small retaining weight with CLReg, while WMDP NPO also rebalances this weight.
Offline paraphrasing is neither performed at inference nor repeated each epoch. TOFU adds a forward pass for the paraphrased input; MUSE constructs views using dropout alone and avoids this extra paraphrase forward pass. Most TOFU methods incur 34%–49% measured wall-clock overhead, but UnDIAL has higher overhead because its implementation does not cache hidden states. That interval is not an upper bound for every method.
Key Experimental Results¶
Main Results¶
TOFU contains fictional-author question–answer pairs. The setup specifies Llama-3.1-8B and Llama-3.2-3B, abbreviated as Llama-3-8B/3B in the tables; MUSE-Books/News use Llama-2-7B. The following selected pairs illustrate gains and costs without equating an aggregate-score improvement with improvement in every raw metric.
ForgetScore converts each forget metric into progress relative to the original fine-tuned and retrained models, then takes a harmonic mean. ForgetQuality is log-transformed because it spans many orders of magnitude. The core definition is:
This expression explicitly incorporates the progress clipping described in the text into the harmonic mean. Absolute values measure how far behavior moves from the original model, without encoding direction or guaranteeing proximity to the target. Clipping after surpassing the target can also conceal over-unlearning. ForgetQuality itself is the p-value of a KS test comparing forget-answer TruthRatio distributions against the retrained reference, not a synonym for ForgetScore.
The source contains an inconsistency between the definition and values of UnlearningScore. Section 4.2 describes it as the harmonic mean of ForgetQuality and ModelUtility, but TOFU table values match the harmonic mean of ForgetScore and ModelUtility: for example, SimNPO+CL combines 0.97182 and 0.69815 into 0.81256, while its ForgetQuality is 0.00229. The reported scores are retained below with this discrepancy explicitly noted, rather than silently repairing the authors’ definition.
| Dataset and method | UnlearningScore: base → +CL | Retain metric: base → +CL | PrivLeak: base → +CL |
|---|---|---|---|
| TOFU-8B, SimNPO | 0.77378 → 0.81256 | ModelUtility 0.67261 → 0.69815 | -48.65591 → 55.03337 |
| TOFU-3B, SimNPO | 0.71348 → 0.78664 | ModelUtility 0.61040 → 0.66531 | -61.99405 → -3.72569 |
| MUSE-Books, GradDiff | 0.41148 → 0.73778 | RetainKnowMem ROUGE 0.59876 → 0.58451 | -60.46598 → -57.32249 |
| MUSE-Books, SimNPO | 0.73961 → 0.74936 | RetainKnowMem ROUGE 0.60877 → 0.59918 | -30.65828 → -21.11686 |
| MUSE-News, SimNPO | 0.46128 → 0.57173 | RetainKnowMem ROUGE 0.47777 → 0.42403 | -90.55416 → 88.65659 |
Source: Tables 1–2. PrivLeak is a membership-inference AUC difference calibrated relative to the retrained reference; it should approach zero, not increase indiscriminately. Positive values indicate over-unlearning and negative values under-unlearning. For TOFU-8B SimNPO, CLReg increases the absolute value, ruling out a claim of universal privacy improvement. MUSE-News SimNPO slightly reduces its absolute value but remains far from zero and shifts into over-unlearning.
WMDP-cyber is discussed only as a high-level safety-unlearning evaluation: on Zephyr-7B-β, SimNPO’s aggregate score changes from 0.633 to 0.674, versus RMU’s published 0.677. This experiment substitutes four-option random accuracy, 0.25, for an unavailable retrained reference and does not report PrivLeak. The proxy comparison does not establish equivalence to genuine retraining.
Ablation Study¶
The first analysis table selects results from Tables 3–4. Entanglement is the ratio of within-group dispersion to separation of the two group means, with lower values preferred. MK-MMD measures multi-kernel distribution discrepancy and 2-Wasserstein measures distribution distance; higher values indicate greater separation for both. These geometric metrics do not directly measure removal guarantees.
| Dataset and model, all SimNPO | Entanglement: base → +CL | MK-MMD: base → +CL | 2-Wasserstein: base → +CL |
|---|---|---|---|
| TOFU-8B | 20.24046 → 5.90206 | 0.01537 → 0.07327 | 0.10001 → 0.31250 |
| TOFU-3B | 17.19105 → 5.62938 | 0.01833 → 0.06930 | 0.06768 → 0.23700 |
| MUSE-News | 92.17580 → 11.08532 | 0.00392 → 0.01775 | 0.04923 → 0.11468 |
The second analysis table summarizes mechanism tests from Tables 5, 10, and 14. These analyses rerun public checkpoints, so absolute scores must not be mixed with main-table scores. Rows from different experiments also do not form one common search space.
| Analysis setting | Config | Reported metric | Interpretation |
|---|---|---|---|
| TOFU-3B, SimNPO two-stage | Fifteen-epoch base method | UnlearningScore 0.734 | Matched-total-epoch control |
| Same setting | Five shaping epochs, then ten base-method epochs | UnlearningScore 0.771 | No CLReg in stage two |
| TOFU-3B, SimNPO negative count | No CLReg | UnlearningScore 0.724 | Matched baseline for this ablation |
| Same setting | 1 / 32 negatives | UnlearningScore 0.803 / 0.807 | Many negatives are not required within this range |
| TOFU-3B, SimNPO three-seed recovery evaluation | Base / +CL | recovery AUC \(0.329\pm0.021\) / \(0.228\pm0.012\) | Mean and standard deviation; lower means slower recovery |
| TOFU-3B, output-quality path | \(\lambda=0\) | UnlearningScore 0.688; P(clean) 0.868 | No regularizer |
| Same setting | \(\lambda=0.3\) | UnlearningScore 0.766; P(clean) 0.887 | Score and coherence both improve |
| Same setting | \(\lambda=1\) | UnlearningScore 0.785; P(clean) 0.585 | Higher score accompanies substantial degeneration |
Recovery AUC normalizes forget-answer probability recovery between the current unlearned checkpoint and the original fine-tuned model, then integrates it trapezoidally over evaluation indices. It is distinct from membership-inference AUC. P(clean) is a gibberish classifier’s predicted probability for its clean-text class, not a human quality assessment.
Key Findings¶
- All five base methods improve their aggregate scores across five task settings in the main tables, but TOFU’s broadly improved retaining utility does not extend uniformly to MUSE: several MUSE pairs have lower retain ROUGE.
- Shaping representations before removing CLReg still helps, strengthening the case that separation facilitates subsequent unlearning beyond correlation in joint training. It does not establish separation as the only causal factor.
- The authors report that 14 of 16 combinations of contrastive form and symmetry outperform the baseline; not every configuration succeeds. Final-layer targets outperform earlier target layers in the layer study, rather than benefiting simply from regularizing more layers.
- Recovery evaluation supports slower recovery for SimNPO and GradDiff, but NPO’s recovery AUC changes from 0.068 to 0.077 and does not improve. The appendix states that recovery curves can eventually converge within 64 steps.
- Across four consecutive MUSE-News requests, retaining-knowledge erosion falls from 37.1% to 3.0%, while CLReg’s per-request ForgetScore weakens. General-capability differences in the appendix should likewise mainly be read as preservation, not significant capability improvement.
Highlights & Insights¶
- Representations need not imitate a retrained model point by point. A more useful objective may first reduce forget–retain interference and then let behavioral objectives perform selective modification.
- Positives constructed from different views of the same input bring contrastive shaping into generative models without class labels. Using only retain-side negatives turns generic dispersion into directional repulsion.
- Two-stage interventions, geometric metrics, and recovery tests jointly support the mechanistic account. They evaluate different questions, and none can replace the others.
Limitations & Future Work¶
- The one-step embedding analysis does not cover shared parameters, repeated updates, or general multi-negative settings. It establishes neither retain-distortion bounds nor certified removal or privacy guarantees.
- A high ForgetScore can result from incoherent output on forget queries. Every configuration in the output-quality path has a refusal rate of 0.000, so degeneration must not be presented as a safety refusal.
- Fixed regularization strength can cause over-unlearning. Main-table GradDiff PrivLeak on MUSE-News changes from -87.84635 to 99.68514, showing that aggregate-score gains can coexist with worse privacy calibration.
- Main tables largely report single runs, with three-seed analysis limited to a representative SimNPO pair. Negative-count robustness and layer-selection conclusions are also bounded by the tested tasks.
- Concept-dependent layer selection and retain-behavior-constrained subspace intervention after shaping are possible directions, not validated results. They require joint evaluation of knowledge recovery, output usability, and utility under repeated requests.
Related Work & Insights¶
- vs NPO / SimNPO: The base methods reduce forgotten content through output-level objectives, while CLReg adds geometric constraints between hidden representations. These are complementary; a similarity objective does not directly replace negative-preference unlearning.
- vs RMU: RMU sends forget activations toward a random direction, whereas CLReg constructs structured repulsion using same-example positives and retain negatives. The paper does not evaluate RMU+CL, so compatibility is not established.
- vs classifier contrastive unlearning: Classifiers can use labels to define positive and negative relationships. This work substitutes same-example augmentation for class supervision and evaluates generative behavior and retained utility.
- Research insight: Selecting regularization strength using aggregate score, PrivLeak, and output coherence together is more reliable than optimizing one score. The best strength for sequential requests may also differ from that for a single request.
Rating¶
- Novelty: 4/5 — Same-example contrastive shaping without labels is a clear LLM-unlearning regularizer, although contrastive learning itself is not new.
- Experimental Thoroughness: 4/5 — Multiple methods, representation analyses, and interventions are covered; main-table seeds and long request streams remain limited.
- Writing Quality: 4/5 — The method is understandable and failure cases are documented, but aggregate-metric definitions and some broad claims require careful checking.
- Value: 4/5 — A reusable approach to reducing unlearning collateral damage, not evidence of compliant data deletion.