UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?¶
Conference: NeurIPS2026
arXiv: 2609.37076
Area: LLM Safety
Keywords: machine unlearning, weight interpolation, model merging, proxy metrics, model selection
TL;DR¶
UnlearningSoup replaces some repeated tuning in training-based large language model unlearning with evaluation-guided weight interpolation: EfficientSoup searches using the original model and two trained unlearned models, while PerformanceSoup merges existing candidates, improving empirical forgetting–retention quality and selection costs across benchmarks without eliminating base unlearning training or providing certified forgetting.
Background & Motivation¶
Training-based machine unlearning must reduce the influence of designated data on a large language model (LLM) while preserving other capabilities. GradDiff combines gradient ascent on forget examples with cross-entropy on retain examples; NPO, WGA, and SatImp further control unlearning updates, whereas LUNAR modifies relevant internal representations. The difficulty is not simply whether a better loss exists: small changes in learning rate, unlearning strength, or random seed can move the same method from insufficient forgetting to severe damage to general utility. A strong final model often hides the cost of several preceding training and evaluation runs.
The paper observes that training loss is not a reliable selection signal: it can continue improving while the evaluated forgetting–retention trade-off first improves and then deteriorates. Repeated tuning therefore uses expensive training trajectories to search indirectly for favorable evaluation outcomes. Meanwhile, candidates derived from the same original model frequently occupy a shared high-performance region in two-dimensional weight-interpolation slices. A better solution sometimes lies between trained endpoints rather than at any endpoint. This empirical shared evaluation basin does not imply that models with arbitrary architectures or initializations can be averaged directly.
The authors consequently shift attention from designing another unlearning update to selecting combinations of directions already produced by training. Early in the process, few candidates are available and additional training should be reduced; later, many candidates exist and another expensive run may deliver little improvement. Core idea: use evaluation signals that reflect both forgetting and retention to guide weight-space search, applying two-stage interpolation with an original-model anchor when candidates are scarce and performance-weighted greedy merging when many candidates already exist.
Method¶
Overall Architecture¶
UnlearningSoup is a selection layer around a base unlearning algorithm, not a standalone unlearning loss. Its inputs are parameter-compatible candidates derived from a common original checkpoint and evaluation data; its output is one merged model checkpoint. Experiments with different model families test the method separately, rather than merging LLaMA and Qwen parameters together.
The pipeline first performs base unlearning training, then scores candidates and interpolated models using a validation proxy. With few candidates, EfficientSoup searches two sequentially determined edges; with many candidates, PerformanceSoup sorts by performance and attempts to add models to a recipe. They are alternative strategies selected according to candidate availability, not mandatory successive stages.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
O["Common original model"] --> B["Base unlearning training<br/>Forget and retain data"]
B --> C["Parameter-compatible candidates"]
C --> V["Validation proxy<br/>Score candidates and mixtures"]
Q["Validation evaluation data"] -.-> V
V -->|Two unlearned candidates and original| E["Two-stage edge search<br/>EfficientSoup"]
V -->|Multiple existing candidates| P["Performance-weighted greedy merging<br/>PerformanceSoup"]
E --> F["Final checkpoint<br/>Full evaluation and deployment"]
P --> F
F --> I["User input to single-model output<br/>No search during inference"]
Solid arrows show training, selection, and final-checkpoint processing; the dashed arrow supplies validation data only to scoring. Forget and retain data supervise base training, not additional deployment inputs. Proxy evaluation belongs to offline search, not an inference-time ensemble.
Key Designs¶
1. Validation proxy: do not substitute training loss for forgetting–retention quality
Evaluating only the forget set can favor destructive updates, whereas evaluating only retention can select a model that has barely forgotten anything. The authors use a deviation score whose ideal point has zero extractability on forget data and unit extractability on retain data. For Extraction Strength (ES), the main text gives the following definition, with lower values preferred:
ES measures how much target content remains extractable, so the same metric covers both forgetting targets and retained content. On TOFU candidates, the authors observe a correlation with aggregate statistical quality and lower evaluation cost than computing all metrics repeatedly. The repeated-tuning baseline uses the same proxy for selection, preventing all proxy-related savings from being attributed exclusively to interpolation. Only the selected final model receives full evaluation. This correlation is evidence from the reported settings, not a guarantee for arbitrary removal requests.
The terminology also needs care: the main text calls these test-time performance measures, while the theoretical appendix limits monotonicity to a validation proxy. Experimental details specify evaluation counts but do not clearly establish whether repeatedly consulted proxy examples and final reported examples are independent. This note does not claim a verified independent validation–test protocol; deployment should include held-out evaluation data not used for selection.
2. Two-stage edge search: correct excessive updates with an anchor, then use a second forgetting direction
EfficientSoup first trains two unlearned candidates from the original model using the same base method. In the main experiments, both follow the original method's recommended configuration and differ only in random seed. The better candidate is selected by the proxy. Stage one performs binary-search-style interpolation between that candidate and the original model, retaining the best observed mixture. Stage two repeats the search between this intermediate checkpoint and the other unlearned candidate. Averaging parameters neither averages generated answers nor requires new gradient updates.
The original model anchors retention because base training may overshoot and damage behavior that should remain unchanged. Moving toward the anchor can reduce that damage. The second candidate supplies a different update residual, potentially producing a mixture better than a single training endpoint. However, the anchor also contains information targeted for forgetting: moving toward it can reintroduce target knowledge. Both forgetting and retention must therefore be checked; proximity to the original model is not itself an improvement in safety.
This is a finite-depth, two-stage heuristic edge search, not a global optimizer over the entire triangle. The accessible second-stage segment depends on the point selected in stage one. Membership of the output in the three-model convex hull does not establish global optimality on that hull. The appendix's local convex-quadratic analysis explains interpolation and complementary-error benefits under its assumptions; it does not prove that actual LLM evaluation landscapes are globally convex.
3. Performance-weighted greedy merging: exploit complementary candidates without assigning equal shares
PerformanceSoup targets a stage where several training and evaluation runs have already been completed. It ranks candidates by proxy performance and initializes the ingredient set with the best candidate. It then attempts to add each weaker candidate, recomputes weights using all accepted ingredients' scores, and evaluates the new recipe. A candidate is accepted only when the mixture is no worse than the current recipe. This is not unconditional averaging of all seven models. The main text and Algorithm 1 specify:
The intuition is that stronger candidates are closer to favorable regions and should receive larger shares. A weaker candidate can still contribute complementary errors, but should not immediately receive the same share as stronger members, as under uniform averaging. Greedy acceptance prevents an addition from directly worsening the measured proxy. It does not guarantee monotonic improvement in other metrics or independent test performance, and it is not proof of permanently deleting data influence.
The formula has an important scale ambiguity. The main-text DS includes a factor of 100, and Table 3 reports 42.58, whereas the corresponding entry in appendix Table 12 is 0.4258. Substituting 42.58 directly produces a negative performance score. When all scores and their sum are negative, normalized weights can remain positive, but the relationship “better candidate receives greater weight” can reverse; mixed-sign scores can produce negative weights or a zero denominator. Proposition D.8 explicitly assumes \(P>0\) for every candidate. Although the appendix contains DS values without the factor of 100, this does not establish how the implementation normalizes, clips, or ensures positive scores. This note preserves the published rule and flags the ambiguity rather than adding a correction on the authors' behalf.
A Worked Example¶
Consider GradDiff unlearning on TOFU-10% with LLaMA-3.2-3B: the goal is to remove information about designated fictional authors while retaining other authors and general knowledge. Only defensive model selection is described here; private question–answer content is not reproduced.
The starting point is not zero training: two candidates with different seeds are trained from the same original model. The validation proxy decides which candidate is interpolated with the original first. Stage one retains a better intermediate checkpoint; stage two combines it with the other candidate. The main result table does not disclose final mixing coefficients, so result values cannot justify invented coefficients or a fabricated step-by-step search trajectory.
Table 6 supplies actual evidence of the effect of changing search resolution. At depths 2, 3, and 4, DS is 46.34, 29.03, and 25.19, while SBQ is 7.144, 7.654, and 7.776. Depth 5 yields no further improvement. SBQ here is a final aggregate statistic from full evaluation, not the base training loss: finer weight selection can improve an endpoint, but more search is not always better.
If seven trained candidates already exist, PerformanceSoup instead keeps the proxy-best candidate and attempts to add the remaining six in order. Evaluation of each proposed mixture determines acceptance. Deployment uses one merged parameter set and a single model forward process, not seven candidate generations followed by voting.
Loss & Training¶
The merging stage introduces no new back-propagation loss. Base candidates still use the original objectives of GradDiff, NPO, SimNPO, WGA, SatImp, LUNAR, or BS-T. The overall procedure first produces forgetting directions and then chooses their combination through evaluation; it does not achieve unlearning from the original model alone without training.
Appendix E.4 specifies AdamW, a TOFU batch size of 32, 10 epochs, a linear schedule, and one warm-up epoch. The standard learning-rate exploration is \(\{5\times10^{-6},10^{-5},2\times10^{-5}\}\), with separate settings for LUNAR. The experimental platform has eight NVIDIA A100 GPUs; this is not a statement that every task simultaneously occupies all eight. BS-T lacks a public implementation and is reproduced from its paper. The cache states that code is included in supplementary material but provides no verifiable public repository URL.
The TOFU cost ledger is central: repeated tuning uses 7 training runs, 7 DS evaluations, and 1 full evaluation; EfficientSoup uses 2 training runs, 8 DS evaluations, and 1 full evaluation; PerformanceSoup uses 7 training runs, 13 DS evaluations, and 1 full evaluation. The latter's small additional overhead is relative to an already-trained candidate pool, not an end-to-end cost consisting only of merging.
WMDP and MUSE do not use a fast single-metric proxy; full evaluation is used during search. The three schemes respectively require “training runs / full evaluations” of 6/6, 2/6, and 6/11. MUSE also saves a checkpoint after every epoch and selects among ten checkpoints, adding within-run checkpoint selection. The relationship of these choices to an independent final test is not sufficiently detailed in the cache.
Key Experimental Results¶
Main Results¶
TOFU contains 4,000 question–answer pairs about 200 fictional authors; the paper focuses on Forget 5% and 10%. Higher SBQ is better. Erasing Quality (EQ) is 10 times the harmonic mean of four inverted forgetting statistics. Retention Quality (RQ) is 10 times the harmonic mean of Model Utility and retain ES. SBQ is the root mean square of EQ and RQ. LBQ uses GPT-4o judgments of fluency, relevance, absence of misleading content, and correctness in avoiding target content, aggregated by a harmonic mean. The authors run LBQ three times and report averages.
The following entries are selected from main-text Tables 2 and 4 for TOFU-10%. The seven-run selection baseline is not presented as a universal SOTA across prior work. GPU-hours is the unit labeled in these main tables.
| Model / base method | Selection scheme | SBQ ↑ | LBQ ↑ | GPU-hours ↓ |
|---|---|---|---|---|
| LLaMA-3.2-3B / GradDiff | Repeated tuning | 5.635 | 0.049 | 2.172 |
| LLaMA-3.2-3B / GradDiff | EfficientSoup | 7.776 | 2.150 | 0.675 |
| LLaMA-3.2-3B / GradDiff | PerformanceSoup | 7.454 | 0.516 | 2.185 |
| LLaMA-3.1-8B / LUNAR | Repeated tuning | 7.259 | 7.123 | 4.596 |
| LLaMA-3.1-8B / LUNAR | EfficientSoup | 7.565 | 7.354 | 1.407 |
| LLaMA-3.1-8B / LUNAR | PerformanceSoup | 7.358 | 7.102 | 4.640 |
| Qwen2.5-7B / BS-T | Repeated tuning | 8.082 | 7.416 | 5.345 |
| Qwen2.5-7B / BS-T | EfficientSoup | 8.249 | 7.744 | 1.624 |
| Qwen2.5-7B / BS-T | PerformanceSoup | 8.159 | 7.525 | 5.394 |
For GradDiff here, EfficientSoup increases SBQ by 38.0% and reduces time by 68.9%, but absolute LBQ remains only 2.150. Large relative improvements from a near-zero baseline do not mean that language quality is fully restored. PerformanceSoup on LUNAR increases SBQ while slightly decreasing LBQ, directly illustrating that aggregate statistical quality and linguistic behavior are not identical.
Cross-benchmark results do not unconditionally improve every component either. In main-text Table 5, MUSE-News / SimNPO MemQ rises from 30.485 to 30.725, slightly worsening forgetting, while UtilPres rises from 36.630 to 40.235. For WMDP / BS-T, EfficientSoup reduces Forget Acc from 0.266 to 0.259, increases MMLU from 0.526 to 0.541, and reduces cost from 6.728 to 2.492 GPU-hours. These are defensive capability-evaluation results, without reproducing sensitive questions.
Ablation Study¶
This analysis table combines summaries from main-text Tables 3 and 6. The two groups use different backbones and must not be compared directly. DS follows the main-text scale with the factor of 100.
| Experiment group | Config | DS_ES ↓ | SBQ ↑ | Evidence |
|---|---|---|---|---|
| Qwen2.5-7B / TOFU-10% | Uniform merging | 80.05 | 5.982 | Table 3 |
| Qwen2.5-7B / TOFU-10% | Uniform greedy merging | 46.69 | 6.952 | Table 3 |
| Qwen2.5-7B / TOFU-10% | Reweighted greedy merging | 42.58 | 7.108 | Table 3 |
| LLaMA-3.2-3B / TOFU-10% | Depth 2 | 46.34 | 7.144 | Table 6 |
| LLaMA-3.2-3B / TOFU-10% | Depth 3 | 29.03 | 7.654 | Table 6 |
| LLaMA-3.2-3B / TOFU-10% | Depth 4 | 25.19 | 7.776 | Table 6 |
| LLaMA-3.2-3B / TOFU-10% | Depth 5 | 25.19 | 7.776 | Table 6 |
The gain from uniform merging to uniform greedy merging is substantially larger than the subsequent gain from performance weighting. Rejecting bad recipes after evaluation is therefore an important component; the entire advantage should not be attributed to reweighting. Search depth reaches diminishing returns around 3–4. Changing the candidate pairing through seeds, learning rates, or strength also yields similar SBQ, but only within the reported experiment.
Key Findings¶
- Efficiency comes from fewer base training runs, not eliminating training. Main TOFU tables show roughly 70% savings, and WMDP / MUSE roughly 60%. Appendix experiments with full-metric search still report over 50% savings, but more expensive evaluation weakens the advantage.
- Time units and summary scopes conflict. Table 1 explicitly specifies minutes, whereas nearby prose says GPU hours. Table 15 retains a performance-table-style caption but displays only time-like values. These numbers are not forcibly converted or combined into one precise timing table here.
- The prose summarizes EfficientSoup gains as 0.9%–32.3%, but Table 2 already contains 38.0% and Table 8 contains 44.4%. The PerformanceSoup paragraph also mistakenly names EfficientSoup, and its stated range does not cover all table entries. This note reports original values for specified models and SBQ / LBQ rather than imposing one overall gain range.
- Appendix E.2 defines Forget Acc as the root mean square of Bio and Cyber accuracy, while main-text §5.1 calls it a harmonic mean. Table results are retained with this definition conflict disclosed. MUSE MemQ is the root mean square of VerbMem and KnowMem, with lower values preferred.
- Appendix Table 19's LLaMA-3.2-3B / 5% entries visibly differ from Table 9. Table 20's entire Qwen2.5-7B / 10% block repeats the 3B values and conflicts with Table 4. Some 1.5B / 5% rows in Table 18 also appear to have shifted RQ / SBQ columns; they are not used to establish new conclusions.
Highlights & Insights¶
- The paper compares the total cost of obtaining a useful unlearned model, rather than only its final checkpoint. This perspective exposes the practical maintenance cost of repeated tuning and prevents low marginal post-processing overhead from being mistaken for low end-to-end cost.
- The original-model anchor makes update strength selectable after training, avoiding retraining for every learning-rate adjustment. It can correct moderate over-unlearning, but needs an explicit forgetting threshold so recovered general utility does not conceal reintroduced target knowledge.
- Merging benefits depend on complementary candidate errors, not simply candidate count. The local appendix theory explains this through average member risk minus a diversity term; practical evaluation is still required to determine whether that diversity helps.
Limitations & Future Work¶
- The shared basin depends on aligned parameters, a common starting point, and non-collapsed candidates. Different architectures or severe training collapse have no general guarantee. If every candidate under-unlearns, convex combinations cannot create a missing forgetting direction.
- The positive-score assumption for weighting is not reconciled with the main-text DS scale; implementation verification is required before reproducing weights. Rankings can be invariant to monotonic rescaling while normalized weights remain sensitive to scale and offset, so these should not be conflated.
- ES can miss residual knowledge outside its measurement scope. SBQ's root mean square also permits strength on one side to compensate for weakness on the other. Proxy optimality or higher SBQ cannot replace independent leakage assessment or establish certified forgetting.
- The authors assess cross-lingual, adversarial-prompt, and subsequent-training stress conditions, with broadly comparable behavior to repeatedly tuned models, but recovery risk remains. In appendix Table 11, post-stress VerbMem for BS-T is 23.1191, versus 24.4796 for EfficientSoup and 24.8208 for PerformanceSoup, showing worse results in some settings. No attack or recovery procedure is provided here.
- Validation–test separation, selection counts, checkpoint selection, and variance reporting need clearer treatment. Fixed validation budgets, independent final testing, and separate minimum forgetting requirements could distinguish selection overfitting, utility recovery, and actual target removal.
Related Work & Insights¶
- vs ModelSoups: Uniform or uniform greedy merging exploits candidates from a shared starting point; this paper adds an unlearning-specific proxy, an original-model anchor, and performance weighting. It does not prove that ModelSoups fails in every unlearning scenario, but demonstrates better selection in the reported candidate pools and evaluations.
- vs GradDiff / NPO / WGA / LUNAR: These methods generate base unlearning updates, while UnlearningSoup searches checkpoint combinations around them. The approaches are complementary; the paper does not bypass base training on forget and retain data.
- vs Task Arithmetic / NegMerge: These approaches design editing or deletion mechanisms from task vectors or update directions; this paper primarily reduces selection costs in an existing training pipeline. The transferable lesson is to verify parameter compatibility and proxy reliability before evaluation-guided combination, not to merge arbitrary checkpoints directly.
Rating¶
- Novelty: 3/5 — Model merging is established, but the early- and late-stage strategies specifically address unlearning selection.
- Experimental Thoroughness: 4/5 — Broad benchmarks, models, base methods, and ablations, with reproducibility weakened by split, scale, and duplicated-column issues.
- Writing Quality: 3/5 — Clear motivation and workflow, but several inconsistencies in summary ranges, units, metric definitions, and appendix tables.
- Value: 4/5 — Useful for practical selection efficiency when independent forgetting evaluation is retained and candidate training costs are counted.