LLM Alignment–Utility Asymmetry under Semantic-Preserving Transformations¶
Conference: NeurIPS2026
arXiv: 2609.32717
Area: LLM Safety
Keywords: alignment generalization, task utility, distribution shift, paired evaluation, safety supervision
TL;DR¶
Through paired evaluation of original and semantic-preserving representations, this paper finds that task capability transfer does not guarantee corresponding transfer of safety behavior, and qualifies this empirical finding through benign-data conditions, protocol controls, and defensive supervision analysis.
This note covers only defensive research questions, aggregate evidence, and experimental boundaries; it provides no transformation rules, codebooks, reversible procedures, adaptation recipes, prompts, or attack examples.
Background & Motivation¶
Instruction tuning and reinforcement learning from human feedback (RLHF) make large language models (LLMs) more responsive to user needs and safety requirements under familiar language inputs. However, refusal tests in standard form cannot answer a different question: does existing safety behavior persist when task meaning stays unchanged but the surface input distribution changes substantially? Previous studies have reported safety failures across languages or unusual representations, but many primarily record failure rates without also verifying whether the model can still perform ordinary tasks under the same representation conditions. If a model cannot understand the input at all, a low failure rate may reflect missing capability rather than robust safety generalization.
The paper therefore shifts attention from isolated safety failures to joint changes in task utility and aligned behavior. Natural-language variation is often covered by pretraining, making existing language knowledge difficult to separate from behavioral transfer to a new representation; the authors use controlled synthetic representation shifts to reduce this confound. Nevertheless, this controlled experiment does not fully simulate real user distributions: processing a new representation, retaining ordinary task capability, and maintaining safety constraints are three claims requiring separate verification. Adaptation can also change original-input performance, so the original model's capability scores cannot directly serve as controls for the adapted model.
Core idea: measure original- and transformed-condition task utility and safety failure under the same representation shift and adaptation conditions, testing whether capability generalization is accompanied by equally robust alignment generalization rather than treating task competence as evidence of continued safety.
Method¶
Overall Architecture¶
This is a controlled behavioral evaluation study, not a new defensive network. The experiments are organized around paired behavioral evaluation, condition separation, protocol controls, and evidence stratification: they first compare the same model across two representation conditions, then examine whether differences arise from adaptation, task format, or scoring. The outputs are task-specific utility changes, safety-failure changes, and their scope, not a single score guaranteeing safety.
Original and transformed conditions use corresponding evaluation questions, with outputs placed in a commonly interpretable evaluation form before scoring. This prevents output-format differences from being counted directly as behavioral differences; this note describes only the evaluation purpose, not representation procedures. Correct interpretation of the question, task correctness, and maintenance of safety constraints still require separate judgments. As a mechanism and evaluation analysis, the paper does not warrant turning its section outline into an operational flowchart.
Key Designs¶
1. Paired behavioral evaluation: compare changes rather than combine different metrics
The authors use MMLU, ARC-Challenge, and GSM8K to assess utility on knowledge, reasoning, and mathematics tasks, and AdvBench to assess safety behavior. \(U_o\) and \(U_t\) denote task accuracy under original and transformed conditions, while \(A_o\) and \(A_t\) denote the corresponding safety-failure rates. Safety failure follows the paper's ASR definition: the proportion of evaluation questions judged to produce unsafe responses; lower values are better. Task accuracy is better when higher, so the two metrics have opposing interpretations and do not become the same quantity merely because both are percentages.
The paper's asymmetry concerns different stability of the two behaviors under the same distribution shift, not a universal law defined in advance. Results can be read through the following separate signed changes; this restates the source's change quantities and does not introduce a composite metric:
Negative \(\Delta U\) indicates lower utility, and positive \(\Delta A\) indicates more safety failures. These changes should be reported task by task rather than subtracting accuracy from ASR to construct a supposed overall safety score. The metrics come from different question sets, and contrasting their change magnitudes does not establish equally difficult capabilities or a shared internal mechanism.
2. Condition separation: distinguish acquiring representation competence from deteriorating safety behavior
The authors examine parameter updates and in-context learning (ICL) as two adaptation categories. The former can change behavior in both the original and new representation spaces, while the latter leaves weights unchanged but remains affected by context content and commercial serving mechanisms. Each condition therefore requires its own original-representation control; scores from different models and adaptation methods should not be assembled into a universal safety ranking.
A key control uses only benign data: does capability–safety separation persist without transformed unsafe supervision examples? Under this condition, Llama3-8B retains utility relatively close to its original-representation performance on three tasks, but its safety-failure rate changes from 1.7% to 40.3%. This supports the claim that direct unsafe supervision is not necessary for the phenomenon, not the claim that benign adaptation necessarily produces the same outcome in every model. The main commercial ICL experiment includes additional contextual factors, and the authors explicitly acknowledge that it cannot be interpreted as a pure representation-change effect; the benign condition adds evidence without automatically eliminating every confound.
3. Protocol controls: avoid mistaking output format and task differences for generalization mechanisms
Ordinary tasks and safety tasks in the main experiment do not have identical output protocols, and utility scoring and safety judgments also use different procedures. Because these choices can affect the measured gap, the authors add matched-output-format controls and benign counterfactual questions matched in source domain, length, structure, and open-ended response format. Counterfactual evaluation uses human judgments of coherent ordinary instruction completion, reducing task-format differences between multiple-choice accuracy and open-ended safety behavior.
Across four models in the matched counterfactual evaluation, ordinary instruction completion decreases by only 0.7–4.4 percentage points, while corresponding safety failure increases by 51.0–64.0 percentage points. These findings strengthen the evidence for different relative robustness without making utility and safety metrics interchangeable. Matched-output-format experiments still reveal divergence, but more explicit content interpretation substantially reduces safety failure, showing that the measured effect size depends on the evaluation protocol. A high failure rate under one protocol should not be treated as an invariant property across serving interfaces, output requirements, and protection configurations.
4. Evidence stratification: separate behavioral change, defensive conditions, and mechanistic interpretation
The authors check for formatting failures, uninterpretable responses, and task-irrelevant outputs to avoid mistaking scoring-pipeline failures for actual behavioral changes. Format problems are rare in a representative run, supporting a behavioral explanation for part of the phenomenon; however, category proportions from one representative run cannot replace three-run averages. The appendices additionally validate some automatic safety labels through human review and examine directional stability with larger evaluation sets. These controls improve evidence reliability without certifying every model, task, and output through human assessment.
The defensive analysis compares whether safety supervision covers the new representation distribution, associating such coverage with lower safety failure while broadly retaining task utility. This is better interpreted as evidence that safety training requires distribution-coverage verification than as a directly deployable defense recipe. Serving-side refusal and filtering can also change end-to-end behavior and should count as effective protection, not as constraints to circumvent. Representational probing and causal analyses in the appendix support weaker engagement of safety-related pathways, but do not identify a complete circuit or establish one mechanism shared by different adaptation methods.
Key Experimental Results¶
Main Results¶
The table selects results from source Tables 1–2. Utility and safety failure are percentages; original → transformed compares two evaluation conditions within the same adapted model. Each main-text evaluation uses 100 held-out examples; results generally aggregate three independent runs, and entries are mean ± standard deviation, not full-benchmark scores or confidence intervals. Paired values are retained without constructing a cross-model overall ranking; the main experimental conditions are not benign-only conditions.
| Model / adaptation category | MMLU utility: original → transformed | ARC-Challenge utility: original → transformed | GSM8K utility: original → transformed | AdvBench safety failure: original → transformed |
|---|---|---|---|---|
| Llama3-8B-Instruct / parameter updates | 35.0 ± 4.0 → 37.0 ± 8.9 | 57.3 ± 6.7 → 72.3 ± 6.7 | 63.3 ± 4.7 → 64.3 ± 5.1 | 3.7 ± 0.6 → 67.7 ± 3.5 |
| Qwen2.5-7B-Instruct / parameter updates | 54.7 ± 1.5 → 40.0 ± 4.4 | 75.0 ± 10.4 → 68.3 ± 11.5 | 75.0 ± 1.0 → 4.7 ± 1.5 | 3.7 ± 1.5 → 67.0 ± 5.3 |
| Mistral-7B-Instruct / parameter updates | 32.3 ± 5.9 → 29.0 ± 2.6 | 75.0 ± 7.8 → 68.7 ± 8.6 | 52.7 ± 11.1 → 54.0 ± 6.2 | 80.3 ± 1.5 → 78.0 ± 7.8 |
| GPT-4.1 mini / parameter updates | 62.7 ± 3.8 → 53.0 ± 2.6 | 89.0 ± 1.7 → 90.0 ± 4.6 | 83.0 ± 3.0 → 88.0 ± 2.6 | 13.3 ± 0.6 → 74.3 ± 2.5 |
| Gemini 3 Flash / ICL | 89.7 ± 1.5 → 82.3 ± 1.5 | 98.3 ± 0.6 → 98.3 ± 0.6 | 97.7 ± 1.5 → 95.7 ± 0.6 | 2.3 ± 1.5 → 43.0 ± 8.9 |
| Claude 4 Sonnet / ICL | 72.3 ± 7.8 → 52.0 ± 3.6 | 88.7 ± 9.3 → 89.7 ± 2.5 | 96.3 ± 0.6 → 86.3 ± 2.1 | 0.0 ± 0.0 → 12.0 ± 1.7 |
Source Table 2 reports a 7.3-percentage-point absolute MMLU gap for Gemini 3 Flash, whereas subtraction of the displayed rounded means yields 7.4; the original paired means are retained here without silently correcting the source's gap. Qwen2.5 on GSM8K is a clear utility-collapse exception, while Mistral has a ceiling effect because original-condition safety failure is already high. These rows prevent reading a frequently observed asymmetry as a claim that utility never falls and safety always deteriorates in every model.
Ablation Study¶
The following defensive-condition analysis selects Llama3-8B results from source Tables 3 and 19; entries again show percentages as mean ± standard deviation. The first row is a benign-adaptation reference; the latter two match the representation condition of safety supervision, so all three rows should not be treated as a perfectly single-variable three-way experiment. Only research conditions concerning supervision coverage and aggregate outcomes are listed, without data-construction instructions, prompts, adaptation settings, or operational steps.
| Defensive research condition | MMLU utility: original → transformed | ARC-Challenge utility: original → transformed | GSM8K utility: original → transformed | AdvBench safety failure: original → transformed |
|---|---|---|---|---|
| Benign-only adaptation reference | 42.3 ± 2.5 → 41.7 ± 3.1 | 69.7 ± 6.4 → 72.7 ± 2.5 | 67.0 ± 3.5 → 64.3 ± 2.1 | 1.7 ± 0.6 → 40.3 ± 6.0 |
| Safety supervision limited to original representation | 42.0 ± 2.1 → 42.3 ± 3.3 | 68.0 ± 5.0 → 71.7 ± 1.4 | 67.3 ± 3.1 → 64.0 ± 3.5 | 1.7 ± 0.6 → 38.3 ± 2.3 |
| Safety supervision covers new representation | 43.0 ± 1.9 → 40.7 ± 1.5 | 67.3 ± 4.0 → 73.3 ± 1.8 | 66.7 ± 4.1 → 65.3 ± 1.9 | 1.0 ± 0.0 → 8.3 ± 3.0 |
The latter two rows show that rehearsing safety behavior only in the original representation and covering the new distribution with safety supervision cannot be treated as equivalent protection. The lower 8.3% rate is still nonzero risk, and experiments under these limited conditions do not establish defensive generalization guarantees for unknown representations.
Key Findings¶
- Utility and safety should be reported together: GPT-4.1 mini's safety failure increases by 61.0 percentage points, while absolute utility changes across the three tasks are 9.7, 1.0, and 5.0 percentage points.
- Benign adaptation also requires independent safety acceptance checks: direct unsafe supervision is not necessary for the observed generalization gap.
- Capability decline is not evidence of safety: under strong structural perturbations in the appendix, exact-answer utility can approach zero while some open-ended instruction-related behavior persists.
- Commercial serving protections have practical effects: some versions refuse evaluation participation or return empty responses, which is end-to-end protective behavior and does not establish the underlying model's latent mechanism.
Highlights & Insights¶
- Evaluating utility and safety together under distribution shift better excludes false robustness caused by failure to understand inputs than reporting safety failure alone. Paired observations constrain interpretation rather than produce a total score replacing comprehensive acceptance checks.
- Output-format controls, benign counterfactuals, and defensive supervision conditions examine alternative explanations from different directions. They support independent verification of alignment transfer while showing that effect sizes depend on protocol.
- Additional evaluations include sycophancy resistance and fairness, extending the research question beyond refusal behavior. Evidence on these limited objectives does not imply that every alignment objective degrades by the same magnitude.
Limitations & Future Work¶
- Semantic preservation is an empirical working definition, not a theorem; recoverable information does not automatically establish equivalent internal semantic representations.
- Each main-text task uses only 100 examples, and three-run standard deviations do not fully quantify sampling uncertainty or ensure detection of rare behaviors.
- Safety sensitivity analyses with 200 / 300 questions use generated rewrites of the original held-out questions, not the same number of newly independent original items; example dependence remains.
- Human validation accuracy for expanded GPT-4.1 mini labels is 96.0%, 94.0%, and 93.3%, covering only that selected condition rather than certifying labels for every model.
- Main-experiment ICL context and task-output protocols introduce confounds; commercial API outcomes also include unobservable serving-side filtering and version changes.
- Benign ICL results in Appendix H.3 lack standard deviations and cannot be directly pooled with main-text three-run summaries; reversals across conditions do not establish universal monotonic relationships.
- Future defensive research should verify safety coverage across representations, false refusals on ordinary requests, and generalization to independent question sets while retaining input and output protections, rather than optimizing a single failure rate.
Related Work & Insights¶
- Compared with shallow-alignment research: the paper provides paired behavioral evidence under distribution shift consistent with incomplete safety generalization, but behavioral gaps do not prove that all safety mechanisms depend only on surface patterns.
- Compared with multilingual safety research: synthetic representations reduce the influence of known natural-language coverage but are less realistic and do not directly represent multilingual deployment risk.
- Compared with evaluations reporting safety failure alone: capability controls and protocol analyses improve interpretation of the relationship between high failure rates and genuine capability transfer, rather than creating a new attack-effectiveness ranking.
- Defensive insight: release acceptance checks should separately examine ordinary task completion, refusal correctness, and unusual-representation behavior; passing one cannot replace checking the other two.
Rating¶
- Novelty: 4/5 — The main contribution is paired evaluation and identification of a generalization question, not a new representation operation.
- Experimental Thoroughness: 3/5 — Models and controls are broad, but small fixed question sets, protocol confounds, and unobservable commercial systems constrain interpretation.
- Writing Quality: 4/5 — Empirical patterns are distinguished from theoretical guarantees; individual rounded gaps and appendix statistical conventions require care.
- Value: 4/5 — Adds an important dimension to defensive acceptance checks without providing a broadly validated general-purpose defense.