Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models¶
Conference: NeurIPS2026
arXiv: 2609.36254
Area: LLM Safety / Alignment & RLHF
Keywords: safety awareness, chain-of-thought, reasoning–answer consistency, process rewards, over-refusal
TL;DR¶
The paper measures inconsistent safety judgments between reasoning traces and final answers with DSAR and introduces SARA, an on-policy reinforcement learning method combining early safety awareness, full-trace safety, and answer safety; it improves consistency under perturbed reasoning on two DeepSeek-based models, but does not outperform answer-only rewards in every setting or safety metric.
Background & Motivation¶
Large reasoning models generate a chain-of-thought before their final answers, whereas training rewards often concentrate on the final output. A safe answer therefore need not reflect a safe generation process: the model may follow an unsafe direction during reasoning and refuse only at the end, or identify a risk but fail to preserve safety in its final answer. Answer-only evaluation conceals the former case; searching only for safety-related language in reasoning cannot exclude the latter.
SafeChain, SafePath, and STAR-1 use supervised fine-tuning to learn prepared safety trajectories, but fixed training trajectories differ from reasoning generated by the model itself. RECAP reduces this mismatch through on-policy reinforcement learning and perturbed-trajectory augmentation, yet mainly rewards final answers. This paper argues that the model should identify risks early in its own generated trajectories while both the entire reasoning trace and the answer are evaluated, rather than receiving safety feedback only at the endpoint.
“Deceptive safety alignment” is the authors' term for inconsistent observable safety signals, not evidence of actual deceptive intent, hidden objectives, or faithful chain-of-thought. Core Idea: reward early recognition and rejection of risk within reasoning, and jointly optimize that signal with full-trace and final-answer safety so that safety feedback influences on-policy generation earlier.
Method¶
Overall Architecture¶
SARA trains on requests labeled harmful or benign, with half of the training examples receiving perturbed-trajectory augmentation. The policy produces several candidates per request, separating each into a reasoning trace and final answer; different judges provide safety or refusal signals, and DAPO updates the policy. At inference time, the updated model generates reasoning and an answer without requiring the training judges to act as sentence-by-sentence blocking mechanisms.
The method has three designs: Dual-Branch Safety Evaluation establishes the signals being monitored, Early-Awareness Joint Reward constructs the harmful-request reward, and Benign Protection and On-Policy Training prevents safety learning from collapsing into universal refusal. Dashed edges indicate training supervision or the relationship to trained weights; solid edges indicate evaluation data flow or inference-time generation.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Training candidates:<br/>reasoning and answer"] --> B["Dual-Branch<br/>Safety Evaluation"]
B -.->|Judgments for harmful requests| C["Early-Awareness<br/>Joint Reward"]
C -.->|Training reward| D["Benign Protection and<br/>On-Policy Training"]
E["Benign-request refusal score"] -.->|Training supervision| D
D -.->|Trained weights| F["Updated policy model"]
G["Inference-time request"] --> F --> H["Reasoning trace and final answer"]
Key Designs¶
1. Dual-Branch Safety Evaluation: distinguish reasoning safety awareness from answer safety
The authors do not equate refusal words with safety awareness. The sentence-level judge, GPT-oss-safeguard-20B, must determine both whether a sentence recognizes harmful intent in the input and whether it uses that recognition to refuse, stop, or redirect toward a safe alternative. Acknowledging risk while continuing in an unsafe direction does not qualify. The reasoning-awareness indicator \(r_{\mathrm{SA}}\) is 1 if at least one sentence meets the criteria; SAR is the proportion of traces with this indicator, not the proportion of safe sentences.
A second reasoning branch uses Qwen3Guard to judge whether the entire trace is safe, producing \(r_{\mathrm{safe}}\); the answer branch uses the same guard to obtain final-answer safety indicator \(s\). The two reasoning signals are combined with logical OR to accommodate traces that remain safe without explicitly refusing. The inconsistency metric is then:
The expectation covers test requests and model generations. SS is the proportion of final answers judged safe; \(1-\mathrm{DSAR}\) in the main results is consistency, reported as a percentage. SAR is not interchangeable with \(r\), so DSAR is not simply the difference between SAR and SS.
The definition has an important boundary: a trace containing one qualifying safety sentence may be labeled reasoning-safe through the OR even if its later continuation is unsafe. Conversely, when both reasoning and answer are judged unsafe, DSAR is still zero. Low DSAR therefore indicates agreement between judgments and must be interpreted alongside SS, SAR, and full-trace safety analysis, not as an independent safety guarantee.
2. Early-Awareness Joint Reward: move risk recognition toward the beginning of reasoning
Full-trace safety rewards alone may miss unsafe reasoning tendencies when a trace does not contain explicit unsafe details. For harmful requests, SARA uses IBM Granite-Guardian-3.1-8B to score reasoning safety probability \(R_i^{\mathrm{cot}}\) and answer safety probability \(R_i^{\mathrm{ans}}\), then locates the first safety-aware sentence with the sentence-level judge. A trace has \(N_i\) sentences indexed from 0, with the first qualifying position denoted \(k^\star\); if none qualifies, \(k^\star=N_i\).
Recognition in the first sentence receives the maximum awareness reward, whereas no recognition produces no reasoning-side reward. Multiplication requires both high full-trace safety probability and early awareness for a large reasoning reward, rather than optimizing either proxy alone; the separate answer term preserves endpoint supervision. Compared with answer-only rewards, the design places learning pressure on reasoning rather than merely increasing the probability of a final refusal.
However, multiplication is a soft reward, not a hard constraint. A zero reasoning term still permits a total reward of up to 0.5 from a safe answer, and judge probabilities are not reliability certificates. Relative sentence-position rewards can also depend on sentence segmentation and trace length; an early qualifying sentence does not ensure that every subsequent step remains safe.
3. Benign Protection and On-Policy Training: prevent safety optimization from becoming universal refusal
For benign requests, the model does not receive the harmful-request consistency reward. Instead, DS-Qwen2-32B scores final-answer refusal severity from 0 to 10, with reward equal to 1 minus this score divided by 10. Complete refusal receives zero, while less refusal receives higher reward. The policy must distinguish requests requiring refusal from those warranting assistance, rather than achieving good scores by refusing everything.
Training uses DAPO: the old policy samples a group of trajectories for each request, scalar rewards are normalized within the group into advantages, and a token-level policy objective with asymmetric clipping updates the policy. Although safety awareness is judged sentence by sentence, optimization still uses a scalar trajectory-level reward, not independent dense rewards for every sentence. The paper also does not directly minimize DSAR; its reward shaping improves consistency indirectly.
Half the examples receive perturbed-trajectory augmentation, requiring recovery from reasoning contexts that deviate from the target behavior. This note discusses only that defensive training arrangement and does not reproduce perturbation text or construction procedures. Training and main evaluation use the same class of augmentation mechanism, making additional unseen-perturbation evaluation important, but not representative of all out-of-distribution safety settings.
A Worked Example¶
Suppose a training trace initially recognizes a risk, subsequently fails to remain safe, and recovers safety in its final answer. This illustrates judgment logic without reproducing a request or trace. DSAR's OR may mark the reasoning as safe and, with a safe answer, classify the pair as consistent. SARA still uses full-trace safety probability in its reward: initial recognition alone cannot maximize the reasoning reward. Consistency must therefore be interpreted alongside full-trace and answer safety, and the judges are not inference-time hard blockers.
Loss & Training¶
The training set contains 2K requests: 1K harmful requests from SafeChain and 1K benign requests from FalseReject, with no overlap with evaluation samples. Training experiments cover only DeepSeek-R1-0528-Qwen3-8B and DeepSeek-R1-Distill-Qwen2-14B; the broader five-model survey does not mean that SARA was trained and validated on five models.
Appendix B.1 specifies LoRA rank 8, scaling factor 16, and all linear layers as targets; learning rate \(3\times10^{-5}\), weight decay 0.1, 10 warmup steps, 1 training epoch, prompt batch size 32, and 4 rollouts per request. Generation temperature and top-p are both 1.0, with a maximum prompt length of 3972 tokens.
DAPO uses lower and upper clipping parameters of 0.2 and 0.28, zero KL coefficient, and token-level loss averaging. Training runs use two H100 80GB GPUs, with reward models hosted separately; the authors report approximately 5 and 6 GPU hours per run for the two models, respectively. These figures should not be extrapolated into a complete cost estimate for every ablation.
Key Experimental Results¶
Main Results¶
The table selects Original, RECAP, and SARA from source Table 2. StrongReject uses 313 samples under perturbed evaluation, while SafeChain uses 500 samples under standard evaluation; values are percentages except Avg., which also uses a 0–100 scale. Each safety entry lists SAR / SS / consistency \(1-\mathrm{DSAR}\), all higher-is-better.
| Model | Method | StrongReject: SAR / SS / consistency | SafeChain: SAR / SS / consistency | Avg. |
|---|---|---|---|---|
| DS-Qwen3-8B | Original | 53.40 / 87.86 / 66.77 | 56.60 / 79.80 / 85.00 | 72.23 |
| DS-Qwen3-8B | RECAP | 57.80 / 99.40 / 70.93 | 71.00 / 98.80 / 96.30 | 77.61 |
| DS-Qwen3-8B | SARA | 75.10 / 97.44 / 85.30 | 72.80 / 95.40 / 96.40 | 82.76 |
| DS-Qwen2-14B | Original | 32.90 / 53.67 / 74.44 | 33.80 / 64.00 / 79.80 | 68.98 |
| DS-Qwen2-14B | RECAP | 70.90 / 99.04 / 80.51 | 75.80 / 96.60 / 95.80 | 81.67 |
| DS-Qwen2-14B | SARA | 77.00 / 95.85 / 84.66 | 80.40 / 95.00 / 94.60 | 81.97 |
Under perturbed evaluation on Qwen3, SARA exceeds RECAP by 17.30 percentage points in SAR and 14.37 points in consistency, but loses 1.96 points in SS. Corresponding changes on Qwen2 are +6.10, +4.15, and −3.19 points. Under standard evaluation, Qwen2 consistency instead drops from RECAP's 95.80 to SARA's 94.60, so improvement is not universal across settings.
Utility does not improve uniformly: SARA's GSM8K / MMLU-Pro accuracy is 86.28 / 61.74 on Qwen3 and 81.27 / 56.58 on Qwen2; the latter GSM8K result is slightly below Original's 81.58. Qwen3 OR-Bench HS increases from RECAP's 67.02 to 91.21, but HS measures non-refusal, not correctness or assistance quality.
Avg. first averages results within safety, helpfulness, and utility, then takes the harmonic mean of those three scores. The safety dimension mixes SAR, SS, consistency, and SS under unseen perturbations, so a better aggregate cannot replace checking individual safety metrics. SafePath, RECAP, and SARA use the same 2K examples; SafeChain replaces the benign portion, and STAR-1 uses its own dataset. The baselines therefore do not all use identical training data.
Ablation Study¶
The reward ablation in source Table 3 uses DS-Qwen3-8B; the table preserves its one-decimal precision. The complete reward improves awareness and consistency under perturbation, but is not the configuration with the highest final-answer safety.
| Reward configuration | StrongReject: SAR / SS / consistency | SafeChain: SAR / SS / consistency |
|---|---|---|
| Original | 52.4 / 87.9 / 66.8 | 56.6 / 79.8 / 85.0 |
| Answer safety only (RECAP) | 57.8 / 99.4 / 70.9 | 71.0 / 98.8 / 96.3 |
| Full-trace reasoning safety only | 64.2 / 99.4 / 84.4 | 65.4 / 97.4 / 96.0 |
| Safety awareness only | 71.6 / 93.3 / 81.5 | 71.4 / 87.4 / 91.8 |
| Full SARA reward | 75.1 / 97.4 / 85.3 | 72.8 / 95.4 / 96.4 |
Awareness-only rewards yield higher SAR than full-trace reasoning rewards, but lower SS and consistency. Producing safety-aware language and maintaining full-trace safety are different supervision targets. The complete reward combines them while retaining a trade-off in answer safety.
The augmentation ablation in source Table 6 further demonstrates that robustness gains depend on the evaluation condition.
| Model | SARA training augmentation | StrongReject: SAR / SS / consistency | SafeChain: SAR / SS / consistency |
|---|---|---|---|
| DS-Qwen3-8B | Without | 62.9 / 92.1 / 78.0 | 89.0 / 97.8 / 98.0 |
| DS-Qwen3-8B | With | 75.1 / 97.4 / 85.3 | 72.8 / 95.4 / 96.4 |
| DS-Qwen2-14B | Without | 76.4 / 90.2 / 81.2 | 90.2 / 95.2 / 95.4 |
| DS-Qwen2-14B | With | 77.0 / 95.9 / 84.7 | 80.4 / 95.0 / 94.6 |
For both models, augmentation improves all three metrics under perturbation but lowers all three under standard evaluation. Without augmentation, SARA still exceeds RECAP in perturbed SAR by 11.1 and 22.7 points, respectively, supporting a reward-design contribution independent of augmentation.
Key Findings¶
- Unseen-perturbation evaluation reports answer SS only: on Qwen3, ICL / H-CoT results are 100.0 / 50.00 for RECAP and 99.25 / 52.00 for SARA; on Qwen2, they are 96.60 / 56.00 and 96.75 / 56.00. These results do not demonstrate improved reasoning–answer consistency under those conditions.
- Human validation uses only 100 Qwen2 generations under standard StrongReject evaluation. Appendix D reports 93%, 95%, and 92% agreement for safety awareness, answer safety, and DSAR categories, respectively. The inconsistency categories contain few examples, limiting claims about stable accuracy across models and perturbations.
- Held-out judges support perturbed SAR gains: RECAP / SARA score 67.1 / 82.4 with gpt-oss-120B and 54.3 / 72.5 with Qwen2.5-32B. Under standard evaluation, Qwen2.5-32B reverses the ordering, with RECAP at 74.8 and SARA at 72.0; superiority is not judge-independent.
Source conflicts must be retained: Table 1 gives Qwen2 StrongReject DSAR as 26.20→25.56, contradicting the claim that perturbation increases DSAR for every model. The text attributes SafeChain's 15.00→34.40 to Qwen2, whereas Table 1 assigns it to Qwen3; Qwen2 instead has 20.20→22.00.
Tables 3 and 6 give Qwen3 Original SAR as 52.4, versus 53.40 in Tables 1 and 2, which is not a rounding difference. The main text claims 95% human agreement for the sentence-level judge, whereas Appendix D reports 93% for awareness and 95% for answer judgments. The text also claims SafeChain improves Qwen3 standard SS, but Table 2 gives 79.20, below Original's 79.80. This note preserves each table's values rather than silently correcting them.
Highlights & Insights¶
- Separating final safety from process-level awareness identifies models that recover safety only at the endpoint. DSAR is useful as a supplementary diagnostic, not a replacement for safety scores.
- Early-awareness rewards consider the position of a safety decision, rather than merely whether safety language appears. Multiplication by full-trace safety probability makes the local signal contribute jointly with a trajectory-level signal.
- A separate benign-request refusal reward directly addresses over-refusal introduced by safety fine-tuning. Because this reward targets reduced refusal, correctness and assistance quality still require additional supervision.
Limitations & Future Work¶
- The authors acknowledge training only two DeepSeek-based models, without covering larger models, closed-source models, long-horizon strategic deception, or tool use. The five-model descriptive survey does not fill these training-validation gaps.
- DSAR's OR permits a safety sentence to mask later unsafe continuation, while jointly unsafe reasoning and answers count as consistent. Future evaluation could report all four reasoning–answer categories separately and measure whether generation deviates again after recognizing risk; these are not results established by this paper.
- Judge-based rewards may respond to surface language patterns, and the same awareness judge participates in training and main evaluation. Held-out judges address part of this concern but evaluate SAR only, not the complete DSAR pipeline.
- Representation analysis finds stronger benign–harmful separation at the answer stage, but Appendix F constructs that stage by directly switching stage tags and skipping actual reasoning. It reveals stage-conditioned differences, not causal tracking through complete generation, and does not establish that answer-only rewards caused the difference.
- Relative sentence-position rewards require further checks for segmentation, length, and subsequent deviations. The reward combination and limited tests provide empirical evidence, not a universal proof that supervising intermediate reasoning is necessary for reliable safety.
Related Work & Insights¶
- vs RECAP: Both use on-policy DAPO and perturbed-trajectory augmentation; SARA adds full-trace reasoning safety and early-awareness rewards. Its strongest evidence is improved SAR and consistency under perturbation, with lower answer SS.
- vs SafeChain / SafePath / STAR-1: These methods learn safety trajectories through supervised fine-tuning, whereas SARA evaluates trajectories generated by the current policy. Because some baselines use different training data, the main table compares complete methods rather than isolating the causal effect of rewards.
- vs Compliance Gap / D-REX: The former concerns behavior across monitoring conditions, and the latter uses multiple criteria to evaluate deceptive reasoning; DSAR focuses on agreement between two safety judgments within one generation. Its narrower scope should not be conflated with actual intent, strategic deception, or chain-of-thought faithfulness.
Rating¶
- Novelty: 4/5. Joint early-awareness and trajectory-safety rewards are targeted, but build on existing on-policy RL and judges.
- Experimental Thoroughness: 3/5. Main results, reward and augmentation ablations, and judge validation are substantial; training-model coverage and human-validation scale remain limited.
- Writing Quality: 3/5. The mechanism is clear, but several text–table conflicts and overly strong conclusions reduce reliability.
- Value: 4/5. The paper offers process-safety diagnostics and training ideas; practical use still requires monitoring answer safety, over-refusal, and metric distortion.