Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems¶
Conference: NeurIPS2026 (acceptance information supplied in the task package)
arXiv: 2609.30383
Area: LLM Safety / Safety Evaluation
Keywords: skill cascading, compositional security, shared context, counterfactual verification, safety evaluation
TL;DR¶
Using 213 selected cases in SkillCascade-Bench, the paper shows that passing individual skill reviews does not imply safe joint execution; sandbox stress tests across three agent systems and eight models yield an author-reported average success rate of 89.4%, while a behavior-composition defense achieves only partial improvement.
Background & Motivation¶
A skill is not a simple function but a capability package containing natural-language instructions, scripts, references, and assets. Skill-based agents dynamically load these capabilities for a task and reuse third-party components, making installation review an important supply-chain boundary. However, skills do not operate in completely independent semantic execution spaces: the same large language model reads a shared context, and earlier outputs inform later interpretation and selection. Components therefore pass not only data but also assessments of that data; an installation-time approval cannot automatically establish that a local assessment remains valid for a downstream task.
Existing skill scanners primarily ask whether a component contains suspicious behavior, while runtime monitors often ask whether the current call crosses a permission boundary. The paper examines a different problem: individual component assessments are accepted, but the joint result diverges from the user's actual objective. Such a failure need not involve an overtly anomalous call or a change to model weights; it can arise between factual interpretation, result aggregation, and decision semantics. Local compliance and end-to-end correctness are therefore distinct propositions, and adding the former does not establish the latter.
The study turns cross-component risk into testable conditions rather than relying on a concerning example alone. It requires individual review approval, an unsafe joint result, and disappearance of that result when any one modification is reverted, separating a compositional effect from a single-component problem. Core Idea: analyze joint behavior and component-reversion counterfactuals to test whether “each component passes review” supports the conclusion that “the entire workflow is safe.”
Method¶
Overall Architecture¶
SkillCascade is a research framework for constructing and validating stress-test cases, while SkillCascade-Bench contains the resulting verified and selected cases. This note covers threat conditions, evidence structure, and defense evaluation only; it does not provide skill modifications, generation prompts, attack-construction procedures, or appendix payloads. The contribution is evaluation and mechanism analysis, not training a new model or providing an attack-deployment guide, so the research stages are not depicted as an execution network.
The study first bounds the threat: two or more controlled skills from a third-party supply chain are co-installed and subsequently invoked in an ordinary user workflow. The attacker cannot change model weights, the system prompt, or platform mechanisms, manipulate user input or the user session, or modify skills outside its control. Co-installation is an important prerequisite; the paper discusses suites from a common vendor but does not measure how frequently attacker-controlled co-installation occurs in real marketplaces.
The framework uses five research roles: scenario discovery, candidate-case construction, independent scanning, test-case generation, and outcome judging. From a defensive auditing perspective, the key issue is not how these roles create candidate content but what evidence they establish: whether the task is realistic, components are independently approved, sandbox traces meet the evaluation criteria, and the same result persists after component reversion. The independent scanner sees one skill at a time; the outcome judge evaluates outputs and traces without receiving the attack plan, and human reviewers further check realism, joint dependence, and outcome evidence. These information boundaries reduce some evaluation contamination but cannot remove selection effects or correlated errors from shared models.
Key Designs¶
1. Compositional security conditions: separate component approval from workflow safety
“Individual stealth” is operationally defined by scanner approval, not by proof that a modification is harmless in every context. Let the evaluated modified system be \(\mathcal{A}^{*}\) and the system reverting modification \(i\) be \(\mathcal{A}^{*\setminus i}\); \(h\) measures outcome harm, \(\tau\) is the joint-harm threshold, and \(\varepsilon\) is a much smaller tolerated residual. The definition can be expressed through the following three conditions, retaining the source notation:
The first condition concerns component reviews, the second joint execution, and the third counterfactual outcomes after individual reversions. They answer different questions: observing a joint failure does not establish independent approval, and independent approval does not establish that harm depends on the entire combination. Indispensability requires reverting a modification rather than deleting the entire skill; deletion could merely disable legitimate functionality and obscure attribution. Likewise, a workflow with three skills does not necessarily contain three modifications; the appendix includes unmodified rendering components.
The definition provides a system-level auditing target rather than showing that a particular scanner can never improve. Every judgment remains conditional on the task, harm criterion, and scanner; approval by one tool cannot be elevated into a universal security certificate. Defenders need an evidence chain from the user objective to the final output, with explicit boundaries on the validity of each local assessment.
2. Cascade patterns: distinguish sequential dependence from aggregation dependence
A causal cascade depends on execution order: an earlier result must enter the shared context before a later step can use it for further interpretation. A compositional cascade emphasizes relatively independent contributions that jointly create the final problem, without requiring a fixed order of appearance in the theoretical account. A hybrid cascade combines sequential propagation and convergence of multiple contributions and cannot be reduced to either pattern alone. “Compositional security” is the umbrella concept, whereas “compositional cascade” is only one of the three patterns; “causal” also does not mean rigorous causal identification has been completed.
Shared context makes these distinctions operationally important: security review must establish not only what a component produces but whether downstream components treat it as a fact, assessment, or authorization. Sequential dependence calls for checking inherited premises across steps; aggregation dependence calls for checking whether the final result violates global constraints; hybrid dependence requires both. Placing multiple skill texts in a single query does not guarantee that a reviewer reasons about these relationships. “Seeing more content” and “understanding the consequences of its composition” are therefore different capabilities.
Pattern labels also require caution. For one compositional example, the appendix emphasizes independent contributions but also describes a downstream component depending on upstream fields, suggesting that actual cases can combine dataflow dependence with contribution convergence. A robust classification should distinguish execution order, field dependence, and contributions to harm rather than infer complete exchangeability from the authors' label alone. This note does not repeat the modification details; it retains only this limitation of classification and evidential interpretation.
3. Verification and selection: treat the successful-case bank as a conditional stress test
The benchmark contains 213 cases across 10 domains, seven objectives, and three patterns, with an average workflow length of 3.25 skills. Objectives cover output corruption, decision manipulation, functional denial, quality degradation, scope drift, information exfiltration, and unauthorized action. These are categories of evaluated outcomes, not skill contents or executable behaviors; their shared purpose is to assess end-to-end harm to the current user request. Cases undergo automated outcome judgment followed by independent review by three experts, and only unanimous cases enter the benchmark.
Human review checks joint dependence, workflow realism, and whether outputs satisfy the evaluation criteria; candidate-review agreement is Fleiss's \(\kappa=0.82\). Each case is executed three times for each target-system configuration, and the mean success rate across those runs is reported. A separate human check of the outcome judge covers only 30 judgments and reports Cohen's \(\kappa=0.86\); this supports a limited agreement check rather than domain-wide accuracy guarantees. The judge uses each case's own criteria, so overly narrow criteria or checks tied to particular wording can diverge from the harm that matters to the user.
The appendix states that the default limits for four feedback budgets are all 3, and total project API expenditure is approximately USD 8,000, including preliminary and failed experiments. These figures describe selection conditions and resource context, not adjustment or replication instructions. Successful cases undergo budgeted selection, making the reported rate a conditional result from rerunning verified cases rather than the failure probability of random skill suites. The paper supplies no co-installation denominator sufficient to estimate marketplace incidence, and total API expenditure does not establish per-case cost.
4. Behavior-composition review: assess the joint consequences of cross-skill commitments
The authors propose a Behavior-Composition Scanner as a proof-of-concept defense. It summarizes what behavior each skill commits to under particular conditions, reasons about the joint consequences of those commitments, and assesses whether a hypothetical scenario would produce an unsafe overall result. Unlike simple joint-text review, its object of analysis shifts from whether a component looks problematic to what happens when the local commitments hold together. This does not require reproducing any original attack content in this note and does not constitute a complete protection scheme.
The direction can be understood as generative counterfactual auditing: checking constraints that the system should preserve end to end rather than classifying isolated text. The paper reports a reduction of more than 30 percentage points in scanner approval for causal cascades, with smaller improvement for compositional cascades and still weaker improvement for hybrid cascades. The text cache lacks the figure's complete numerical values, so precise detection rates or between-pattern gaps cannot be reconstructed. The authors explicitly position the scanner as a proof of concept; greater sensitivity than existing checks does not mean compositional security has been solved.
Key Experimental Results¶
Main Results¶
The following table reproduces all model configurations from source Table 2; ASR is the fraction of cases judged to achieve the evaluated objective, with higher values indicating greater vulnerability. All values are percentages; the final column retains the displayed source row averages rather than replacing them with recalculations.
| Model | OpenClaw ASR | Claude Code ASR | Codex ASR | Source Avg |
|---|---|---|---|---|
| GPT-5.4 | 82.6 | 78.4 | 80.9 | 80.6 |
| GPT-5.4 Mini | 100.0 | 95.8 | 94.3 | 96.7 |
| Claude Opus 4.6 | 78.5 | 74.2 | 76.8 | 76.5 |
| Claude Sonnet 4.6 | 91.7 | 89.2 | 88.5 | 89.8 |
| Gemini 2.5 Pro | 90.2 | 86.4 | 88.9 | 88.5 |
| Gemini 2.5 Flash | 94.1 | 91.3 | 92.6 | 92.7 |
| Qwen 3 72B | 99.1 | 96.7 | 95.4 | 97.1 |
| Kimi K2.5 | 95.2 | 92.4 | 93.1 | 93.6 |
Reported-versus-tabulated check: the text reports a global ASR of 89.4%; the equal-weight mean of the 24 displayed table values is approximately 89.43%, also rounding to 89.4% at one decimal place. The mean of the displayed values agrees with the authors' report, but those values are not raw per-run data. The rounded table cannot replace original run statistics, so this consistency check is limited to the precision of the displayed values.
Ablation Study¶
The next table combines pre-deployment review results from source Tables 3–4; SR is the approval fraction, not detection accuracy. Per-skill review uses skill submissions as its denominator, while joint review uses entire combinations, so subtracting across granularities does not directly measure defense improvement. \(\Delta\) retains the source's modified-minus-original percentage-point difference; rejection of original skills is not a false-positive rate established against complete ground truth.
| Review granularity | Scanner / mode | Original SR (%) | Modified SR (%) | \(\Delta\) (percentage points) |
|---|---|---|---|---|
| Per-skill | Skill-Scanner / Static | 96.8 | 94.3 | -2.5 |
| Per-skill | SkillScan / Static | 95.4 | 92.1 | -3.3 |
| Per-skill | Skill-Scanner / LLM | 91.3 | 86.6 | -4.7 |
| Per-skill | Skill-Vetter / LLM | 94.2 | 90.5 | -3.7 |
| Per-skill | AIG-Scanner / LLM | 93.0 | 89.2 | -3.8 |
| Joint | Skill-Scanner / Static | 82.7 | 82.7 | 0.0 |
| Joint | SkillProbe / Hybrid | 86.0 | 81.7 | -4.3 |
| Joint | Joint-LLM / LLM | 88.2 | 83.0 | -5.2 |
The following results come from source Table 5 and apply only to runtime-defense evaluation on OpenClaw + GPT-5.4-mini. DER is the fraction of cases not detected by a runtime defense; it is not residual ASR after deploying that defense and should not be conflated with SR.
| Runtime defense | Monitoring paradigm | DER (%) |
|---|---|---|
| DefenseClaw | Static and dynamic policy | 90.6 |
| ClawGuard | Per-call rule enforcement | 88.9 |
| HITL Defense | Risk-tiered staged gating | 91.8 |
| AgentDoG | Trajectory-level inspection | 82.5 |
The indispensability analysis examines three-skill workflows: reverting one modification leaves 13–18% residual ASR, and reverting two reduces it to single digits. This supports joint dependence for most harm but does not establish that every run meets the definition's requirement of nearly eliminated residual harm. The text cache does not provide complete per-pattern figure values, so a more detailed ablation table cannot be constructed.
Key Findings¶
- Model row averages span 76.5–97.1%, while the same model varies by fewer than 6 percentage points across hosts; model differences are larger than host differences on this selected benchmark.
- Both individual and joint scanners still approve most modified samples, but this does not show that all future joint defenses will fail or constitute a formal test of statistical indistinguishability.
- Mean DER across the four runtime defenses is approximately 88.5%; AgentDoG's 82.5% is the lowest value in this table, not evidence that it is best across all systems and tasks.
- Reversion experiments leave nonzero residuals, and the Behavior-Composition Scanner offers only partial improvement; security conclusions should retain both non-ideal findings.
Highlights & Insights¶
- Changing the security unit: the paper separates component approval from joint outcomes. For third-party skill suites, component-review records cannot replace workflow-level evidence.
- Value of counterfactual verification: reverting modifications while preserving skill functionality helps establish whether a failure depends on composition. Residual success rates also caution against conflating a strong definition with its empirical approximation.
- Joint visibility is not joint understanding: a reviewer can see multiple skills and still miss their overall consequences. Behavioral commitments, data provenance, and end-to-end constraints warrant more defensive attention than simple concatenation.
Limitations & Future Work¶
- Selection bias and extrapolation: only successful, unanimously approved cases enter the benchmark, and evaluation uses sandboxes; 89.4% is neither a real-world incident rate nor a risk estimate for arbitrary skill combinations.
- Limited judge coverage: a human check of 30 judgments cannot establish reliability for every domain, objective, and model. Stratified blind review, confidence intervals, and harm criteria independent of specific output wording would strengthen the evidence.
- Gap between definition and verification: single reversion leaves 13–18% residual ASR; counterfactual results should be reported per case and modification, distinguishing indispensable modifications from unmodified supporting components.
- Ambiguous pattern classification: contribution independence and field dependence can coexist. Future cases should describe explicit dataflow relationships and harm-aggregation rules rather than rely on labels alone.
- Incomplete defense evaluation: the Behavior-Composition Scanner lacks a complete numerical table and systematic comparisons of legitimate-task utility, false-positive cost, and workflow length; superiority or full coverage cannot be claimed.
Related Work & Insights¶
- vs Skill-Inject / SkillJect: these studies primarily examine risks at the individual-skill level; this paper emphasizes multiple modifications acting together and outcomes after individual reversions, focusing on compositional security judgments rather than a local payload.
- vs STAC and tool-chain research: the source distinguishes its setting from studies that leave components unchanged and primarily alter calls or messages. Individual threat models still require paper-by-paper verification; this note adopts the proposed comparison dimension without treating the source's survey generalization as a universal fact.
- vs SkillProbe / Joint-LLM: these approaches provide cross-skill visibility but retain high approval rates on this benchmark; the Behavior-Composition Scanner additionally reasons about joint commitments, showing a useful direction rather than a mature solution.
- Defensive research direction: prioritize the provenance and validity bounds of intermediate assessments, preservation of critical facts through workflows, and final-result consistency with user objectives rather than increasing the number of component approvals.
Rating¶
- Novelty: 4/5 — Combines local review, joint harm, and indispensability into an explicit research object.
- Experimental Thoroughness: 3/5 — Broad cross-system coverage, but successful-case selection, judge-check size, and defense-utility evaluation constrain the conclusions.
- Writing Quality: 3/5 — A clear central argument, with table references, pattern boundaries, and correspondence between strict definitions and residual outcomes needing clarification.
- Value: 4/5 — Provides evidence and direction for workflow-level security auditing without a complete defense guarantee.