Researchers have demonstrated “skill cascading” attacks in which several agent extensions look benign on their own but cause harm when executed together. A skill can include instructions, scripts, and reference files, making reusable agent ecosystems productive but also creating dependencies that single-package scanners may not understand.

In one example, separate medical skills weaken evidence that a drug was discontinued, downgrade interactions tied to it, and suppress the resulting low-priority alert. No individual change states the complete malicious objective, yet the chain removes a serious warning before it reaches a physician. The team built an automated red-teaming framework and a benchmark with 213 validated cascades across several domains.

Tests against representative agents and model backbones, including OpenClaw, Claude Code, and Codex, found that these cross-skill combinations could induce harmful behavior while avoiding existing per-skill checks and runtime monitors. Constructed benchmark attacks do not measure how often such chains occur in real marketplaces. They do show that signing or scanning each component separately cannot establish the safety of the composed workflow. Defenses need to inspect data flow, shared state, ordering, and the combined effect of every skill an agent loads.