ToxicSkills: What the First Large-Scale Agent Skill Audit Found
Snyk scanned 3,984 AI agent skills: 36.82% had a security flaw, 13.4% a critical one, 76 carried confirmed malware and 8 were still live at publication.
Snyk’s ToxicSkills report scanned 3,984 AI agent skills from ClawHub and skills.sh as of February 5, 2026. It found at least one security flaw in 1,467 of them (36.82%), a critical-severity issue in 534 (13.4%), and 76 confirmed malicious payloads. Eight were still publicly installable at publication.
Updated September 2026: the finding has since been corroborated at larger scale. OWASP now tracks Malicious Skills as AST01 in its Agentic Skills Top 10, citing a USENIX Security 2026 study that scanned 98,380 skills across public marketplaces and confirmed 157 malicious ones carrying 632 vulnerabilities. The guidance below is unchanged.
Until ToxicSkills, nobody had independently audited the agent skill ecosystem at scale. There were incident reports — the ClawHavoc campaign documented 1,184 malicious skills across 12 publisher accounts and made the threat concrete. A systematic scan of the full public corpus had not happened.
Key Figures
| Figure | What it measures | Source |
|---|---|---|
| 3,984 | Skills scanned from ClawHub and skills.sh, as of Feb 5, 2026 | Snyk ToxicSkills |
| 1,467 (36.82%) | Skills with at least one security flaw | Snyk ToxicSkills |
| 534 (13.4%) | Skills with at least one critical-severity issue | Snyk ToxicSkills |
| 76 | Confirmed malicious payloads after human-in-the-loop review | Snyk ToxicSkills |
| 8 | Confirmed-malicious skills still live on clawhub.ai at publication | Snyk ToxicSkills |
| 1,184 | Malicious skills in the earlier ClawHavoc campaign, across 12 accounts | OWASP AST01 |
| 98,380 / 157 | Skills scanned and malicious skills confirmed in the USENIX Security 2026 follow-up | OWASP AST01 |
| 49,592 | ClawHub skills evaluated by the SkillSieve triage framework, at $0.006 per skill | arXiv 2604.06550 |
The research team used human-in-the-loop review to validate automated findings. The 76 confirmed payloads are skills with deliberate backdoors, credential stealers, or data exfiltration logic embedded directly in the Markdown instruction files. They were not false positives or ambiguous patterns.
Snyk’s summary of how low the bar sits is one sentence:
The barrier to publishing a new agent skill on ClawHub? A
SKILL.mdMarkdown file and a GitHub account that’s one week old.— Snyk, ToxicSkills: Malicious AI Agent Skills in the ClawHub Supply Chain
A skill is that folder and nothing more — a SKILL.md plus any scripts and references it ships alongside — which is what separates it from an MCP server or a plugin.
The ToxicSkills Threat Taxonomy
Snyk organized their findings into eight threat categories. The breakdown shows what attackers are actually doing versus what developers typically worry about — the most common finding is not a backdoor, it is a skill that pulls untrusted web content into the agent’s context.
| Threat category | Severity | Share of scanned skills flagged |
|---|---|---|
| Third-party content exposure | Medium | 17.7% |
| Suspicious downloads | Critical | 10.9% |
| Hardcoded secrets | High | 10.9% |
| Direct financial access | Medium | 8.7% |
| Credential handling | High | 7.1% |
| Malicious code | Critical | 5.3% |
| Unverifiable dependencies | Medium | 2.9% |
| Prompt injection | Critical | 2.6% |
Figure: categories overlap, which is why the per-category rates sum to more than the 36.82% headline. The two largest bands are exposure patterns, not payloads.
Critical Severity
Prompt injection (2.6%) — Hidden instructions embedded in skill Markdown that redirect agent behavior. This includes base64-obfuscated payloads, Unicode smuggling, and “ignore previous instructions” patterns. These attacks exploit the gap between what the skill description says and what the agent actually receives and processes. The SkillJect research documented similar injection patterns achieving 97.5% attack success against Claude Code — this category is not theoretical. The same primitive is what the MCP security checklist exists to cover on the tool side.
Malicious code (5.3%) — Actual backdoors, credential stealers, and remote code execution payloads embedded in skill setup scripts or inline shell commands. The ClawHavoc campaign’s credential-stealing skills fell into this category: the skills looked legitimate and their .md files directed agents to download and execute external binaries.
Suspicious downloads (10.9%) — Skills that fetch executables or archives from unknown domains, GitHub releases from unfamiliar accounts, or password-protected ZIP files. The suspicious download pattern is a delivery mechanism, not a payload class on its own — it is how malicious code gets onto the machine without being embedded directly in the skill file where scanners can see it.
High Severity
Credential handling issues (7.1%) — Skills that instruct agents to echo API keys, embed credentials in commands, or ask users to paste secrets into agent outputs. These are not malicious by design, but they create exposure. A skill that logs your AWS credentials to confirm a configuration step has done real damage even if the author had no ill intent.
Hardcoded secrets (10.9%) — API keys, access tokens, and private credentials left directly in skill files. ToxicSkills pairs this with 1Password research it cites, in which 23% of organizations reported that their agents had been tricked into leaking credentials. Hardcoded secrets in installed skills are a direct path to that outcome.
Medium Severity
Third-party content exposure (17.7%) — Skills that fetch arbitrary web content, parse social media feeds, or clone external repositories and pass the results directly to the agent. This enables indirect prompt injection: a malicious actor controls a webpage that a skill fetches, and embeds agent instructions in that page’s content. The agent treats the fetched content as trusted input. The Cloud Security Alliance describes the same mechanism as agent context poisoning.
Unverifiable dependencies (2.9%) — Skills that load instructions from remote URLs at runtime (curl | bash equivalents, dynamic imports, remote Markdown files). The skill you installed is not the skill that runs — its behavior is determined by whatever URL it fetches at execution time.
Direct financial access (8.7%) — Skills with hardcoded access to trading platforms, crypto wallets, or payment systems. The risk here is obvious: a skill with credentials for a brokerage account that also contains a prompt injection vulnerability is a high-value target for attackers who can exploit the injection.
Why These Numbers Matter
The 36.82% figure is the headline, but the more important number is 13.4% critical. Snyk’s team explicitly tuned their detectors to minimize false positives on widely adopted legitimate skills. A 13.4% critical rate after conservative tuning suggests the real-world risk is higher, not lower, than reported.
There is a structural reason the ecosystem looks this way. Publishing an agent skill requires a Markdown file and a GitHub account that is a week old. There is no mandatory review, no security scan at upload time, no code signing. The barrier is set at the same level as early npm and PyPI — before those ecosystems learned hard lessons about what happens when popular packages get compromised.
The difference is that agent skills do not run in a sandboxed interpreter. They run inside AI agents that typically have shell access, filesystem read/write permissions to the project directory and often the home directory, access to environment variables and credential files, and the ability to send messages through connected integrations. As Snyk puts it in its companion analysis of the agentic attack surface: when you install a malicious skill, the agent becomes the attacker, with the developer’s full environmental access.
This is why the ClawHavoc campaign — which we analyzed in depth — was so effective. The malicious skills did not need an exploit. They just needed instructions that told the agent to read ~/.ssh/ and send the contents somewhere.
The 8 Skills That Stayed Live
One detail from the ToxicSkills report deserves more attention: at publication time, eight of the 76 confirmed malicious skills were still publicly available on clawhub.ai.
That is not a criticism of ClawHub’s response time — coordinated disclosure and takedown take time. The point is that the window between “attacker publishes malicious skill” and “malicious skill is removed” is a real window. Developers installing skills during that window get the malicious version. And if they installed it, they still have it: removal from the registry does not remove it from their machine.
Figure: takedown is the only control most registries have, and it acts at the far end of the window. A pre-install scan is the one that runs on day zero.
This is the core argument for pre-install scanning rather than post-publish moderation. Reactive removal is valuable, but it has an inherent latency. The ToxicSkills data gives that latency a face: at the moment of publication, 76 confirmed-malicious skills had been available, and 8 remained live after the research team’s disclosure process. Every developer who installed one of those 76 during their active window was exposed with no warning.
What to Do About It
Before installing any skill:
Run a scanner on it first. For GitHub-hosted skills, paste the repository URL into SkillSafe’s web scanner — it produces a structured report with severity ratings before you activate anything, no account required. If you are installing from a URL or a registry without built-in scanning, download the skill files and inspect them; at minimum, grep for subprocess, os.system, eval, exec, and outbound HTTP calls.
When evaluating a skill from a marketplace:
Check whether the registry performs pre-share scanning. The scanning architecture comparison breaks down what different approaches catch and miss. A registry that scans at publish time (once) is better than no scanning, but it does not catch updates that introduce malicious code after the initial review — the pattern TeamPCP used against LiteLLM, covered in when trusted packages turn hostile.
For skills already installed:
The ToxicSkills findings are retrospective. If you installed skills from ClawHub before February 2026 without verifying them, it is worth scanning them now. At a 13.4% critical rate, the prior probability that a randomly selected installed skill carries a critical finding is not negligible.
For skill authors:
The credential handling and hardcoded secrets categories represent unintentional risk, not malice — together they account for 18% of flagged skills. A skill that instructs an agent to print an API key for verification creates real exposure even if the author never considered the threat model. Scanning your own skills before sharing — SkillSafe requires a scan report before a skill can be shared — catches these issues while you can still fix them privately. Everything published through the SkillSafe registry carries one.
The Ecosystem Is at an Inflection Point
Snyk’s framing in the ToxicSkills report is worth repeating: the agent skills ecosystem is at the same inflection point npm and PyPI hit before security became a first-class concern. The patterns are identical — typosquats, malicious maintainers, post-install scripts as attack vectors. The difference is that the privilege model is worse. A compromised npm package can do damage; a compromised agent skill can do that same damage while impersonating your development agent.
The remediation path is also the same as it was for package ecosystems: mandatory signing, pre-distribution scanning, integrity verification at install time, and audit trails. Those mechanisms exist today for AI agent skills. The question is whether developers adopt them before the incident rate climbs to where it became impossible to ignore in package ecosystems.
Detection research is moving in that direction. SkillSieve, a three-layer triage framework published in April 2026, ran over 49,592 real ClawHub skills and reported an F1 of 0.929 at $0.006 per skill — evidence that scanning a whole corpus is now a budget line, not a research project.
The ToxicSkills data suggests the current trajectory is not favorable. 534 critical findings in under 4,000 skills is a high baseline for an ecosystem that is still in early growth. As the corpus grows into the tens of thousands, the absolute numbers will scale with it unless the security infrastructure does too.
Frequently Asked Questions
What percentage of AI agent skills contain security flaws?
36.82% in the largest published audit. Snyk’s ToxicSkills team scanned 3,984 skills from ClawHub and skills.sh in February 2026 and flagged 1,467 with at least one security issue. 534 of those (13.4%) carried a critical-severity finding, and 76 contained confirmed malicious payloads after manual review.
Are agent skills a bigger risk than npm or PyPI packages?
The exposure is worse for the same class of compromise. A malicious npm package runs inside an interpreter; a malicious skill runs inside an agent that already holds shell access, filesystem write permissions and your environment variables. Snyk’s framing: when you install a malicious skill, the agent becomes the attacker with the developer’s full access.
How do I check a skill for malware before installing it?
Scan the repository before activation. SkillSafe’s web scanner takes a GitHub URL and returns a severity-rated report without an account. If you are installing from an unscanned source, read the files first and grep for subprocess, os.system, eval, exec, remote curl calls and any URL the skill fetches at runtime.
What is the OWASP Agentic Skills Top 10?
An OWASP project cataloguing the ten highest-impact risks in agent skill ecosystems. Malicious Skills is AST01, the first entry, defined as skills that appear legitimate but hide credential stealers, reverse shells, backdoors or social-engineering instructions in SKILL.md prose. It cites both ClawHavoc’s 1,184 malicious skills and a USENIX Security 2026 scan of 98,380 skills.
If a malicious skill is removed from the registry, am I safe?
No. Delisting stops new installs; it does not touch a copy already on disk, and most registries send no notification. Eight of the 76 skills Snyk confirmed malicious were still live at publication, and every machine that installed one during the window still has it. Re-scan installed skills rather than assuming a takedown reached you.
Sources
- Snyk: ToxicSkills — Malicious AI Agent Skills in the ClawHub Supply Chain
- OWASP Agentic Skills Top 10, AST01 — Malicious Skills
- Cloud Security Alliance: SKILL.md and Agent Context Poisoning
- Snyk: Your AI “Skills” Are the New Agentic Attack Surface
- arXiv: SkillSieve — A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
- arXiv: Formal Analysis and Supply Chain Security for Agentic AI Skills
- ITSecurity.network: Agent Skills — The New Supply Chain Attack Vector
Figures cited from the Snyk ToxicSkills report unless otherwise attributed. SkillSafe did not independently replicate the scan.