Why Scanning Architecture Matters: Comparing Skill Registry Security
Three registry security models - reactive moderation, install-time scanning, and dual-side verification with a tree hash - and the attack each one misses.
AI skill registries run one of three security architectures: reactive community moderation, install-time-only scanning, or dual-side verification with a cryptographic tree hash. What separates them is not scanner quality. It is when and where the check runs, which decides whether a tampered artifact or a novel payload is ever caught.
Updated September 2026: the architectural comparison below is unchanged, and the evidence for it has grown. Unit 42 documented five more malicious ClawHub skills found between February and May 2026 - after the registry’s post-incident controls were in place, because every one of those controls runs after publication.
This is a technical comparison, not a competitive review. The numbers come from the four independent audits of the ClawHavoc campaign, which is the largest public test any of these models has had: Antiy CERT catalogued 1,184 malicious skill packages across 12 author IDs, and Snyk’s ToxicSkills study found that 36.82% of 3,984 skills carried at least one security issue.
The three models at a glance
| Model | Where the check runs | What it proves | What defeats it |
|---|---|---|---|
| Reactive moderation | After publication, after a user complains | Someone eventually noticed | Moving faster than the community. ClawHavoc ran ~6 days before the first audit |
| Install-time scanning | On the consumer’s machine, on the bytes delivered | The delivered file matches no known-bad signature today | A tampered artifact, a novel variant, or a payload fetched after the scan |
| Dual-side verification | Before sharing AND after download, plus a hash comparison | The artifact is byte-identical to the one the publisher scanned | Scanner evasion at authoring time; composition-level attacks |
The Three Models
Model 1: Reactive Community Moderation
How it works: Users flag suspicious skills. Above a report threshold, the skill is auto-hidden or sent for manual review. Some registries run periodic scans after flagging.
Examples in practice: ClawHub used this model before ClawHavoc. Three community reports triggered auto-hide.
What it catches:
- Skills that have already been flagged by at least one user
- Behaviors so obviously malicious that users recognize them quickly
What it misses:
- Everything before the first user notices and reports
- Sophisticated attacks designed to look entirely legitimate
- Novel techniques with no existing reports
- The window between upload and first install - which can be minutes for targeted attacks
The structural problem: The system is entirely reactive. A malicious skill has already been distributed and potentially installed before any protection activates. The attacker’s winning condition is simply: move faster than the community notices. In ClawHavoc, Bitdefender attributed 354 malicious packages to a single actor, and roughly six days passed between the first poisoned upload on January 27, 2026 and the first public audit. A community reporting system is not designed for that rate. The full account is in our ClawHavoc post-mortem.
Model 2: Install-Time-Only Scanning
How it works: When a user installs a skill, the registry or install toolchain scans the downloaded files against detection rules or known-malware databases. If the scan flags something, the install is blocked or warned.
Examples in practice: Skills.sh integrated Snyk, Socket and generative AI analysis for install-time scanning as of February 2026; we looked at that registry’s API surface separately in skills.sh and the API supply chain.
What it catches:
- Known malware signatures present in the downloaded files
- Declared dependency vulnerabilities
- Behavioral patterns matching known malicious code
What it misses:
The tamper window. Install-time scanning checks the artifact as received. It cannot detect whether that artifact matches what the publisher originally submitted. If the registry is compromised, the CDN is tampered with, or the distribution pipeline is attacked between publish and install, the scanner will analyze the tampered version - and may pass it as clean, because the modified artifact hasn’t been seen before.
Novel variants. Snyk and Socket maintain databases of known-bad packages. AMOS-based malware (the payload in ClawHavoc) is actively developed to evade signature detection. The specific AMOS variant used in ClawHavoc was not in VirusTotal’s database at the time of initial distribution. Novel variants pass signature checks for days or weeks before detection databases catch up - and Snyk’s own audit shows how much a signature-first view misses: 91% of the confirmed malicious skills combined code with prompt-injection text, which no malware signature describes.
Deferred payload delivery. Some attacks don’t embed malware in the skill archive at all. The skill contains instructions that cause the AI agent to download and execute an external payload during normal use. Install-time scanning sees a clean archive. The malicious payload is fetched later, entirely outside the scan window.
The fundamental problem: Install-time scanning verifies one artifact at one moment against one set of signatures. It cannot establish that the artifact is identical to what the publisher submitted, and it cannot catch attacks that defer payload delivery past the scan.
Model 3: Dual-Side Verification with Cryptographic Tamper Detection
How it works: Security checks run on both sides of the distribution pipeline.
-
Publisher side: Before sharing, the publisher runs a local security scan. The scan report and a SHA-256 cryptographic tree hash of the archive are stored on the registry. Skills with critical findings cannot be shared - the gate is enforced before the skill is published to the index.
-
Registry: The tree hash is stored as an immutable record at publish time.
-
Consumer side: When installing, the consumer independently re-scans the downloaded files. The server compares the consumer’s scan report against the publisher’s original report, and the consumer’s computed hash against the stored hash.
Only if both reports agree (or diverge within acceptable bounds) and the hashes match does the install receive a verified verdict. A hash mismatch produces an automatic critical verdict and blocks installation - no human review required, no database update needed.
What it catches:
Registry-level tampering. If the stored archive is modified after publish - by a compromised admin, a breach, or an infrastructure attack - the hash comparison fails immediately. The consumer’s computed hash will not match the stored hash. Installation is blocked before any human reviewer notices anything is wrong.
CDN and transit tampering. If the distribution network delivers a modified file, the hash comparison catches it regardless of what the consumer’s scan result is. The cryptographic check is independent of detection databases.
Publisher-side attack staging. Because the publisher-side scan runs before the skill can be shared, skills containing instructions to download external payloads, access credential files, or execute obfuscated code are blocked at the source. The attack cannot reach the registry index at all.
What it doesn’t guarantee:
- A publisher who crafts a skill specifically to evade the scanner’s detection rules - scanner evasion is possible in any static analysis system
- Composition-level attacks that use entirely normal-looking operations in combination to achieve malicious outcomes
Figure: the same pipeline, three architectures. Only the third one has a check upstream of publication and a comparison across the two ends.
Attack Vectors: Side-by-Side
| Attack Vector | Reactive | Install-time | Dual-side |
|---|---|---|---|
| Known malware | ❌ (depends on reports) | ✅ (signature match) | ✅ (publisher-side blocked) |
| Novel malware variant | ❌ | ❌ (not in DB yet) | ✅ (behavioral rules catch staging) |
| Registry infrastructure breach | ❌ | ❌ (scans tampered file) | ✅ (hash mismatch) |
| CDN/transit tampering | ❌ | ❌ (scans delivered file) | ✅ (hash mismatch) |
| Fake prerequisite / external payload fetch | ❌ | ⚠️ (depends on analysis depth) | ✅ (publisher scan catches staging) |
Credential file access (~/.ssh, .env) | ❌ | ⚠️ (may catch) | ✅ (publisher scan + behavioral rules) |
Prompt injection in .md instructions | ❌ | ❌ (not a code scanner target) | ✅ (dedicated prompt injection rules) |
Agent persona/memory persistence (SOUL.md) | ❌ | ❌ | ✅ (publisher scan) |
| Social engineering (ClickFix-style UX) | ❌ | ❌ | ⚠️ (partially detectable) |
| Typosquatting | ❌ | ❌ | ✅ (verified namespace enforcement) |
The Cryptographic Guarantee
The core distinction between install-time scanning and dual-side verification comes down to what each approach proves about the artifact.
Install-time scanning proves: “This artifact, as I received it right now, does not match any known-bad signatures in the current database.”
Dual-side verification with tree hashing proves: “This artifact is byte-for-byte identical to the artifact the publisher submitted. The publisher’s pre-share scan passed. The consumer’s post-download scan independently agrees. Any deviation - in the archive, in transit, in the registry - would have broken the chain and blocked installation.”
These are categorically different security properties. The first is a probabilistic check against an external database that may or may not include the specific threat. The second is a cryptographic binding between the submitted and delivered artifact - a SHA-256 comparison over a 2^256 output space - with independent verification at both ends that does not depend on any external database being current.
Figure: two independent scans and one hash comparison. A mismatch is a critical verdict with no human in the loop.
This distinction is most consequential in three scenarios that have all occurred in adjacent ecosystems:
-
Registry and account compromise. npm’s
ua-parser-jsis the canonical case: three versions (0.7.29, 0.8.0 and 1.0.0) were published with embedded malware on October 22, 2021, rated CVSS 8.8 as CVE-2021-4229, with the advice that “any computer that has this package installed or running should be considered fully compromised”. Install-time scanning of the compromised artifact provides no protection when the artifact is the published one; a publisher-side gate is what has to fire. -
Pipeline and CDN tampering. Codecov’s Bash Uploader was modified in cloud storage after an attacker extracted a GCS HMAC key from a Docker image layer. The compromise was discovered on April 1, 2021 and disclosed on April 15 - and it was found because a customer compared the SHA256 published on GitHub against the SHA256 they computed locally. That is precisely the check dual-side verification performs automatically on every install.
-
Sophisticated novel malware. As AI agents become high-value targets, attackers invest in evasion. Snyk confirmed 76 malicious payloads in its ClawHub and skills.sh scan, 8 of which were still live at publication; the ClawHavoc AMOS variant was absent from VirusTotal at distribution time. A skill built to avoid known signatures is exactly the case where behavioral publisher-side scanning - which targets what the skill does, not what it looks like - provides durable protection.
The practical formulation: install-time scanning tells you the artifact looks clean today. Dual-side verification tells you the artifact is what the publisher submitted and both parties agree on its contents. For a security-sensitive distribution channel, the second property is the one that matters. The mechanics are documented in how dual-side verification works.
Why Install-Time Scanning Is Still Worth Having
This is not an argument that install-time scanning is worthless. It provides genuine value:
- Defense in depth: Multiple independent checks catch more than any single check alone
- Known-malware database coverage: For previously documented malware families, signature matching is fast and reliable with well-maintained databases
- Dependency vulnerability detection: Tools like Snyk excel at identifying vulnerable declared dependencies - a different threat class from malicious skill behavior
The argument is that install-time scanning alone leaves structural gaps that grow more dangerous as the ecosystem matures and AI agents become more autonomous and more credentialed. The same layering logic runs through NIST’s Secure Software Development Framework and SLSA’s build levels: provenance and integrity are separate requirements from vulnerability scanning, and neither substitutes for the other.
What a Fourth Model Could Look Like
The three models above operate at the artifact level: they analyze what a skill contains before and after distribution. A fourth model - currently at the frontier of registry security research - would operate at the runtime level: enforcing what a skill is actually permitted to do during execution.
This approach would look something like a declared permission manifest:
# skill.permissions.yaml (hypothetical)
network:
outbound: false
filesystem:
read: ["./"] # project directory only
write: ["./output/"]
shell:
allowed: false
agent:
persona_modification: false
memory_write: false
At install time, the skill’s declared permissions are cryptographically bound to the archive. At runtime, the AI agent enforces the manifest - refusing to execute any operation the skill did not declare. An attempt to read ~/.ssh/ by a skill that declared filesystem.read: ["./"] would be blocked at the agent level, not the scan level.
This model shifts the guarantee from “we scanned this and it looked clean” to “this skill is structurally incapable of performing undeclared operations”. It’s analogous to the evolution from antivirus scanning to OS-level sandboxing in mobile app ecosystems - a transition that took years but materially changed the threat model.
Nothing in the vendor documentation contradicts the need for it. Anthropic’s own guidance on Agent Skills is blunt about what an installed skill can do:
a malicious Skill can direct Claude to invoke tools or execute code in ways that don’t match the Skill’s stated purpose
Several practical challenges remain: defining a permission grammar expressive enough to cover legitimate skill behaviors without being so broad it provides no constraint; getting AI agent runtimes to enforce the manifest consistently; handling skills that legitimately need broad access (e.g., a deployment helper that needs SSH and network access). These are solvable engineering problems, not fundamental blockers.
Runtime enforcement via declared permissions would not replace dual-side verification - it would complement it. Publisher-side scanning and cryptographic tamper detection remain necessary because they operate before the skill ever runs. Runtime enforcement adds a third layer: even if a malicious skill somehow reaches the consumer’s machine, it cannot exceed its declared permissions.
This is Phase 3 thinking. The industry isn’t there yet. But the trajectory from reactive moderation to dual-side verification to runtime permission enforcement is the logical progression as AI agents become more capable and the consequences of a compromised skill grow more severe. The same argument applies to connectors and servers, which we cover in the MCP security guide.
The Practical Implication for Developers
When you install a skill from a registry that uses install-time-only scanning, you are trusting:
- That the registry has not been compromised since the skill was published
- That the delivery infrastructure has not been tampered with
- That the current detection database covers the specific attack technique in the skill
- That the attacker did not defer payload delivery past the scan window
When you install a skill from a registry using dual-side verification, you additionally verify:
- That the archive you received is byte-for-byte identical to what the publisher submitted
- That the publisher’s pre-share scan passed before the skill was made available
- That any tampering in transit would be caught by hash mismatch before installation completes
The extra verification takes seconds. The guarantee it provides is structural rather than probabilistic - and structural guarantees hold regardless of whether the attacker anticipated your defenses.
Frequently Asked Questions
How do I check whether an AI skill is safe before installing it?
Look for a scan report that was produced before publication, a recorded hash of the archive, and a namespace tied to an authenticated publisher. If the registry only scans at install time, you are trusting that nothing changed between publish and download. You can run any GitHub-hosted skill through the SkillSafe scanner without an account.
What is dual-side verification?
The publisher scans the skill and stores a SHA-256 tree hash before sharing; the consumer re-scans after downloading and recomputes the hash; the server compares both. Matching hashes plus agreeing reports produce a verified verdict. A hash mismatch is an automatic critical verdict that blocks installation, with no human review step.
Does install-time scanning catch a compromised registry?
No. It scans the bytes it receives, so a tampered artifact is scanned in its tampered form and can pass. Codecov’s 2021 breach was caught only because a customer compared the published SHA256 against their own computed one - a comparison dual-side verification performs on every install.
What share of public AI skills carry security issues?
In the largest public audit to date, Snyk scanned 3,984 skills across ClawHub and skills.sh and found 36.82% carried at least one security issue, 13.4% a critical-severity one, and 76 confirmed malicious payloads. Antiy CERT’s historical count of one registry reached 1,184 malicious packages from 12 author IDs.
Is a signed skill the same as a verified skill?
No. A signature proves who published an artifact; it says nothing about what the artifact does. Provenance systems such as Sigstore and SLSA answer “did this come from the claimed build”, while a pre-share scan answers “does this read credentials or fetch a payload”. A registry needs both.
Conclusion
Security architecture is not a product feature. It is a set of commitments about what an adversary would have to do to defeat your defenses - and whether those commitments are enforced by code or depend on human reaction speed.
Reactive moderation requires attackers to be slow. Install-time scanning requires attackers to be unsophisticated or their malware to be old. Dual-side verification with cryptographic tamper detection requires attackers to simultaneously compromise the publisher, evade static analysis, and defeat a hash comparison that the consumer computes locally against an immutably stored reference. The latter is a meaningfully harder problem.
The ClawHavoc campaign demonstrated what happens when a high-velocity attack meets a system that depends on human reaction speed. As AI agents accumulate more credentials, more file access, and more autonomous capability, the cost of getting registry security architecture wrong increases proportionally.
The minimum viable security model for a skill registry today includes publisher-side scanning, cryptographic tamper detection, and independent consumer verification. Install-time scanning is a useful addition to that baseline, not a substitute for it. Runtime permission enforcement is the next evolution - and the ecosystem should be building toward it now, before the attack surface matures faster than the defenses.
SkillSafe’s scanner ruleset is publicly documented at skillsafe.ai/security. We publish this comparison to make the architectural tradeoffs concrete - not to position our implementation as the final word on registry security.
Related reading: Supply chain posts - ClawHavoc post-mortem - MCP security guide