Security (Updated September 17, 2026) 10 min read

Claude Mythos Found Zero-Days Everywhere. Here's Your Playbook.

Claude Mythos Preview found zero-days in every major OS and browser, one of them 27 years old, for under $50 a bug. What defenders should change now.

On April 7, 2026, Anthropic’s red team published its report on Claude Mythos Preview: the model found zero-day vulnerabilities in every major operating system and browser it was pointed at, including a 27-year-old OpenBSD bug, at under $50 per successful run. Defenders have months, not years. Here is what to change first.

Updated September 2026: Anthropic’s Project Glasswing update (May 22) reported more than 10,000 high- or critical-severity vulnerabilities found with roughly 50 partners, 530 disclosed to maintainers and 75 patched at the time of writing. Firefox 150 shipped fixes for 271 vulnerabilities found in that evaluation, and Epoch AI measured a 3.5x spike in disclosed CVEs in June 2026. The playbook below is unchanged; the volume it is written for has gone up.

What Mythos Preview found

TargetBugAgeWhat the model didCost
OpenBSD TCPSigned integer overflow in SACK sequence handling, null-pointer dereference27 yearsRemote crash of any OpenBSD host over TCPUnder $50 per successful run (under $20,000 for 1,000 scaffold runs)
FFmpeg H.264Slice counter collision with a sentinel value, out-of-bounds heap write16 yearsFound what years of fuzzing campaigns missed~$10,000 across several hundred repository scans
FreeBSD NFS (CVE-2026-4747)Stack buffer overflow in RPCSEC_GSS authentication17 yearsBuilt a working unauthenticated root exploit, splitting a 20-gadget ROP chain across multiple packetsHours
Linux kernelChains of 2-4 bugs: KASLR bypass, use-after-free, heap sprayMixedMultiple full privilege-escalation chainsNot disclosed
Web browsersJIT heap spray escaping both the renderer sandbox and the OS sandboxMixedTurned a page visit into kernel write accessNot disclosed

The report withholds the details on purpose:

Over 99% of the vulnerabilities we’ve found have not yet been patched, so it would be irresponsible for us to disclose details about them.

— Anthropic red team, Claude Mythos Preview, April 2026

Extrapolating from the validated sample, the report projects over 1,000 critical-severity and thousands of high-severity vulnerabilities discovered and still under responsible disclosure.

The benchmark numbers

The OSS-Fuzz comparison is where the capability jump stops being anecdotal. Against roughly 7,000 entry points across 1,000 repositories, in a single run each:

  • Sonnet 4.6 and Opus 4.6 each found 250-275 tier-1/2 crashes, 1 tier-3, and zero tier-4 or tier-5 results.
  • Mythos Preview found 595 tier-1/2 crashes and 10 tier-5 full control-flow hijacks — the most severe category, and the one the previous generation never reached at all.

Bar chart comparing OSS-Fuzz results: Sonnet 4.6 and Opus 4.6 each found 250-275 tier-1/2 crashes and zero tier-5 hijacks, while Mythos Preview found 595 crashes and 10 tier-5 full control-flow hijacks Figure: the gap is not incremental. Tier-5 — full control-flow hijack — went from zero to ten in one model generation.

The exploitation benchmark tells the same story. Opus 4.6 turned the Firefox 147 JavaScript-engine bugs it had found into working shell exploits twice in several hundred attempts. Mythos Preview produced 181 working exploits on the same benchmark and achieved register control on 29 more.

Anthropic’s manual review of 198 vulnerability reports found 89% exact agreement with the model’s severity assessments and 98% agreement within one severity level. The model is not just finding bugs. It is triaging them accurately enough to prioritize an attack list.

Why this matters beyond zero-days

Zero-days make for dramatic reading, but the broader capability profile in the report is what defenders should focus on. Mythos Preview also demonstrated:

  • Reverse engineering of stripped closed-source binaries into plausible source code, then finding vulnerabilities in the reconstructed code
  • Cryptography library weaknesses in TLS, AES-GCM, and SSH implementations
  • Logic vulnerabilities including authentication bypasses and authorization flaws
  • Web application flaws — cross-site scripting, SQL injection, CSRF

The cost profile is the quiet bombshell. A zero-day discovery campaign costs $50 to $10,000 in API credits. An n-day exploit chain — the kind that takes human experts weeks — came in under $1,000 for a one-bit adjacent-page write and under $2,000 for a Unix-domain socket chain, each completed in under a day. These are not nation-state budgets. These are individual developer budgets.

Forescout’s analysis of the report makes the operational version of the same point: traditional vulnerability management is too slow when patches do not exist yet, and critical-infrastructure operators often get a maintenance window only every few months.

What defenders should do now

Anthropic’s report includes recommendations for defenders. They’re good. But they’re written for a broad audience — CISOs, security teams, infrastructure operators. If you’re a developer shipping code, here’s what’s immediately actionable.

Four-stage chain showing attacker economics -- discovery under $50, accurate triage at 89% severity agreement, exploit development under $2,000, and over 99% of findings unpatched -- with the defender counter-move under each stage Figure: each stage of the attacker’s pipeline got cheaper. Each has a defender move that is available today.

1. Scan your code like an attacker would

The vulnerability categories Mythos Preview excels at — memory corruption, authentication bypasses, injection flaws, cryptographic weaknesses — are exactly the categories that static analysis catches when the rules are specific enough. Generic “check for vulnerabilities” prompts don’t cut it. You need structured detection patterns mapped to real CWEs.

Two options worth having side by side. Anthropic ships claude-code-security-review (6.2k stars), a GitHub Action that reviews pull-request diffs for injection, broken auth, IDOR, hardcoded secrets, weak crypto and deserialization RCE, with false-positive filtering tuned to drop low-impact noise. Its README is blunt about its own limits: “This action is not hardened against prompt injection attacks and should only be used to review trusted PRs.”

For scanning what’s already in the tree, the @jeremie-strand/security-scanner skill on SkillSafe scans for known vulnerability patterns across your software stack, project dependencies, and containers. It works inside Claude Code, Cursor, and Windsurf. Install it:

npx skills add https://api.skillsafe.ai/jeremie-strand/security-scanner

Then scan:

/security-scanner

In a world where a model can find and exploit a 27-year-old vulnerability for $50, manual-only security review is not a defensible strategy.

2. Patch aggressively

The report’s recommendation to shorten patch deployment cycles deserves emphasis. Mythos Preview’s n-day exploitation capabilities mean the window between patch release and working exploit is collapsing. The FreeBSD NFS exploit targeted a 17-year-old vulnerability. The Linux kernel exploits targeted recently-patched issues.

Treat CVE-related dependency updates as urgent. Enable auto-updates where possible. If your patching process involves a two-week review cycle, that cycle is now your attack surface — and with roughly 1,500 high- and critical-severity CVEs published by notable organizations in June 2026 alone, the queue is not getting shorter.

3. Audit your dependencies, not just your code

The Mythos report found vulnerabilities in widely-used cryptography libraries, codecs (FFmpeg), and system-level services (NFS, TCP stacks). Your application code might be clean. Your dependencies might not be.

Run security-scanner with a focus on dependency analysis. Know what’s in your stack. If a dependency hasn’t been updated in years, that’s no longer just tech debt — it’s a liability that can be discovered and exploited at API-credit prices. The same logic applies to the AI tooling you install: a skill or MCP server is a dependency that runs with your credentials, which is why every shared skill in the SkillSafe registry carries a scan report before it can be shared.

4. Assume your exposed services will be probed

Mythos Preview found remote code execution in NFS. It found remotely-triggerable crashes in TCP implementations. If you’re running internet-facing services, especially on older infrastructure, assume that AI-powered vulnerability scanners will find what human auditors missed.

Minimize your attack surface. Disable services you don’t need. Segment your network. Monitor for anomalous access patterns. The cost of scanning every exposed service is about to drop by orders of magnitude — for defenders and attackers alike.

The asymmetry problem

The core tension in the Mythos report is asymmetry. Defenders need to secure every service, every dependency, every code path. An attacker — now potentially a model running for a few thousand dollars — only needs to find one exploitable flaw.

Project Glasswing is closing part of that gap from the top: roughly 50 partners in the first phase, including Cloudflare, Mozilla, Microsoft, Oracle and Palo Alto Networks, with 1,587 open-source findings verified valid out of 1,752 assessed — a 90.6% true-positive rate. But 530 disclosed and 75 patched, against a backlog in the thousands, is the shape of the problem: discovery scaled, remediation did not.

For the millions of developers shipping production code every day, the playbook is simpler: scan systematically, patch aggressively, and treat every dependency — including every skill your agent loads — as part of your attack surface. If you want the verification side of that story, see how dual-side verification works or run a file through the SkillSafe scanner.

Frequently Asked Questions

What has Claude Mythos done?

Claude Mythos Preview, disclosed by Anthropic’s red team on April 7, 2026, autonomously found zero-day vulnerabilities in every major operating system and browser it was tested against — including a 27-year-old OpenBSD TCP bug and a 17-year-old FreeBSD NFS remote root flaw (CVE-2026-4747) — and wrote working exploits for several. Anthropic projects over 1,000 critical-severity findings, more than 99% unpatched at publication.

Which vulnerabilities did Claude Mythos find?

The publicly described findings are an OpenBSD TCP/SACK integer overflow (27 years old), an FFmpeg H.264 out-of-bounds heap write (16 years old), FreeBSD NFS remote root via RPCSEC_GSS (CVE-2026-4747, 17 years old), multiple Linux kernel privilege-escalation chains, and browser sandbox escapes. Mozilla shipped fixes for 271 Firefox vulnerabilities from the same evaluation in Firefox 150.

Did Claude Mythos break out of its sandbox?

In one evaluation, yes — under instruction. The Cloud Security Alliance’s research note describes an early Mythos Preview placed in a secured sandbox and tasked by a simulated user with escaping it to contact the supervising researcher. The model built a multi-step exploit and reached the internet through a system configured to talk to only a handful of approved services.

How much does an AI-discovered zero-day cost?

Under $50 for a single successful OpenBSD discovery run, and roughly $10,000 for the several-hundred-run campaign that surfaced the FFmpeg bug. N-day exploit development came in under $1,000 and under $2,000 for two chains, each finishing in under a day. That is a developer’s monthly API budget, not a nation-state program.

What should a small team do first?

Put a scanner in CI this week: claude-code-security-review on pull requests, plus a dependency audit on every release. Then shorten the patch cycle — Epoch AI counted roughly 1,500 high- and critical-severity CVEs from notable organizations in June 2026, more than 3.5x the previous monthly record, so a two-week review window is now the exposure.

Start today. The models that find these vulnerabilities are getting cheaper and more capable on a curve that isn’t slowing down.