Anthropic's Claude Mythos finds 23,019 vulnerabilities in first month of Project Glasswing
Claude Mythos discovered 23,019 vulnerabilities across 1,000+ open-source projects with 90.6% confirmation rate in Project Glasswing's first month.
Claude Mythos Finds 23,019 Vulnerabilities in First Month
Anthropic’s Project Glasswing delivered dramatic results in its opening month. According to the first-month report from May 22, 2026, Claude Mythos identified 23,019 vulnerabilities across 1,000+ open-source projects, with 90.6% confirmed as real on independent sampling.
Project Glasswing: A Defensive Cybersecurity Initiative
Anthropic launched Project Glasswing on April 7, 2026, distributing Claude Mythos Preview to roughly 50 partner organizations—including AWS, Apple, Google, Microsoft, NVIDIA, CrowdStrike, and JPMorgan Chase—exclusively for defensive cybersecurity work.
Prompt Engineering Breakthroughs Drive AI Model Performance
Across the industry, structured prompting techniques continue to unlock significant performance gains. Eight well-written examples in a prompt were enough for a 540-billion parameter model to beat fine-tuned GPT-3 on the GSM8K math benchmark. A 2024 study by OpenAI showed that structured prompts produce output preferred by human annotators in 73% of cases compared to unstructured prompts on free-form writing tasks.
Advanced reasoning frameworks show even more dramatic improvements. In ‘Game of 24’, GPT-4 with standard Chain-of-Thought solved only 4% of problems, while Tree of Thoughts achieved 74%. ReAct beats imitation learning and reinforcement learning methods by 34% in absolute success rate on ALFWorld with one or two in-context examples.
Self-Refine demonstrates an average 20% performance improvement across 7 different tasks with GPT-3.5, ChatGPT, and GPT-4 without training or supervised data.
Practical applications confirm these gains. A test of two follow-up email prompts on 50 real clients in 2025 found the structured prompt (CREATE + Chain-of-Thought) got a reply 64% of the time versus 19% for the basic prompt.
Frameworks and Tools Mature
Anthropic put context engineering in writing on their engineering blog in ‘Effective context engineering for AI agents’. Stanford released DSPy in late 2023 and by 2026 it has reached version 2.x with the MIPROv2 optimizer. DSPy shows 10–40% quality improvements compared to manual prompting on RAG pipelines and classifiers.
Gemini 3.5 Flash Goes GA
Google brought Gemini 3.5 Flash to general availability on May 19 at Google I/O 2026. The model beats Gemini 3.1 Pro at roughly 4x the speed on coding and agentic benchmarks. API pricing stands at $1.50 per million input tokens and $9.00 per million output tokens.
Google also announced Gemini 3.5 Pro at I/O, with Sundar Pichai stating ‘give us until next month to get it to you’. The model is confirmed for June but with no specific date provided.
Anthropic Roadmap Signals
A Sonnet 4.8 source map accidentally shipped with @anthropic-ai/claude-code npm v2.1.88 on March 31, 2026, containing source filter lists with the strings sonnet-4-8, opus-4-7, and mythos. Polymarket closed at 3% on a Sonnet 4.8 ship date of May 24, 2026.
Grok 5 Still Training
xAI’s Grok 5 is still training on Colossus 2, expanded from 1 GW to 1.5 GW in April. Reported specs include ~6 trillion parameters MoE, 1.5M context, and native multimodal. Polymarket contract for Grok 5 public release by June 30, 2026 collapsed from 68 cents in February to 12% by early April and currently sits in the 12–33% range.
Source: WaveSpeed Blog
Developments since publication
-
Anthropic announced Project Glasswing on April 7, 2026, bringing together Amazon Web Services, Anthropic, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, N Source
-
Claude Mythos Preview has found thousands of high-severity vulnerabilities, including some in every major operating system and web browser. Source
-
Mythos Preview found a 27-year-old vulnerability in OpenBSD that allowed an attacker to remotely crash any machine running the operating system. Source
-
Mythos Preview discovered a 16-year-old vulnerability in FFmpeg in a line of code that automated testing tools had hit five million times without catching the problem. Source
-
Mythos Preview autonomously found and chained together several vulnerabilities in the Linux kernel to allow escalation from ordinary user access to complete machine control. Source
-
Mythos Preview scored 83.1% on CyberGym's Cybersecurity Vulnerability Reproduction benchmark, compared to Opus 4.6's 66.6%. Source
-
Anthropic is committing $100 million in usage credits for Mythos Preview across Project Glasswing efforts and $4 million in direct donations to open-source security organizations. Source
-
Anthropic will donate $2.5 million to Alpha-Omega and OpenSSF through the Linux Foundation, and $1.5 million to the Apache Software Foundation. Source
-
Mythos Preview will be available to participants at $25/$125 per million input/output tokens via Claude API, Amazon Bedrock, Google Cloud's Vertex AI, and Microsoft Foundry. Source
-
A significant number of scanned AI infrastructure hosts had been deployed without authentication in place, with authentication not enabled by default in many projects. Source
-
Of 5,200+ Ollama servers queried without authentication prompts, 31% answered with connected models accessible. Source
-
Exposed agent management platforms n8n and Flowise instances revealed instances without authentication exposed to the internet, with one Flowise instance exposing entire business logic of an LLM chatb Source
-
Intruder identified over 90 exposed instances across government, marketing, and finance sectors with chatbots, workflows, prompts, and outward access openly accessible. Source
-
Intruder found repeated insecure patterns including insecure defaults, misconfigured Docker setups, hardcoded credentials, applications running as root, no authentication on fresh installs, and embedd Source
-
Anthropic announced Claude Mythos Preview and Project Glasswing, a consortium of technology companies formed to find and patch security vulnerabilities in critical software. Source
-
Anthropic committed up to 100 million USD in usage credits and 4 million USD in direct donations to open source security organizations. Source
-
Mythos autonomously found thousands of zero-day vulnerabilities across every major operating system and web browser, including a 27-year-old bug in OpenBSD and a 16-year-old bug in FFmpeg. Source
-
When AISLE tested Mythos's showcase vulnerabilities on small, cheap, open-weights models, eight out of eight models detected Mythos's flagship FreeBSD exploit, including one with only 3.6 billion acti Source
-
A 5.1B-active open model recovered the core chain of the 27-year-old OpenBSD bug in AISLE's testing. Source
-
AISLE has been running a discovery and remediation system against live targets since mid-2025, discovering 15 CVEs in OpenSSL, 5 CVEs in curl, and over 180 externally validated CVEs across 30+ project Source
-
The Intruder team scanned just over 2 million hosts using certificate transparency logs and identified 1 million exposed services. Source
-
The AI infrastructure scanned by Intruder was more vulnerable, exposed, and misconfigured than any other software they have ever investigated. Source
-
A significant number of exposed AI infrastructure hosts had been deployed out of the box with no authentication in place. Source
-
Of 5,200+ Ollama API servers queried that listed a connected model, 31% answered without requiring authentication. Source
-
Of all frontier models identified across Ollama servers, 518 were wrapping well-known frontier models from Anthropic, Deepseek, Moonshot, Google, and OpenAI. Source
-
Within a couple of days of lab work, Intruder researchers found arbitrary code execution in one popular AI project. Source
-
Intruder identified over 90 exposed instances of agent management platforms (n8n and Flowise) across sectors such as government, marketing, and finance. Source
-
Some AI infrastructure projects exhibit poor deployment practices including insecure defaults, misconfigured Docker setups, hardcoded credentials, and applications running as root. Source
-
On a basic security reasoning task, small open models outperformed most frontier models from every major lab in AISLE's testing. Source
-
AISLE describes AI cybersecurity capability as 'jagged' because it does not scale smoothly with model size, model generation, or price. Source
-
Addy Osmani, a Google Chrome engineer, introduced the term 'Loop Engineering' in June 2026 Source
-
Anthropic reported that over 80% of its engineers are already using self-improving loops Source
-
Anthropic's engineers stated that over 80% of Anthropic's engineers are already using self-improving loops, and this will reach 100% within 3 to 6 months Source
-
Loop Engineering removes humans from directly commanding AI line by line, making the system itself the loop rather than humans Source
-
At the April AI Engineer Conference, Anthropic's engineers compared two approaches for Claude developing a retro-style mini-game app: minimal prompts took 20 minutes and cost $9, while an iterative ag Source
-
Jason Wei and colleagues at Google Research proved that eight well-written examples in a prompt were enough for a 540-billion parameter model to beat fine-tuned GPT-3 on the GSM8K math benchmark in th Source
-
The paper 'Tree of Thoughts: Deliberate Problem Solving with Large Language Models' (Princeton/Google DeepMind, May 2023) showed that in 'Game of 24', GPT-4 with standard Chain-of-Thought solved only Source
-
The paper 'ReAct: Synergizing Reasoning and Acting in Language Models' (ICLR 2023) showed that on ALFWorld, ReAct beats imitation learning and reinforcement learning methods by 34% in absolute success Source
-
The paper 'Self-Refine: Iterative Refinement with Self-Feedback' (NeurIPS 2023) showed average 20% performance improvement across 7 different tasks with GPT-3.5, ChatGPT and GPT-4 Source
-
Tom Brown and the OpenAI team published 'Language Models are Few-Shot Learners' in 2020, showing that given between 10 and 100 examples of the desired output type, GPT-3 matched or exceeded fine-tuned Source
-
Anthropic published 'Effective context engineering for AI agents' on their engineering blog in 2026 Source
-
The paper 'Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks' by Patrick Lewis and the Facebook AI team (NeurIPS 2020) is the foundation for RAG, where the model queries a document data Source
-
Stanford team published results showing 10-40% quality improvements with DSPy compared to manual prompting on RAG pipelines and classifiers Source
-
Anthropic's '2026 Agentic Coding Trends Report' states that 2026 is the year of the transition from single-agent workflows to multi-agent ones Source
-
A test of two prompts by an individual (referred to in the article) on 50 real clients in 2025 showed the second prompt (using CREATE + CoT) got a reply 64% of the time, while the first prompt got a r Source
-
Full adoption of Loop Engineering is expected within 3–6 months, driven by efficiency gains and reduced dependence on manual intervention Source