Drooid Logo
Back to story perspectives

Full Breakdown

AI Frontier Models Break Cybersecurity Benchmarks, Prompt Urgent Defense Shift

5/15/2026, 8:28:51 PM

Breakthrough Test Results

The UK AI Security Institute (AISI) tested Anthropic’s Claude Mythos Preview and OpenAI’s GPT-5.5. Both models solved the “The Last Ones” range (6/10) and the previously unsolved “Cooling Tower” range (3/10) within a 2.5 M-token limit.

Evolution and Partnerships

Mythos Preview launched in April 2026 as part of Anthropic’s Project Glasswing with AWS, Microsoft, Google, NVIDIA, CrowdStrike and others. OpenAI’s GPT-5.5-Cyber entered testing under its Trusted Access for Cyber program. AISI’s forecast of a 4.7-month doubling cycle was again outpaced.

Key Figures and Organizations

Anthropic and OpenAI develop the frontier LLMs. AISI runs the benchmarks. Palo Alto Networks supplies test data. Pentagon cyber-policy chief Katherine Sutton highlighted defensive uses.

Performance Metrics

‘The Last Ones’ solved 6/10, ‘Cooling Tower’ 3/10. AISI caps tasks at 2.5 M tokens. Palo Alto Networks found 75 issues linked to 26 CVEs across 130 products, versus its typical five-CVE monthly rate. AISI now estimates the 80 % reliability horizon doubles roughly every four months.

Why It Matters

The models significantly shrink attack cycles to minutes, enabling near-real-time vulnerability discovery and exploit generation. They also promise rapid defensive automation, but the speed threatens to outpace traditional patching and response processes.

Official Statements & Responses

AISI said the models “outperformed” but warned benchmarks are inadequate; Sutton argued Mythos can “identify and remediate vulnerabilities in minutes to seconds,” urging integration; Klarich called outcome “the light at the end of the tunnel” for security.

Criticism & Opposition

Experts note the 2.5 M-token cap likely understates true capability, creating large error bars. Rapid vulnerability discovery could overwhelm patch cycles, and the same technology could be weaponized by adversaries.

Conflicting Reports & Gaps

AISI notes that token caps likely understate the true capabilities of frontier models, widening success-rate error bars. Long-term reliability on extended tasks remains untested, and real-world defensive efficacy is still uncertain.

Verbatim Quotes

  • “The newer Mythos Preview checkpoint completed both our cyber ranges, solving the range 'The Last Ones' in 6 of 10 attempts and the previously unsolved 'Cooling Tower' in 3 of 10 attempts,” — AISI blog authors
  • “I hear a lot of people talking about challenges and threats when they talk about Mythos,” — Katherine Sutton, Assistant Secretary for Cyber Policy, U.S. Department of Defense
  • “This is the light at the end of the tunnel,” Klarich said in the blog.” — Lee Klarich, Chief Product and Technology Officer, Palo Alto Networks
  • “Cybersecurity teams are going to be under a lot of pressure, for sure.” — Katie Moussouris, Founder and CEO, Luta Security
  • “understates what frontier models can do,” — AISI

What’s Next

The 2026 Cyber Summit on May 21 will host briefings from Sutton, AISI researchers, and industry partners, providing a forum for policymakers to discuss the implications. Anthropic is seeking a $30 billion funding round to expand Claude infrastructure, while OpenAI plans further GPT-Cyber refinements.