Mid-tier AI Models Now Rival Frontier Systems for Hacking

Mid-tier AI models now rival top-tier in hacking tasks, raising attack scale risk

Mid-tier AI models are rapidly closing the gap with frontier systems in offensive cybersecurity tasks. Recent XBOW research highlights that AI models such as GPT 5.5 and several open-weight alternatives have achieved significant advances in autonomous web application testing, even without source code access. This development raises urgent security questions as these models’ lower costs may lead to a surge in automated attacks, particularly against small to medium-sized business web assets.

XBOW Research: Mid-tier AI Models’ Surprising Leap in Hacking Capabilities

Earlier this week, XBOW released its mid-year 2026 AI model security research report, drawing attention to a growing class of AI models that are neither the most advanced ‘frontier’ models nor basic earlier systems. This ‘middle class’ includes proprietary and open-source models such as Z.ai’s GLM-5.2, xAI’s Grok 4.5, Anthropic’s Opus 4.7, and Meta’s Muse Spark 1.1. All of these demonstrated strong performance in hacking and exploitation tasks previously considered the domain of top-tier models.

According to XBOW’s head of AI, Albert Ziegler, “It’s not even that the open-source variants or not quite frontline competitors are catching up as such. It’s that they are crossing a certain threshold, which means that suddenly they are providing net value at a cheaper price.” This shift is notable because, only six months ago, these models struggled with moderately complex, multi-step tasks. Now, they can successfully complete them, marking a transformative change in offensive security capabilities.

One of the key factors is cost. These mid-tier models are much cheaper to operate, allowing users to run them repeatedly or in parallel, which was not economically viable with earlier or frontier models. This makes it feasible to scale up offensive security efforts, potentially allowing attackers to automate vulnerability discovery and exploitation at volume.

Technical Findings: GPT 5.5 Sets a New Baseline in Web App Exploitation

GPT 5.5, considered a near-frontier model, stands out for its exceptional performance in exploitation benchmarks. According to XBOW’s findings, the leap between GPT 5 and 5.5 represents one of 2026’s most significant advances in autonomous web application testing. These models now excel in both ‘white box’ (with source code) and ‘black box’ (without source code) scenarios, a critical distinction for real-world attackers who rarely have code access.

  • GPT 5.5 achieved a vulnerability miss rate of just 10 percent, compared to GPT 5’s rate of 40 percent.
  • The model performed even better in black box testing, surpassing previous versions that relied on source code.
  • These improvements mean attackers can use affordable models to probe for weaknesses by interacting live with websites, rather than analysing code alone.

The XBOW report emphasises the significance of models working effectively without source code. “What translated into findings was the ability to reach and prove a vulnerability against the running system, not to infer it from a pattern in the source,” the researchers noted. This shift makes exploitation more accessible and practical for a larger pool of threat actors.

Another critical insight from XBOW’s testing is that access to source code is now less important. Instead, the ability to interact dynamically with web applications or software is a stronger predictor of success. This development could put more organisations with internet-facing systems at risk, since attackers can now automate the testing of live environments without needing inside information.

Multi-Agent Attack Potential: Lower Costs Amplify Threats

Frontier models such as Mythos and GPT 5.6 remain more capable at some individual cybersecurity tasks, but their high operational costs limit widespread offensive use. In contrast, the affordability of mid-tier AI models enables new attack strategies. Adversaries can now:

  • Run multiple agents in parallel to probe web applications at scale.
  • Conduct repeated and persistent attacks without prohibitive costs.
  • Continuously refine their tactics based on model outputs, adapting faster than many defenders can respond.

Anthropic’s recent research into multi-agent swarms reinforces this concern. When multiple AI agents collaborate, they can identify vulnerabilities in open-source software projects far more quickly than a single model operating alone. For defenders, this means the speed and volume of AI-driven attacks could increase significantly in the near future.

Why This Matters: Security Risks for SMBs and Broader Ecosystems

The sudden improvement in mid-tier AI models’ hacking abilities poses a clear risk to small and medium-sized businesses. These organisations often rely on common web applications and may lack the resources to monitor or defend against sophisticated automated attacks. The combination of low-cost AI, autonomy and improved black box performance means more attackers can target more vulnerable assets with less effort and expense.

What Organisations Should Do Now

In light of these findings, organisations should:

  • Prioritise regular security testing of internet-facing web applications, including dynamic testing approaches.
  • Monitor emerging AI-driven attack techniques and update defences accordingly.
  • Work with cybersecurity partners who understand the latest AI exploitation trends.

Staying informed about advances in AI offensive tooling is critical, especially as the barrier to entry for automated attacks continues to drop.

Originally reported by cyberscoop.com.

Share this bulletin

About the Author

Headshot of Jonny Pelter, leading cyber security expert in the UK and CISO

Jonny Pelter

Partner

  • CIPM
  • CIPP/E
  • CISSP
  • CISM
  • CRISC
  • ISO27001
  • Prince2
  • MSc
  • BSc

Jonny Pelter

Jonny is a Founding Partner at CyPro and executive group level CISO who has worked closely with the British intelligence agencies NCSC and GCHQ.

An ex-professional rugby player and originating from KPMG and Deloitte, Jonny has a wealth of experience across numerous sectors including technology, critical national infrastructure, financial services, oil & gas, insurance, betting, pharmaceuticals and utilities.

Jonny is a leading cyber security expert in the UK, having featured on national media for his professional commentary such as BBC News, iPlayer, Telegraph and Times Radio.

View Profile
Back to Bulletins

Related CyPro Services

  • Managed Detection and Response (MDR)

    Managed Detection and Response (MDR) is an end-to-end managed service designed to help organisations detect, analyse and respond to cyber threats quickly and effectively. It...
    View Service
CyPro Cookie Consent

Hmmm cookies...

Our delicious cookies make your experience smooth and secure.

Privacy PolicyOkay, got it!

We use cookies to enhance your experience, analyse site traffic, and for marketing purposes. For more information on how we handle your personal data, please see our Privacy Policy.

Schedule a Call