Cybersecurity

OpenAI’s hacking AI completes 95% of exploit chains — the standard model does 1.5%

Susan Hill

GPT-5.6-Cyber found a high-severity flaw in Chrome’s V8 JavaScript engine during internal testing — one that had not been publicly disclosed. It also identified flaws in mobile operating systems, databases, and kernel software, all previously unknown. OpenAI‘s new model is not a defensive tool. It is built to hack.

The model was trained on GPT-5.6-Sol, OpenAI’s current general-purpose model, then retrained specifically for offensive security work. On OpenAI’s internal Cyber Capability Evaluation benchmark, GPT-5.6-Cyber completed 95% of exploit chain requests — tasks involving privilege escalation, authentication bypass, and arbitrary code execution. The standard version of GPT-5.6 completes 1.5% of the same requests. The previous specialized model, GPT-5.5-Cyber, scored 57.3%.

That jump — from 57.3% to 95% in one generation — is the number that defines this release. It means the gap between what an AI can do and what a skilled human attacker can do closed by roughly 40 percentage points in under a year.

GPT-5.6-Cyber’s specific capabilities include identifying known vulnerabilities, building exploit chains from scratch, and finding zero-days: flaws no one has documented or patched yet. That last capability is the sharpest break from any prior AI security tool. Finding zero-days takes specialized expertise and significant time. GPT-5.6-Cyber can now do it on demand.

OpenAI is not releasing the model publicly. Access runs through Daybreak Red, the higher tier of the company’s Daybreak cybersecurity program. Initial partners include Accenture, IBM, CrowdStrike, Cloudflare, Palo Alto Networks, Sophos, Capgemini, EY, KPMG, and PwC — large consulting and security firms that will use the model to scan clients’ infrastructure before attackers find the same flaws.

What the access controls do not answer is the harder question: OpenAI now decides who gets to use an AI capable of finding vulnerabilities in any software. The vetted list is large firms and top-tier consultancies. Individual security researchers, smaller teams, and academic institutions are not included in the initial rollout. The Chrome V8 flaw found during testing is under coordinated disclosure with Google — the company has not said how many other vulnerabilities the model identified before launch.

OpenAI said it plans to expand the Daybreak Red program to additional vetted partners in the coming months.

For the software industry, the math is blunt: if an AI can scan a codebase for zero-days at a 95% completion rate, every security team will eventually need one — or will need to assume adversaries already have access to one.

Tags: , , , , ,

Discussion

There are 0 comments.