Cybersecurity

Claude Broke Into OpenAI in Days — and the Harder Questions Are Anthropic’s

Adrian Kessler
Add us on Google

A small team pointed Anthropic‘s newest model at OpenAI‘s front door and walked through it. The version most people are reading — clever outsiders embarrassing the maker of ChatGPT — gets the subject wrong. The break-in did not hinge on anything OpenAI forgot to patch. It hinged on the moment Anthropic shipped a more capable Claude. That is the part worth sitting with: a job that had stalled for days solved itself once a new flagship went live.

The people involved were not criminals. Researchers at the security startup Hacktron worked inside OpenAI’s own bug-bounty program, chained together a couple of unglamorous web weaknesses, reached the accounts of OpenAI employees, and got as far as an internal code repository — then stopped, filed what they found, and took a modest award. They left a harmless marker behind as proof and went no further. Read as a crime, it is a non-event. Read as a capability test, it is a flare going up.

Here is the mechanism the wire coverage skated past. The team first ran the job on Claude Opus 4.8, the version tuned for security work, and it stalled across repeated sessions. Then Anthropic released Opus 5. “Opus 4.8 struggled across several sessions to produce a working exploit,” the Hacktron team wrote. “Within hours of Opus 5’s release, we gave it the same problem and it succeeded.” VentureBeat reported the newer model wrote a working exploit for an awkward chip target within hours. Nothing about OpenAI’s defenses changed in that window. The model did.

That is precisely the kind of jump — a leap in what an ordinary user can do with an off-the-shelf subscription — that Anthropic’s own safety machinery is built to detect and hold back. The company keeps its most capable cyber model, Mythos, behind export controls and a short list of vetted defenders, and calls it the strongest security model it has built. None of that mattered here. The publicly available tier was enough. As Gray Swan chief executive Matt Fredrikson put it to TechCrunch, “for $200 a month, anyone can use these tools and hack into a company like OpenAI.”

This is where the accountability settles on Anthropic rather than its rival. Anthropic has built its whole identity on caution — its chief executive, Dario Amodei, speaks openly about the “real dangers” of the technology and has pressed the industry to slow down. A general-release model that measurably lowers the bar for breaking into a competing lab is an uncomfortable counterweight to that message. Asked to account for it, the company said nothing; CBS News reported that Anthropic did not respond to a request for comment. OpenAI, for its part, thanked the researchers, narrowed the permissions the attackers had abused, and revoked the affected sessions.

The specifics are almost mundane. The way in was an outdated image-processing library bundled with the Discourse software that runs OpenAI’s community forum; a second slip in how sign-in tokens were scoped let the researchers ride a compromised session into employee accounts for ChatGPT and Codex, and on to private code. The entire chain, from first foothold to repository access, took under three days. OpenAI paid $6,500. The team says it spent less than $3,000 of Claude time across the whole effort — two months of work, most of it human, compressed at the finish into a machine that no longer needed a specialist.

The scarce resource that used to protect a target was never the patch. It was the expert who could find the flaws, weld them together and see the path through — and the months that took. That expertise just got cheap. This time the team that reached OpenAI’s code filed a report and took the bounty. The next one to clear the same bar, on the same subscription, is under no obligation to.

Tags: , , , , ,

Add us on Google

Discussion

There are 0 comments.