AI

OpenAI calls Astra its most aligned model ever. It won’t show its reasoning

Adrian Kessler

Astra operates across code, security analysis, computer use, and multi-step professional tasks with a speed and accuracy that outranks OpenAI’s previous model, Sol, and Anthropic’s Fable on standard software-engineering benchmarks. The company describes it as the first model it has built that crosses the “Critical” tier of cybersecurity capability: it can identify and develop exploits for previously undiscovered security flaws, at a scale and speed that no human team can match.

The mechanism behind this is a technique called opaque recurrence. Standard language models produce their reasoning as readable text — the chain of thought that researchers use to audit decisions, catch errors, and verify that a model is behaving as intended. Astra works without producing that text. It can solve hard problems while leaving no readable trace of the process. OpenAI’s chief scientist, Jakub Pachocki, acknowledged the shift: more capable models, he said, can “perform harder tasks using fewer language tokens.” Monitoring becomes harder in proportion.

That creates a structural problem that OpenAI’s own claims make visible. The company describes Astra as its “most aligned model” to date. Alignment — the engineering discipline concerned with making AI systems behave the way their designers intend — has historically relied on reading what the model is doing step by step. A model that reasons invisibly makes its own alignment unverifiable by that method. OpenAI has not described in public what alternative method it uses to verify Astra’s behavior before deployment.

Greg Brockman, OpenAI’s president, described Astra as a “generational leap” and was asked directly whether it represents artificial general intelligence. His answer: the contractual AGI definition that governed OpenAI’s partnership with Microsoft no longer applies. AGI is now, he said, a “spiritual concept.” That framing matters. OpenAI’s Microsoft agreement included clauses tied to AGI — conditions under which the terms of the relationship would change. Describing AGI as undefined keeps those clauses dormant, regardless of what Astra is actually capable of.

Access launched first to Daybreak, OpenAI’s vetted cybersecurity program — a structure the company frames as defensive-first deployment. The reasoning is that organizations hunting for vulnerabilities in their own systems get the capability before those hunting vulnerabilities in others’. That logic holds only as long as access controls hold. Astra’s zero-day identification capabilities are, by design, precisely what a well-resourced attacker would want. A model that reasons without a readable chain of thought is also harder to constrain or audit after the fact if something goes wrong.

Broader access rolls out to paid tiers — Pro, Plus, Enterprise, Business — and via the API in the coming week.

Tags: , , , , ,

Discussion

There are 0 comments.