AI

OpenAI declared AGI. The benchmark they cited put the score 37 points lower

Adrian Kessler

Jensen Huang’s post on X ran to one sentence and a congratulation: “From ChatGPT to o1 to Astra in 4 years, AGI has arrived. Congratulations @OpenAI team.” He sent it on September 6. Three days earlier, Greg Brockman, OpenAI’s president, had used almost identical framing with the press: “Welcome to the AGI era.” The declarations came close together, from two men with complementary interests, and they were received largely as confirmation.

The benchmark beneath the claim was ARC-AGI-3. OpenAI said Astra, its new frontier model released under the GPT-6 designation, scored 99.9% on ARC-AGI-3 and 98% on FrontierMath Tier 4. Scores at that level, under standard conditions, would represent a real discontinuity — a model that doesn’t just perform well but lands in territory that earlier benchmarks treated as aspirational. The organization that built ARC-AGI-3 published a score from their standard evaluation harness: 62.7%.

Thirty-seven percentage points separate those two numbers. Both come from real evaluations. OpenAI’s internal harness and the benchmark creators’ reference harness are not the same protocol, and the same model produces different scores under each. What the ARC-AGI-3 organization uses as its reference condition — the one it considers valid for comparing models across the field — produced 62.7%. OpenAI’s internal setup produced 99.9%. The difference between those numbers is the difference between a strong benchmark result and a declaration of arrival.

The chip sale behind the AGI congratulation

Astra was trained on NVIDIA chips. Huang sells NVIDIA chips. When a CEO publicly congratulates a customer for clearing a historic threshold — a threshold that makes the hardware used to build that product the defining infrastructure for what comes next — the announcement is, structurally, a commercial claim. Not a conspiracy. A structural reality: the conflict of interest was visible, the framing was about arrival rather than performance, and most of the coverage received both without comment.

The chip market for AI infrastructure moves on narrative. If Astra is AGI, then whatever trained Astra becomes the hardware of the post-AGI world — and the companies buying chips for the next training run will be buying toward that benchmark. Huang’s congratulations were, in that specific sense, also a product statement. The statement depended on a number that the benchmark organization itself measured differently.

What independent verification shows — and what it doesn’t

As of publication, Astra does not appear in the Artificial Analysis Intelligence Index or in Arena.ai’s rankings — the two most-watched independent evaluation systems for frontier models. Independent verification for major frontier releases typically arrives within days. The absence here is not evidence of poor performance. It means the external validation that would normally accompany a claim of this magnitude has not materialized yet.

The 62.7% from the benchmark creators is not a refutation. A score at that level on ARC-AGI-3 — a test designed to resist the optimization tricks that inflated earlier benchmarks — is genuinely strong. ARC-AGI-3’s predecessors were saturated: models absorbed training data resembling the test questions and pushed numbers up without improving at novel reasoning. ARC-AGI-3 changed the question design to require abstraction over pattern retrieval. Sixty-two point seven percent on that test, under the reference harness, is far above anything that existed two years ago.

The researchers who raised objections were precise about this. The model’s capability was not their target. The jump from performance to declaration was. “Scores well on these benchmarks” and “AGI has arrived” require a step in between: a stable, agreed definition of AGI against which arrival can be measured. No such definition exists across the field. OpenAI has an internal one. By OpenAI’s own definition, Astra may qualify. But a company measuring its product against a definition it wrote is a different kind of result than one measured by an external standard.

The sentence that didn’t come with a definition

Huang’s post had no caveats. “AGI has arrived” is a declarative — not “by some measures” or “under certain definitions,” but a clean arrival. The four-year progression he named — ChatGPT, o1, Astra — is real as a trajectory of capability. Each model in the sequence extended what was possible. Framing that sequence as a journey with an endpoint, and announcing the endpoint, makes a different claim: it says arrival is distinguishable from the journey, that there is a line rather than a slope, and that Astra crossed it.

The benchmark number behind that claim was 62.7%. The people who built the test published it. The question it raises — whether the AGI declaration is a measurement or a market move — was there on September 6. The coverage largely answered it without asking.

Tags: , , , , ,

Discussion

There are 0 comments.