Technology

Gemini 4 Argon is Google’s comeback model — and almost nobody can use it yet

Google DeepMind’s new model writes up to 1M tokens and tops Google’s own charts, but it goes first to vetted cyber defenders
Adrian Kessler
Add us on Google

For most of the past year, Google’s artificial intelligence story was about what it had not shipped. Gemini 4 Argon is the reply, and the most revealing thing about it is not a benchmark score. It is the guest list. Google DeepMind’s new frontier model goes first to a small group of vetted cybersecurity defenders through the company’s Fairwind Program, and to no one else.

Argon is built for long-horizon work: software engineering, legal and financial research, and cyber defense. The key mechanical change is output length: the model can now reason and write hundreds of thousands of tokens in a single pass, enough to carry a large code migration through to the end instead of stitching it together from fragments.

Google says thousands of its engineers already use it. In the announcement, the company describes Argon agents that freed more than 300 TiB of memory across its data centers and are porting C and C++ code to the memory-safe language Rust, up to the 800,000-plus lines of the Fuchsia Zircon kernel. On libgav1, Google’s open-source video decoder, Argon rewrote 32,000 lines of hand-tuned SIMD code into safe Rust that runs 2.7 times faster than the earlier Rust port.

Bar chart comparing Gemini 4 Argon and Gemini 3.8 Flash Cyber on vulnerability discovery and the Wiz penetration test benchmark
Image: Google

The output ceiling rises to 1 million tokens, up from 64,000. The benchmark numbers are Google’s own: 77.9% on DeepSWE v1.1, first place on the Vals Index of economically weighted knowledge work and 91.7% on LVBench for long-video understanding. A tally by VentureBeat counts Argon leading outright on 12 of 18 disclosed benchmarks, while OpenAI’s GPT-6 Astra still beats it by more than ten points on FrontierSWE v2. Until outside testers can run the model, these are claims, not verdicts.

The cybersecurity framing explains the gate. Google trained Argon to find, validate and patch vulnerabilities on its own, and will give trusted defenders a version with the cyber guardrails removed. Through Wiz’s Scan for Good initiative, the model flagged a critical flaw exposing personal data in hospital software that earlier models had missed. The uncomfortable mechanism is that a model this good at fixing holes is, by definition, good at finding them.

So the safeguards are the real story. Google says Argon refuses help with cyber and chemical, biological, radiological or nuclear attacks, is its most resistant model yet to prompt injection, and runs under a monitor that watches its chain of thought and halts execution when it oversteps the user’s intent. Google is also taking part in the U.S. government’s voluntary pre-release model access process.

Who can use Gemini 4 Argon today? Only the Fairwind cohort and Google’s own teams. There is no date for developers, businesses or consumers; access will start with paid API customers and Google AI Ultra subscribers. The introductory price is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterwards.

The model was announced on September 30 by Koray Kavukcuoglu, who took the top job at Google DeepMind after Demis Hassabis stepped down as chief executive in August. It is Google’s first new flagship since Gemini 3 in November 2025, and its real test is not a leaderboard but whether a model released without guardrails stays in the hands it was meant for.

Tags: , , , , , , ,

Add us on Google

Discussion

There are 0 comments.