viberg.tech

Gemini 4 Argon is the cheapest frontier model you cannot buy yet

Google has announced Gemini 4 Argon, priced well below its American rivals and given first to a closed group of cyber defenders. Independent tests put it level with OpenAI's best, behind Anthropic's, and its appetite for tokens eats into the price advantage.

Portrait of Sundar Pichai
Sundar Pichai, now Google's chief executive, at Mobile World Congress in Barcelona in February 2014. Photo: Maurizio Pesce, CC BY 2.0

Google announced Gemini 4 Argon on Wednesday and called it its most capable model so far. Almost nobody can use it. Argon is going first to a group of “trusted cyber defenders” in Google’s Fairwind Program, which gets a version without the usual cyber guardrails. Paying API customers and Google AI Ultra subscribers come next, as soon as possible, with no date given. Google also says it is taking part in the US government’s voluntary process for pre-release access while access widens.

The price is the headline

Google’s own numbers lean heavily on benchmarks, but the figure that matters for most businesses is the price. During an introductory period Argon costs $2 per million input tokens and $10 per million output tokens. After that it rises to $4 and $20. Cached input is 95% cheaper. VentureBeat worked out that the introductory price is a fifth of OpenAI’s GPT-6 Astra and half of Anthropic’s Claude Opus 5.5. At the standard rate Argon costs the same as Opus 5.5. Google has not said how long the introductory period lasts.

The other technical change is output length. Argon can produce up to a million tokens in one go, up from 64,000 in the previous Gemini. That matters less for chat than for agents that run long jobs, such as migrating a codebase, which Google says its own engineers already use it for.

Google says it leads. The independent tests disagree.

Google reports first-place scores on a long list of tests, including 77.9% on the DeepSWE software engineering benchmark against 74.2% for Opus 5.5. VentureBeat counted 18 published benchmarks and found Argon ahead on 12. It also found the gaps Google did not lead with: Astra is well ahead on two coding and science benchmarks, and Opus 5.5 leads on Terminal-bench by nine points.

The first independent results are less flattering. On Artificial Analysis’s Intelligence Index, which runs every model through the same set of tests, Argon scored 53, level with GPT-6 Astra and Anthropic’s Fable 5.1 and five points behind Opus 5.5. The Austrian site Trending Topics reports the same ranking and adds that on realistic work tasks Opus 5.5 is a long way ahead. Argon is 23 points better than Google’s previous model, Gemini 3.1 Pro Preview. It still does not put Google in front.

My reading is that Google has caught up and decided to compete on price, which is good news for buyers. The headline about retaking the lead says more about how benchmarks get chosen for a launch blog.

Cheap per token, hungry per task

There is a catch in the price, and it is the same one this blog described when frontier prices halved last week. What you pay depends on tokens per task as much as price per token. Artificial Analysis found that Argon used about 62,000 output tokens per task in its tests, against 27,000 for Astra, roughly 2.3 times as many.

At introductory prices Argon still comes out cheapest: about $1.99 per task, against $3.26 for Astra and $5.98 for Opus 5.5, according to Trending Topics. When the price doubles to the standard rate, the cost per task roughly doubles too and ends up above Astra’s. The saving is real for as long as the promotion lasts, and the promotion has no published end date.

Defenders first is now the norm

Argon’s release follows a pattern. In April Anthropic kept its Mythos model away from the public and gave it to a group of mostly American companies so they could fix their software before attackers caught up. Google is now doing something similar, and going further: Fairwind members get Argon with its cyber guardrails switched off. On the CWE-bench vulnerability test Argon ties with Astra at 68%.

The case for this is sound. Models that can find and fix holes in software can also be used to exploit them, and giving defenders a head start is better than giving everyone the same tool on the same day. The problem is who counts as a defender. Analysis by FourWeekMBA points out that Google has published no membership criteria or application process for Fairwind, and has named only one member. A private company is deciding who gets the most capable cyber tool it has built, and the decision sits outside any public process, in the US or in Europe.

For European companies that has a practical side. If the security vendors protecting you are not in the programme, the people looking for holes in your systems may get Argon-level tools, legitimately or not, before your defenders do. And the US pre-release review means Washington sees these models before European authorities do, the pattern this blog described when frontier models started needing a US sign-off.

What to do before it reaches your account

Do not plan on Argon yet. There is no release date, no public API model name, and no word on when the EU gets access. Anything you build today should run on a model you can actually call.

When it does arrive, test it on your own tasks and measure cost per completed task, not price per token. Argon’s token appetite means the bill can surprise you, and the introductory price will end.

Ask your security providers whether they are in Fairwind, or in Anthropic’s equivalent programme for Mythos. The answer tells you whether your defence is getting the newest tools or waiting with everyone else.

And use the competition. Google pricing a frontier model at half of Anthropic’s rate, and a fifth of OpenAI’s top tier, gives you something to bargain with in any renewal talk with your current provider this autumn, whether or not you ever switch.

Keep reading

All analysis →