Google’s launch post for Gemini 4 Argon reports eight benchmarks. It does not mention Terminal-Bench, GPQA, Humanity’s Last Exam or SWE-bench, and its only footnote is about pricing. Artificial Analysis, the independent evaluator most people check first, scores Argon 53 on its Intelligence Index and ranks it eighth of 223 models. Both documents are accurate. They are describing different parts of the same model.
My read: Argon is a real, spiky model, and the two things Google leans on to call it frontier, the choice of scorecard and the price, each need an asterisk. The model is not the problem. The framing is.
The scorecard Google chose
The headline number is 77.9% on DeepSWE v1.1, a long-horizon software engineering benchmark. Latent Space puts the comparators at 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra, a lead of about 3.7 points on a single suite that Google picked to lead with. The rest of the post leans on adjectives. Argon is “leading” on the Vals Index, on Vals Finance Agent v2, on Harvey’s legal benchmark and on Gray Swan’s prompt-injection test, and the text gives no score and no named competitor for any of them.
Where Google does print a number, the field is thin. On Zapier’s AutomationBench Argon takes first place at 51.3%, and on LVBench, a long-video test, it posts 91.7%. On CWE-bench v1, a security benchmark, it ties for first at 68%, tied with a prior model the post calls 3.8 Flash Cyber, which also scored 68%. A tie with the previous model is not much of a headline. Latent Space relays the claim that Argon is state of the art on 13 of 19 “credible” benchmarks, which is a fair count if you let the vendor choose the 19.
The claims that come with numbers elsewhere are less flattering. Google says Argon performs similarly well on Harvey’s legal agent benchmark. Latent Space lists Argon at 19.6% there against 25.42% for Muse Spark 1.2, a model that does not feature in Google’s framing at all.
None of this is dishonest. Every lab’s launch post is a curated scorecard, and Google’s is not worse than most. The point is what a curated scorecard cannot tell you, which is where the model loses.
The boards Google didn’t pick
Artificial Analysis has Argon at 52.6 against 52.7 for GPT-6 Astra, a dead heat, and against 57.6 for Claude Opus 5.5 and 56 for Claude Sonnet 5.5. Google’s previous Pro model scored 30 on the same index, so a 23-point jump is the real story of this release. It also puts Argon level with Astra and about five points behind two Anthropic models. One comparison site adds a caveat worth carrying: the Argon figures are vendor-reported, charted at high effort, while Opus 5.5 is charted at adaptive reasoning, max effort. The settings are not matched, and that cuts in both directions depending on how you resolve it.
On Terminal-Bench 4, the agentic terminal benchmark Google left out, Argon scores 57%, behind Sonnet 5.5’s 64%, per Implicator. On Arena.ai’s WebDev board it sits eighth with 1679 points, even as it is first on the Text Arena at 1525. Front-end work is exactly where Bloomberg’s sources point: Implicator summarizes Bloomberg’s reporting that unnamed Google employees say Argon struggles with some coding tasks, front-end design in particular. Google disputes it. Tulsee Doshi, who leads Gemini products, said Googlers have “put the model through its paces in recent weeks, with many relying on it for their hardest coding and research problems.”

I would not treat the unnamed-employee claim as established. It is a secondhand summary of an anonymous report, and the Terminal-Bench and WebDev numbers are doing the real work. But they point the same direction: on agentic and front-end coding, Argon trails the best Anthropic models, and the launch post puts DeepSWE where those results would have been.
The discount is doing the arguing
Artificial Analysis’ headline for Argon was that it matches Astra at roughly 60% of the cost per task. That figure is real and it is conditional. At the introductory rate of $2 per million input tokens and $10 per million output tokens, Artificial Analysis prices a task at $1.99. At the $4 and $20 list price Google says will apply once the introductory period ends, the same task costs $3.98. Astra costs $3.26 and Opus 5.5 costs $5.98. Argon at list is about 22% more expensive than the model it ties, and cheaper than Opus 5.5 only because Opus costs more.
The introductory discount is 50%, and Artificial Analysis describes it as running for “at least a month.” Google’s post names no end date, and Latent Space notes the same. Token appetite matters here too. Artificial Analysis measures Argon at about 62,000 output tokens per task against 27,000 for Astra, and its model page shows 110 million output tokens to run the full index against a median of 82 million for comparable reasoning models. A model that talks more costs more at any per-token price, and the discount is hiding how much more.
Latent Space reports a conflicting figure from Vals, which measures Argon at about a quarter of Sonnet 5.5’s output tokens. Both can be true. The two suites test different work, and an agent that thinks hard on a reasoning puzzle may be terse on a finance workflow. But if you are budgeting a production workload, neither number substitutes for running your own tasks, and Argon is not broadly available to run them on yet. The 1 million output-token limit Google advertises is also conditional: Latent Space reports a standard maximum of 262,000 tokens, with the full million reached through a new “Long Decode Continuation” feature that pauses and resumes a response across separate calls. Google says it is rolling out to trusted cyber defenders through the Fairwind Program and, per The Next Web, to government users, with developers, enterprises and consumers to follow “as soon as possible.”
What survives the audit
The case against Argon is not the whole story, and the evidence for it deserves equal space. On the Vals Index Argon is first at 68.90%, which weights finance, coding, legal and tax work by each sector’s share of US GDP. Vals’ own methodology note calls Sonnet 5.5 and Opus 5.5 effectively tied at about 67%, with scores of 67.04% and 66.97%, so Argon’s lead is about two points on a measure with an interval near one point either side, per OrcaRouter’s figures of ±0.97 for Argon and ±0.92 for Sonnet. That is a lead, not a rout, and it comes on a benchmark that is mostly not coding. It is still a first place that Google did not have to select. Vals adds that Sonnet 5.5 reaches its score at about two-thirds of Opus 5.5’s cost per test, and warns that no single model leads every component of the index.
Artificial Analysis also has Argon first on its own AutomationBench-AA at 77.5%, six points ahead of Sonnet 5.5, and gives it 65% of grading criteria on AA-Briefcase, its knowledge-work test. Those are the kind of agentic business tasks Zapier and Vals measure, and they agree with Google’s picture of where Argon is strongest. Note that the Artificial Analysis figure for AutomationBench is a different variant from the 51.3% Google quotes, so the two cannot be compared directly.
The hallucination result is the most interesting number in any of these documents. On AA-Omniscience Argon hallucinates 15% of the time against 51% for Astra, according to The Decoder. The same source reports Argon’s accuracy at 50%, and Latent Space puts Astra at 63%. My read is that Argon is trading answered questions for fewer wrong ones, and that trade is a product decision rather than a free lunch. For a legal or finance workflow where a confident wrong answer is the expensive outcome, it may be the right trade. For open-ended coding it is a worse one.
Add a perfect 30 of 30 on Vals’ Vibe Code Bench, against 25 for Opus 5 and 24 for Astra per Latent Space, and the coding picture gets muddier rather than clearer. A model can build whole small apps cleanly and still disappoint on front-end polish and terminal agents. Those are different skills, and a single coding label hides the difference.

Spiky is a fine thing to be
Eighth of 223 is not a failure. It is also not a frontier lead. Argon is first on Text Arena, first on the Vals Index, tied with Astra on Artificial Analysis, behind two Anthropic models there, behind Sonnet 5.5 on Terminal-Bench 4 and eighth on WebDev. That is a model with a sharp profile, whose strongest results sit in finance, agentic business workflows, long video and security. Google would be better served selling that profile than implying a clean sweep.
The practical advice is dull and correct. Do not read the launch scorecard, and do not read the single-index headline either. Pull the board closest to your workload, run fifty of your own tasks when access opens, and price the run at $4 and $20 instead of $2 and $10. Argon may well win. It has to win at list price.

AI-generated editorial illustration · TemperatureZero · October 1, 2026
Keep reading the signal
Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.
Subscribe FreeContinue the archive