Gemini 4 Argon: Google's Benchmarks, Checked
SCIENCE & TECHNOLOGY · OCTOBER 8, 2026

On September 30, 2026 Google DeepMind announced Gemini 4 Argon and published a 19-row benchmark table in which it beats Anthropic’s Claude Opus 5.5 and Claude Fable 5.1 and OpenAI’s GPT-6 Astra on 13 rows, ties one and loses five. Almost nobody can use it yet: it is going first to vetted security teams through Google’s Fairwind Program, at an introductory $2 per million input tokens and $10 per million output tokens. The question worth asking about a table like that is not whether the numbers are impressive. They are. It is who ran each test. I went through the methodology document row by row, then checked the leaderboards Google doesn’t control. This is the first entry feeding the Super Intelligence Scorecard, the standing page where I keep track of this.
The name is a nice choice, for what it’s worth. Argon was discovered in 1894 because two measurements of the same thing didn’t agree. Lord Rayleigh found that nitrogen taken from the air was about 0.5% denser than nitrogen made in the lab. He refused to call that experimental error, and the gap turned out to be a gas nobody knew was there. Disagreeing measurements are the subject of this entry too.
What Google's Gemini 4 Argon benchmarks actually show
The wins are not mostly home cooking
My first suspicion with any vendor table is the grading. A company that picks the benchmarks, runs its own model, and pulls competitors’ numbers from wherever is looking at itself in a flattering mirror. Google’s methodology document is unusually plain about which rows are which, and the split doesn’t support that suspicion. On the nine rows where an outside leaderboard scored all four models (Vals AI, Zapier, Proximal, Surge, CWE-bench), Argon wins seven, ties one and loses one. The knowledge-work results are not close. On Harvey’s legal-agent benchmark, scored by Vals AI, Argon gets 19.6% against 3.8% to 6.7% for the others.
The rows that deserve an asterisk are the ones Google ran itself. Long context (GraphWalks) and video understanding (LVBench) are where Argon’s margins are largest, and Google scored every model there. On LVBench the document notes that Gemini watched at one frame per second while GPT-6 Astra was capped at 800 frames, Opus 5.5 at 600 and Fable 5.1 at 300, “due to API limitations.” That may be the only fair way to run it, but on a long video it means the rivals saw less of the film. Treat those two rows as Google’s word until someone else repeats them.
The losses tell a consistent story. Every row Argon drops is an agent doing long, hands-on technical work: Terminal-Bench 4.0 (57.4% against Opus 5.5’s 66.4%), FrontierSWE v2, Terminal-Bench Science, the ML-engineering test PostTrainBench, and the computer-use test OSWorld. At sitting in a terminal and grinding through a software job, Google’s own table says Argon is not the leader yet.
What the scoreboards Google doesn’t control say
LMArena (now at arena.ai) ranks models by blind head-to-head votes from the public. As of October 8 its text leaderboard has gemini-4-argon-high first at 1525, 18 points ahead of Claude Opus 5.5 at 1507. The ratings carry error bars of ±9 and ±8 points and the two ranges don’t overlap, so the lead is real on today’s votes. It rests on 4,892 votes, though, against tens of thousands for the older models behind it, and the site marks it preliminary.
Artificial Analysis runs its own battery of tests and folds them into one Intelligence Index. It had pre-release access, and it scores Argon at its “high” setting at 53. That ties GPT-6 Astra at max and Claude Fable 5.1, and it sits five points behind Claude Opus 5.5 at 58 and three behind Claude Sonnet 5.5 at 56. Two Artificial Analysis findings, as reported by DeepLearning.AI’s The Batch, are good news for Google: Argon has the lowest hallucination rate (15%) of any model scoring 45 or more on the index, and it ranks first on Artificial Analysis’s own version of AutomationBench at 78%.
So which is it, best in the world or tied for fifth? Both, depending on the question. Human voters judge an answer by how it reads, and Argon apparently reads very well. The index is weighted toward multi-step reasoning, agentic tasks and coding, which is where Google’s own table shows Argon’s weak rows. One more caveat runs in Argon’s favor. Google’s table uses its highest thinking setting, and Artificial Analysis tested “high,” so the index number may move once a max setting is measured.
A disclosure, since this is about Super Intelligence models. I use Anthropic’s Claude for nearly everything, and two of the models Argon is measured against are Claude models. That is why this entry, and the scorecard it feeds, leans on other people’s tests and not on my impressions.
What it costs, and why investors are watching
The introductory price is $2 per million input tokens and $10 per million output tokens, with cached input 95% off. Google’s post says it will rise to $4 and $20 later, with no date given. On Artificial Analysis’s cost-per-task measure, Argon at the introductory price comes to $1.99 per index task, against $3.26 for GPT-6 Astra at max and $5.98 for Claude Opus 5.5 at max. At the standard price that roughly doubles to about $3.98. For context, OpenAI’s GPT-6.1 Sol scores 52 for $0.72 a task, so “cheap” has competition too.
The reason the investor press cares is the bill behind it. Alphabet raised its 2026 capital-spending guidance to $195–205 billion in July, part of roughly $735–750 billion planned across Amazon, Microsoft, Alphabet and Meta this year. A model that wins knowledge-work benchmarks and launches at a discount is what that spending is supposed to buy, and the price is a signal that Google intends to compete on cost. None of that says anything about what any stock will do. It is the context the coverage is written in.
What isn’t there yet
- Public access. Only Fairwind Program testers can use it. Google says paid API customers and Google AI Ultra subscribers come first, “as soon as possible,” with no date.
- A model card. Google’s Gemini models page links the benchmark table and a methodology PDF, but when I checked on October 8 it linked no model card or system card for Argon.
- Independent replication of the Google-run rows. GraphWalks, LVBench, LABBench 2 and PostTrainBench are Google’s runs of every model.
- Cyber guardrails for the first users. Google says the defenders in the Fairwind Program get the model without its cyber guardrails, which is the point of the program. Everyone else gets the restricted version.
Where I could be wrong
- LMArena moves. A preliminary rating on under 5,000 votes can shift a lot as votes come in. The 18-point lead is a snapshot from October 8.
- The source tags are Google’s description of its own method. I sorted rows by what the methodology PDF says, not by re-running anything.
- “Tied for fifth” is my count. Artificial Analysis ranks each reasoning setting separately, so Argon (high) is eighth of 226 entries. Counted one model at a time, it shares a score with GPT-6 Astra and Claude Fable 5.1 behind Opus 5.5 and Sonnet 5.5.
- The hallucination and AutomationBench-AA figures are Artificial Analysis results as reported by The Batch. The model page I could read listed those evaluations without the numbers.
Sources
- Google. Gemini 4 Argon (announcement), September 30, 2026. blog.google
- Google DeepMind. Gemini models page, Performance table. deepmind.google/models/gemini
- Google DeepMind. Gemini 4 Argon Model evaluation: approach, methodology & results (PDF). deepmind.google
- LMArena (arena.ai). Text leaderboard, overall, updated October 8, 2026. arena.ai/leaderboard/text
- Artificial Analysis. Gemini 4 Argon (High) model page and Intelligence Index leaderboard. artificialanalysis.ai
- DeepLearning.AI, The Batch. Data Points: Gemini 4 Argon’s benchmarks and availability, October 5, 2026. deeplearning.ai
- Technology.org. Google Gemini 4 Argon Debuts for Cyber Defenders, October 1, 2026. technology.org
- Platformonomics. Follow the CAPEX: Q2 2026 Scoreboard, July 2026. platformonomics.com
- American Physical Society. This Month in Physics History: Lord Rayleigh and the Discovery of Argon, August 13, 1894. APS News, August 2013. aps.org



