Gemini 4 Argon: Google's Benchmarks, Checked

SCIENCE & TECHNOLOGY · OCTOBER 8, 2026

An open nineteenth-century laboratory notebook belonging to the chemist William Ramsay, its pages filled with handwritten notes, resting on a dark cloth
One of William Ramsay's laboratory notebooks. Ramsay and Lord Rayleigh discovered argon in 1894. Photo by UCL Mathematical & Physical Sciences, CC BY 2.0, via Wikimedia Commons.

On September 30, 2026 Google DeepMind announced Gemini 4 Argon and published a 19-row benchmark table in which it beats Anthropic’s Claude Opus 5.5 and Claude Fable 5.1 and OpenAI’s GPT-6 Astra on 13 rows, ties one and loses five. Almost nobody can use it yet: it is going first to vetted security teams through Google’s Fairwind Program, at an introductory $2 per million input tokens and $10 per million output tokens. The question worth asking about a table like that is not whether the numbers are impressive. They are. It is who ran each test. I went through the methodology document row by row, then checked the leaderboards Google doesn’t control. This is the first entry feeding the Super Intelligence Scorecard, the standing page where I keep track of this.

The name is a nice choice, for what it’s worth. Argon was discovered in 1894 because two measurements of the same thing didn’t agree. Lord Rayleigh found that nitrogen taken from the air was about 0.5% denser than nitrogen made in the lab. He refused to call that experimental error, and the gap turned out to be a gas nobody knew was there. Disagreeing measurements are the subject of this entry too.

What Google's Gemini 4 Argon benchmarks actually show

Checked Oct 8, 2026
Argon tops 13 of Google’s 19 benchmarks. Two independent scoreboards disagree about what that adds up to.
LMArena’s human voters put it first, on a preliminary count of 4,892 votes. Artificial Analysis’s test battery scores it 53, tied with GPT-6 Astra and five points behind Claude Opus 5.5.
13 wins · 1 tie · 5 losses Testers only, no public API No model card linked yet
Human votesLMArena
#1on the text leaderboard, 1525 ±9, flagged preliminary (Oct 8)
Test batteryArtificial Analysis
53Intelligence Index, tested at “high”; the leader, Claude Opus 5.5, scores 58
Cost per taskIntroductory price
$1.99per index task, vs $5.98 for Opus 5.5 and $3.26 for GPT-6 Astra at max
Google’s table, with who ran each test
Third-party leaderboard for every modelGoogle ran every modelMixed: Google ran Argon, rivals’ numbers from elsewhere
Benchmark
Argon
Astra
Fable
Opus
Vals IndexKnowledge work 3rd
68.9
63.1
65.8
67.0
AutomationBenchKnowledge work 3rd
51.3
41.4
31.4
42.5
Vals Finance Agent v2Knowledge work 3rd
65.4
53.5
58.9
58.6
Harvey Legal AgentKnowledge work 3rd
19.6
5.4
6.7
3.8
DeepSWE v1.1Agentic coding mixed
77.9
74.1
67.4
74.2
FrontierSWE v2Agentic coding 3rd
55.0
65.5
56.3
62.3
Vibe Code BenchAgentic coding 3rd
91.9
89.6
90.3
90.3
Terminal-Bench 4.0Agentic coding mixed
57.4
58.2
57.9
66.4
PostTrainBenchML engineering Google
45.3
44.3
40.2
49.3
Terminal-Bench ScienceScience & math mixed
57.6
68.1
52.6
63.3
LABBench 2Science & math Google
88.8
85.4
68.6
73.1
RiemannBenchScience & math 3rd
76.0
72.0
65.6
69.6
GraphWalks ≤128kLong context Google
99.7
98.7
91.4
90.6
GraphWalks 256k–1MLong context Google
84.2
71.8
65.0
66.8
Agent’s Last ExamComputer use mixed
39.5
34.2
–
38.2
OSWorld 2.0 (offline)Computer use mixed
69.2
72.6
–
–
ChartographyMultimodal 3rd
71.6
71.0
46.2
66.3
LVBench (video)Multimodal Google
91.7
87.5
79.7
83.7
CWE-bench v1Cybersecurity 3rd
68.0
68.0
58.0
67.0
Scores in percent, from the Performance table on Google DeepMind’s Gemini page. The source tag for each row is from Google’s own methodology PDF. A dash means Google printed no result. Columns: Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5.
Does it matter who ran the test?
Third-party leaderboards (9 rows)7 wins · 1 tie · 1 loss
Google ran every model (5 rows)4 wins · 1 loss
Mixed sources (5 rows)2 wins · 3 losses
Green is a win for Argon, amber a tie, orange a loss. The surprise runs in Google’s favor: Argon does best on the rows where an outside leaderboard scored every model, and worst on the mixed rows, where Google scored Argon itself and took the rivals’ numbers from public leaderboards or system cards. Every mixed loss is an agent working in a terminal or driving a computer.

The wins are not mostly home cooking

My first suspicion with any vendor table is the grading. A company that picks the benchmarks, runs its own model, and pulls competitors’ numbers from wherever is looking at itself in a flattering mirror. Google’s methodology document is unusually plain about which rows are which, and the split doesn’t support that suspicion. On the nine rows where an outside leaderboard scored all four models (Vals AI, Zapier, Proximal, Surge, CWE-bench), Argon wins seven, ties one and loses one. The knowledge-work results are not close. On Harvey’s legal-agent benchmark, scored by Vals AI, Argon gets 19.6% against 3.8% to 6.7% for the others.

The rows that deserve an asterisk are the ones Google ran itself. Long context (GraphWalks) and video understanding (LVBench) are where Argon’s margins are largest, and Google scored every model there. On LVBench the document notes that Gemini watched at one frame per second while GPT-6 Astra was capped at 800 frames, Opus 5.5 at 600 and Fable 5.1 at 300, “due to API limitations.” That may be the only fair way to run it, but on a long video it means the rivals saw less of the film. Treat those two rows as Google’s word until someone else repeats them.

The losses tell a consistent story. Every row Argon drops is an agent doing long, hands-on technical work: Terminal-Bench 4.0 (57.4% against Opus 5.5’s 66.4%), FrontierSWE v2, Terminal-Bench Science, the ML-engineering test PostTrainBench, and the computer-use test OSWorld. At sitting in a terminal and grinding through a software job, Google’s own table says Argon is not the leader yet.

What the scoreboards Google doesn’t control say

LMArena (now at arena.ai) ranks models by blind head-to-head votes from the public. As of October 8 its text leaderboard has gemini-4-argon-high first at 1525, 18 points ahead of Claude Opus 5.5 at 1507. The ratings carry error bars of ±9 and ±8 points and the two ranges don’t overlap, so the lead is real on today’s votes. It rests on 4,892 votes, though, against tens of thousands for the older models behind it, and the site marks it preliminary.

Artificial Analysis runs its own battery of tests and folds them into one Intelligence Index. It had pre-release access, and it scores Argon at its “high” setting at 53. That ties GPT-6 Astra at max and Claude Fable 5.1, and it sits five points behind Claude Opus 5.5 at 58 and three behind Claude Sonnet 5.5 at 56. Two Artificial Analysis findings, as reported by DeepLearning.AI’s The Batch, are good news for Google: Argon has the lowest hallucination rate (15%) of any model scoring 45 or more on the index, and it ranks first on Artificial Analysis’s own version of AutomationBench at 78%.

So which is it, best in the world or tied for fifth? Both, depending on the question. Human voters judge an answer by how it reads, and Argon apparently reads very well. The index is weighted toward multi-step reasoning, agentic tasks and coding, which is where Google’s own table shows Argon’s weak rows. One more caveat runs in Argon’s favor. Google’s table uses its highest thinking setting, and Artificial Analysis tested “high,” so the index number may move once a max setting is measured.

A disclosure, since this is about Super Intelligence models. I use Anthropic’s Claude for nearly everything, and two of the models Argon is measured against are Claude models. That is why this entry, and the scorecard it feeds, leans on other people’s tests and not on my impressions.

What it costs, and why investors are watching

The introductory price is $2 per million input tokens and $10 per million output tokens, with cached input 95% off. Google’s post says it will rise to $4 and $20 later, with no date given. On Artificial Analysis’s cost-per-task measure, Argon at the introductory price comes to $1.99 per index task, against $3.26 for GPT-6 Astra at max and $5.98 for Claude Opus 5.5 at max. At the standard price that roughly doubles to about $3.98. For context, OpenAI’s GPT-6.1 Sol scores 52 for $0.72 a task, so “cheap” has competition too.

The reason the investor press cares is the bill behind it. Alphabet raised its 2026 capital-spending guidance to $195–205 billion in July, part of roughly $735–750 billion planned across Amazon, Microsoft, Alphabet and Meta this year. A model that wins knowledge-work benchmarks and launches at a discount is what that spending is supposed to buy, and the price is a signal that Google intends to compete on cost. None of that says anything about what any stock will do. It is the context the coverage is written in.

What isn’t there yet

Where I could be wrong

Sources

  1. Google. Gemini 4 Argon (announcement), September 30, 2026. blog.google
  2. Google DeepMind. Gemini models page, Performance table. deepmind.google/models/gemini
  3. Google DeepMind. Gemini 4 Argon Model evaluation: approach, methodology & results (PDF). deepmind.google
  4. LMArena (arena.ai). Text leaderboard, overall, updated October 8, 2026. arena.ai/leaderboard/text
  5. Artificial Analysis. Gemini 4 Argon (High) model page and Intelligence Index leaderboard. artificialanalysis.ai
  6. DeepLearning.AI, The Batch. Data Points: Gemini 4 Argon’s benchmarks and availability, October 5, 2026. deeplearning.ai
  7. Technology.org. Google Gemini 4 Argon Debuts for Cyber Defenders, October 1, 2026. technology.org
  8. Platformonomics. Follow the CAPEX: Q2 2026 Scoreboard, July 2026. platformonomics.com
  9. American Physical Society. This Month in Physics History: Lord Rayleigh and the Discovery of Argon, August 13, 1894. APS News, August 2013. aps.org

Keep reading