The Super Intelligence Scorecard: Which Model Is Best, by Whose Test

SCIENCE & TECHNOLOGY · OCTOBER 8, 2026 · updated October 8, 2026

Ten Super Intelligence company logos in two rows of five on a dark background: Anthropic, OpenAI, Google Gemini, Meta, xAI, Alibaba Qwen, DeepSeek, Xiaomi, Mistral and Z.ai
Logos are trademarks of their owners, shown only to identify the companies; image files via Wikimedia Commons.

This is a living page. Every few weeks a Super Intelligence company announces a model and a table showing it beating everyone else, and every company tops its own table. The question a reader actually has, which Super Intelligence model is best right now, can't be answered from those tables. So this page answers it from the scoreboards the companies don’t run: the human-vote leaderboard LMArena, Artificial Analysis’s test battery, the ARC Prize puzzles, METR’s task-length measurements and Epoch AI’s trend work. I keep it current, write a separate entry when something big lands, and link it from here. Last updated October 8, 2026.

The logos at the top are ten of the labs competing for the top of these boards. The rule here is the one you would use on any machine that claims to be smart: before believing it, find out who is doing the scoring.

Which Super Intelligence model is best right now?

Standing page · updated Oct 8, 2026
It depends whose test you trust. Human voters put Google’s new Gemini 4 Argon first. The biggest test battery puts Anthropic’s Claude Opus 5.5 five points clear.
Three independent scoreboards, three different leaders, and a fourth that has run out of ruler. The disagreement is the useful part: each one measures something different, and the “best model” changes with the job.
Human votes: Gemini 4 Argon Test battery: Claude Opus 5.5 New puzzles: GPT-6 Astra
Human votesLMArena
1525Gemini 4 Argon’s rating, #1, but preliminary on 4,892 votes
Test batteryArtificial Analysis
58Claude Opus 5.5 on the Intelligence Index, the top score of 226 entries
New puzzlesARC-AGI-3
62.7%GPT-6 Astra on the standard harness; 99.9% when allowed its own provider’s tools
Open-model lagEpoch AI
~4 mohow far the best open-weight models trail the frontier (2026 average)
Price spreadCost per task
8×what the top score costs ($5.98) vs a model six points behind it ($0.72)
Super Intelligence spendingBig-four capex
$172Bin Q2 2026 alone, more than their $155B for all of 2023
The scoreboard of scoreboards
LMArena · human votes
Gemini 4 Argon
Blind head-to-head votes from the public: which answer reads better. #2 Claude Opus 5.5. Rewards answers people like. As of Oct 8.
Artificial Analysis · test battery
Claude Opus 5.5
A battery of hard evaluations run by an independent firm, folded into one index. #2 Claude Sonnet 5.5. Weighted to reasoning, agents and code. As of Oct 8.
ARC-AGI-3 · new puzzles
GPT-6 Astra
Puzzle games the model has never seen, scored against 500 members of the public. The harness changes the score from 62.7% to 99.9%. As of Sep 3.
METR · task length
Ruler ran out
How long a task, measured in human working time, a model can finish half the time. The newest model measured was Claude Mythos Preview in May, and METR says anything above 16 hours is now beyond its tasks. Last update May 8.
AnthropicGoogleOpenAIOpen-weight models
Human votes: LMArena’s top eight Oct 8, 2026
Gemini 4 Argon (high)1525 ±9 · prelim.
Claude Opus 5.5 (high)1507 ±8
Claude Opus 4.6 (high)1504 ±3
Claude Fable 5 (high)1504 ±4
Claude Opus 4.7 (high)1501 ±4
Claude Fable 5.1 (max)1501 ±6
Claude Opus 4.61498 ±3
Gemini 3.8 Flash (high)1497 ±5 · prelim.
Each bar is a rating’s confidence range on an axis from 1480 to 1540. Ranges that overlap are a statistical tie. Argon’s range clears everyone else’s, so its lead is real on today’s votes, but it rests on 4,892 votes against tens of thousands for most of the models below it. Note how little separates places two through eight.
Test battery: Artificial Analysis Intelligence Index Oct 8, 2026
Claude Opus 5.558
Claude Sonnet 5.556
Claude Fable 5.153
GPT-6 Astra53
Gemini 4 Argon (tested at high)53
GPT-6.1 Sol52
MiMo-V2.6-Pro (Xiaomi) · best open46
GLM-5.3 (Z AI) · open45
Kimi K3 (Moonshot) · open44
Each model’s best reasoning setting, on a scale starting at zero so the gaps aren’t exaggerated. Artificial Analysis lists every setting separately; I keep one bar per model. Anthropic holds the top two places outright and shares third with Google and OpenAI, and the best open-weight model sits 12 points back.
What a point of intelligence costs Oct 8, 2026
Claude Fable 5.1 (max) · scores 53$7.63
Claude Opus 5.5 (max) · 58$5.98
Claude Sonnet 5.5 (max) · 56$5.46
GPT-6 Astra (max) · 53$3.26
Claude Sonnet 5.5 (xhigh) · 52$2.01
Gemini 4 Argon (high) · 53 · intro price$1.99
GPT-6.1 Sol (max) · 52$0.72
Artificial Analysis’s cost to run one Intelligence Index task, in US dollars, with each model’s index score in the label. This measure folds in how many tokens a model burns thinking, which a per-token price list hides. The top score costs about eight times the cheapest score six points below it. Argon’s $1.99 is an introductory price that Google says will double.
Flagship list price, per million input tokens
GPT-4 · Mar 2023$30
Claude Opus 4.1 · 2025$15
Claude Fable 5.1 · 2026$10
Gemini 4 Argon · 2026, standard$4
The newest flagship lists at under a seventh of what GPT-4 cost three and a half years ago, and it is far more capable. The full story of the two speeds at which this price falls is in Six Million Tokens in Eleven Minutes.
Who pays for it: big-four capex
All of 2023$155B
Q2 2026, one quarter$172B
2026 guidance$735–750B
Capital spending by Amazon, Microsoft, Alphabet and Meta, per the Platformonomics tracker. 2026 is guidance as of late July, and the total has risen at each quarterly update this year. Q3 updates arrive with late-October earnings.
How fast is it improving? METR, Jan–May 2026
Mar 2025
METR: the length of tasks Super Intelligence agents can finish has doubled about every 7 months for six yearsMeasured in human working time, at 50% reliability, on software, ML and security tasks.
Jan 29, 2026
New task suite: since 2023 the doubling time is about 131 days, roughly four monthsTop model then: Claude Opus 4.5, about 5 hours 20 minutes (with a 95% range of about 3 to 12 hours).
May 8, 2026
METR adds Claude Mythos Preview and warns that anything above 16 hours is unreliable on its tasksThe ruler is running out. No newer model has been added since, including Opus 5.5, GPT-6 Astra or Argon.
Claim check
Holds up independentlyTrue, with a big asteriskVendor’s own numberDoesn’t hold up
Gemini 4 Argon beats Claude and GPT-6 on 13 of 19 benchmarksGoogle’s table, Sep 30. On the nine rows scored by outside leaderboards Argon still wins seven. The biggest margins (long context, video) are rows Google ran itself. Full check.
Vendor table
Gemini 4 Argon is the best model in the world#1 on LMArena human votes (preliminary). 53 on Artificial Analysis, five behind Claude Opus 5.5. True on one scoreboard, not the other.
Depends on the test
Argon hallucinates less than any other top modelArtificial Analysis: 15%, the lowest of models scoring 45+ on its index (as reported by The Batch).
Independent
GPT-6 Astra has essentially solved ARC-AGI-399.9% only with OpenAI’s own context tools; 62.7% on the neutral harness everyone else is scored on. ARC Prize calls both state of the art.
Depends on harness
Open models have nearly caught the frontierOne article in circulation put the gap at 6 points. Artificial Analysis’s own board shows 12 (46 vs 58), and Epoch AI found the lag slightly widened in 2026, to about four months.
Not on current data
Super Intelligence’s task length is doubling every four monthsMETR’s post-2023 estimate on its new suite: 131 days, with a range of 107 to 161. METR’s own caveat is that its suite is near its ceiling.
Independent
Who leads at what by the cited test, not my opinion
Coding agent in a terminal
Claude Opus 5.5
66.4% on Terminal-Bench 4.0, as printed in Google’s own table; Argon 57.4%.
Office and professional work
Gemini 4 Argon
Leads Vals Index, Finance Agent and Harvey’s legal benchmark, all scored by Vals AI.
Very long documents
Gemini 4 Argon
84.2% on GraphWalks at 256k–1M tokens vs 71.8% next best. Google’s own run.
Puzzles never seen before
GPT-6 Astra
62.7% on ARC-AGI-3’s standard harness, verified by ARC Prize.
Near-frontier on a budget
GPT-6.1 Sol
Index score 52 for $0.72 a task, per Artificial Analysis.
Weights you can download
MiMo-V2.6-Pro
Xiaomi’s model, 46 on the index, top of 81 open-weight entries.
What I am watching for
Argon’s LMArena rating as votes pile up
Preliminary ratings often settle. If the lead holds past 20,000 votes, the “best on human votes” chip stays.
Argon’s public release and model card
Also an Artificial Analysis run at its highest setting, which could move the 53.
Q3 earnings, late October
Whether the big four raise 2026 capex again, and what they say about 2027.
A new ruler from METR
The time-horizon chart is the clearest progress measure there is, and it has been quiet since May.
The story so far
Oct 8, 2026 · newGemini 4 Argon: Google’s Benchmarks, Checked13 wins of 19, which rows Google ran itself, and why the leaderboards disagree. Sep 22, 2026Claude Opus 5.5 Is Cheaper and FasterThe current Artificial Analysis leader, and a launch that opened with an alignment score. Sep 21, 2026The Most-Downloaded AI Models in the World Are Alibaba’sThe open-weight side of the race, and Qwen’s three billion downloads. Sep 12, 2026The Week the AI Labs Asked to Be RegulatedGPT-6 Astra’s launch, and OpenAI’s turn on safety rules. Sep 10, 2026 · the LedgerSix Million Tokens in Eleven MinutesWhat intelligence costs, and who pays for the free version.
Every number on this page is from a named outside source with the date I read it. A lab’s own claim is labelled as the lab’s until someone else reproduces it. Scores move daily on some of these boards, so check the date in each panel heading.

How this page works

Three rules. Every ranking comes from someone else’s test, named and dated. A company’s own benchmark table is a claim, and it goes in the claim check until an independent scoreboard confirms or contradicts it. No single number decides “best”, because the scoreboards measure different things, and their disagreements are where the information is.

One disclosure belongs here. I use Anthropic’s Claude for nearly everything, and Claude models do very well on several of these boards. That is the reason the page runs on other people’s tests. If Claude slips, the bars will say so.

The four scoreboards in one line each. LMArena shows two anonymous answers side by side and asks a person which is better. It measures what people prefer, which rewards style as well as substance. Artificial Analysis runs a fixed battery of hard tests itself and averages them into one index, and it also prices each run. ARC-AGI-3 drops a model into puzzle games it has never seen and compares it with ordinary people. It is the closest thing to a test of learning on the spot. METR times how long the tasks a model can finish would take a skilled human. It is the most intuitive measure of progress there is, and the hardest to keep running.

Why the scoreboards disagree

Gemini 4 Argon is the clean example this week. People voting on LMArena prefer its answers to everyone else’s. On Artificial Analysis’s battery, which leans heavily on multi-step reasoning, agent work and code, it ties for the second tier. Google’s own table says the same thing in its own way: Argon wins the knowledge-work and long-document rows and loses the terminal and coding-agent rows. There is no contradiction there. A model can write the answer you’d pick and still be second-best at grinding through a software job.

The ARC-AGI-3 result for GPT-6 Astra is the other lesson: the harness is part of the score. The same model scores 62.7% when every lab plays by the same neutral rules and 99.9% when it can use OpenAI’s own context-management features. Both are real, and they answer different questions. When a headline quotes one number, it is worth asking which harness produced it.

Update log

Where I could be wrong

Sources

  1. LMArena (arena.ai). Text leaderboard, overall view, updated October 8, 2026. arena.ai/leaderboard/text
  2. Artificial Analysis. LLM leaderboard (Intelligence Index, cost per task, speed), read October 8, 2026. artificialanalysis.ai/leaderboards/models
  3. Artificial Analysis. Open-weights models leaderboard, read October 8, 2026. artificialanalysis.ai/models/open-source
  4. ARC Prize. OpenAI’s GPT-6 Astra on ARC-AGI-3, September 3, 2026. arcprize.org/blog/astra
  5. METR. Time Horizon 1.1, January 29, 2026. metr.org
  6. METR. Time-horizons page, last updated May 8, 2026. metr.org/time-horizons
  7. METR. Measuring AI Ability to Complete Long Tasks, March 19, 2025. metr.org
  8. Epoch AI. Data Insight: open-weight vs closed models on the Epoch Capabilities Index, May 29, 2026. epoch.ai
  9. Google DeepMind. Gemini models page (Performance table) and Gemini 4 Argon Model evaluation methodology PDF. deepmind.google/models/gemini
  10. DeepLearning.AI, The Batch. Data Points: Gemini 4 Argon’s benchmarks and availability, October 5, 2026. deeplearning.ai
  11. Platformonomics. Follow the CAPEX: Q2 2026 Scoreboard, July 2026. platformonomics.com
  12. The Ledger. Six Million Tokens in Eleven Minutes (September 10, 2026), for the flagship price history and its sources.

Keep reading