Logos are trademarks of their owners, shown only to identify the companies; image files via Wikimedia Commons.
This is a living page. Every few weeks a Super Intelligence company announces a model and a table showing it beating everyone else, and every company tops its own table. The question a reader actually has, which Super Intelligence model is best right now, can't be answered from those tables. So this page answers it from the scoreboards the companies don’t run: the human-vote leaderboard LMArena, Artificial Analysis’s test battery, the ARC Prize puzzles, METR’s task-length measurements and Epoch AI’s trend work. I keep it current, write a separate entry when something big lands, and link it from here. Last updated October 8, 2026.
The logos at the top are ten of the labs competing for the top of these boards. The rule here is the one you would use on any machine that claims to be smart: before believing it, find out who is doing the scoring.
Which Super Intelligence model is best right now?
Standing page · updated Oct 8, 2026
It depends whose test you trust. Human voters put Google’s new Gemini 4 Argon first. The biggest test battery puts Anthropic’s Claude Opus 5.5 five points clear.
Three independent scoreboards, three different leaders, and a fourth that has run out of ruler. The disagreement is the useful part: each one measures something different, and the “best model” changes with the job.
Human votes: Gemini 4 ArgonTest battery: Claude Opus 5.5New puzzles: GPT-6 Astra
Human votesLMArena
1525Gemini 4 Argon’s rating, #1, but preliminary on 4,892 votes
Test batteryArtificial Analysis
58Claude Opus 5.5 on the Intelligence Index, the top score of 226 entries
New puzzlesARC-AGI-3
62.7%GPT-6 Astra on the standard harness; 99.9% when allowed its own provider’s tools
Open-model lagEpoch AI
~4 mohow far the best open-weight models trail the frontier (2026 average)
Price spreadCost per task
8×what the top score costs ($5.98) vs a model six points behind it ($0.72)
Super Intelligence spendingBig-four capex
$172Bin Q2 2026 alone, more than their $155B for all of 2023
The scoreboard of scoreboards
LMArena · human votes
Gemini 4 Argon
Blind head-to-head votes from the public: which answer reads better. #2 Claude Opus 5.5. Rewards answers people like. As of Oct 8.
Artificial Analysis · test battery
Claude Opus 5.5
A battery of hard evaluations run by an independent firm, folded into one index. #2 Claude Sonnet 5.5. Weighted to reasoning, agents and code. As of Oct 8.
ARC-AGI-3 · new puzzles
GPT-6 Astra
Puzzle games the model has never seen, scored against 500 members of the public. The harness changes the score from 62.7% to 99.9%. As of Sep 3.
METR · task length
Ruler ran out
How long a task, measured in human working time, a model can finish half the time. The newest model measured was Claude Mythos Preview in May, and METR says anything above 16 hours is now beyond its tasks. Last update May 8.
AnthropicGoogleOpenAIOpen-weight models
Human votes: LMArena’s top eight Oct 8, 2026
Gemini 4 Argon (high)1525 ±9 · prelim.
Claude Opus 5.5 (high)1507 ±8
Claude Opus 4.6 (high)1504 ±3
Claude Fable 5 (high)1504 ±4
Claude Opus 4.7 (high)1501 ±4
Claude Fable 5.1 (max)1501 ±6
Claude Opus 4.61498 ±3
Gemini 3.8 Flash (high)1497 ±5 · prelim.
14801495151015251540
Each bar is a rating’s confidence range on an axis from 1480 to 1540. Ranges that overlap are a statistical tie. Argon’s range clears everyone else’s, so its lead is real on today’s votes, but it rests on 4,892 votes against tens of thousands for most of the models below it. Note how little separates places two through eight.
Test battery: Artificial Analysis Intelligence Index Oct 8, 2026
Claude Opus 5.558
Claude Sonnet 5.556
Claude Fable 5.153
GPT-6 Astra53
Gemini 4 Argon (tested at high)53
GPT-6.1 Sol52
MiMo-V2.6-Pro (Xiaomi) · best open46
GLM-5.3 (Z AI) · open45
Kimi K3 (Moonshot) · open44
015304560
Each model’s best reasoning setting, on a scale starting at zero so the gaps aren’t exaggerated. Artificial Analysis lists every setting separately; I keep one bar per model. Anthropic holds the top two places outright and shares third with Google and OpenAI, and the best open-weight model sits 12 points back.
What a point of intelligence costs Oct 8, 2026
Claude Fable 5.1 (max) · scores 53$7.63
Claude Opus 5.5 (max) · 58$5.98
Claude Sonnet 5.5 (max) · 56$5.46
GPT-6 Astra (max) · 53$3.26
Claude Sonnet 5.5 (xhigh) · 52$2.01
Gemini 4 Argon (high) · 53 · intro price$1.99
GPT-6.1 Sol (max) · 52$0.72
$0$2$4$6$8
Artificial Analysis’s cost to run one Intelligence Index task, in US dollars, with each model’s index score in the label. This measure folds in how many tokens a model burns thinking, which a per-token price list hides. The top score costs about eight times the cheapest score six points below it. Argon’s $1.99 is an introductory price that Google says will double.
Flagship list price, per million input tokens
GPT-4 · Mar 2023$30
Claude Opus 4.1 · 2025$15
Claude Fable 5.1 · 2026$10
Gemini 4 Argon · 2026, standard$4
The newest flagship lists at under a seventh of what GPT-4 cost three and a half years ago, and it is far more capable. The full story of the two speeds at which this price falls is in Six Million Tokens in Eleven Minutes.
Who pays for it: big-four capex
All of 2023$155B
Q2 2026, one quarter$172B
2026 guidance$735–750B
Capital spending by Amazon, Microsoft, Alphabet and Meta, per the Platformonomics tracker. 2026 is guidance as of late July, and the total has risen at each quarterly update this year. Q3 updates arrive with late-October earnings.
How fast is it improving? METR, Jan–May 2026
Mar 2025
METR: the length of tasks Super Intelligence agents can finish has doubled about every 7 months for six yearsMeasured in human working time, at 50% reliability, on software, ML and security tasks.
Jan 29, 2026
New task suite: since 2023 the doubling time is about 131 days, roughly four monthsTop model then: Claude Opus 4.5, about 5 hours 20 minutes (with a 95% range of about 3 to 12 hours).
May 8, 2026
METR adds Claude Mythos Preview and warns that anything above 16 hours is unreliable on its tasksThe ruler is running out. No newer model has been added since, including Opus 5.5, GPT-6 Astra or Argon.
Claim check
Holds up independentlyTrue, with a big asteriskVendor’s own numberDoesn’t hold up
Gemini 4 Argon beats Claude and GPT-6 on 13 of 19 benchmarksGoogle’s table, Sep 30. On the nine rows scored by outside leaderboards Argon still wins seven. The biggest margins (long context, video) are rows Google ran itself. Full check.
Vendor table
Gemini 4 Argon is the best model in the world#1 on LMArena human votes (preliminary). 53 on Artificial Analysis, five behind Claude Opus 5.5. True on one scoreboard, not the other.
Depends on the test
Argon hallucinates less than any other top modelArtificial Analysis: 15%, the lowest of models scoring 45+ on its index (as reported by The Batch).
Independent
GPT-6 Astra has essentially solved ARC-AGI-399.9% only with OpenAI’s own context tools; 62.7% on the neutral harness everyone else is scored on. ARC Prize calls both state of the art.
Depends on harness
Open models have nearly caught the frontierOne article in circulation put the gap at 6 points. Artificial Analysis’s own board shows 12 (46 vs 58), and Epoch AI found the lag slightly widened in 2026, to about four months.
Not on current data
Super Intelligence’s task length is doubling every four monthsMETR’s post-2023 estimate on its new suite: 131 days, with a range of 107 to 161. METR’s own caveat is that its suite is near its ceiling.
Independent
Who leads at what by the cited test, not my opinion
Coding agent in a terminal
Claude Opus 5.5
66.4% on Terminal-Bench 4.0, as printed in Google’s own table; Argon 57.4%.
Office and professional work
Gemini 4 Argon
Leads Vals Index, Finance Agent and Harvey’s legal benchmark, all scored by Vals AI.
Very long documents
Gemini 4 Argon
84.2% on GraphWalks at 256k–1M tokens vs 71.8% next best. Google’s own run.
Puzzles never seen before
GPT-6 Astra
62.7% on ARC-AGI-3’s standard harness, verified by ARC Prize.
Near-frontier on a budget
GPT-6.1 Sol
Index score 52 for $0.72 a task, per Artificial Analysis.
Weights you can download
MiMo-V2.6-Pro
Xiaomi’s model, 46 on the index, top of 81 open-weight entries.
What I am watching for
Argon’s LMArena rating as votes pile up
Preliminary ratings often settle. If the lead holds past 20,000 votes, the “best on human votes” chip stays.
Argon’s public release and model card
Also an Artificial Analysis run at its highest setting, which could move the 53.
Q3 earnings, late October
Whether the big four raise 2026 capex again, and what they say about 2027.
A new ruler from METR
The time-horizon chart is the clearest progress measure there is, and it has been quiet since May.
Every number on this page is from a named outside source with the date I read it. A lab’s own claim is labelled as the lab’s until someone else reproduces it. Scores move daily on some of these boards, so check the date in each panel heading.
How this page works
Three rules. Every ranking comes from someone else’s test, named and dated. A company’s own benchmark table is a claim, and it goes in the claim check until an independent scoreboard confirms or contradicts it. No single number decides “best”, because the scoreboards measure different things, and their disagreements are where the information is.
One disclosure belongs here. I use Anthropic’s Claude for nearly everything, and Claude models do very well on several of these boards. That is the reason the page runs on other people’s tests. If Claude slips, the bars will say so.
The four scoreboards in one line each.LMArena shows two anonymous answers side by side and asks a person which is better. It measures what people prefer, which rewards style as well as substance. Artificial Analysis runs a fixed battery of hard tests itself and averages them into one index, and it also prices each run. ARC-AGI-3 drops a model into puzzle games it has never seen and compares it with ordinary people. It is the closest thing to a test of learning on the spot. METR times how long the tasks a model can finish would take a skilled human. It is the most intuitive measure of progress there is, and the hardest to keep running.
Why the scoreboards disagree
Gemini 4 Argon is the clean example this week. People voting on LMArena prefer its answers to everyone else’s. On Artificial Analysis’s battery, which leans heavily on multi-step reasoning, agent work and code, it ties for the second tier. Google’s own table says the same thing in its own way: Argon wins the knowledge-work and long-document rows and loses the terminal and coding-agent rows. There is no contradiction there. A model can write the answer you’d pick and still be second-best at grinding through a software job.
The ARC-AGI-3 result for GPT-6 Astra is the other lesson: the harness is part of the score. The same model scores 62.7% when every lab plays by the same neutral rules and 99.9% when it can use OpenAI’s own context-management features. Both are real, and they answer different questions. When a headline quotes one number, it is worth asking which harness produced it.
Update log
October 8, 2026. Page created. Scoreboards read the same day: LMArena text (updated Oct 8), the Artificial Analysis Intelligence Index and open-weights leaderboard, ARC Prize’s Sep 3 post on GPT-6 Astra, METR’s time-horizon page (last updated May 8) and Jan 29 TH1.1 post, Epoch AI’s May 29 open-vs-closed analysis, and the Platformonomics Q2 capex tracker. First entry: Gemini 4 Argon, checked.
Where I could be wrong
These boards move daily. LMArena in particular re-rates as votes arrive, and a preliminary rating can shift a lot. Every panel is a snapshot from the date in its heading.
“One bar per model” is my simplification. Artificial Analysis scores each reasoning setting separately, and some models were tested at fewer settings than others. Argon was only tested at “high.”
Cost per task is one firm’s measurement on one set of tasks. Real-world bills depend on the job, caching and how much thinking you let a model do.
The flagship price line mixes labs and compares list prices for input tokens only. It shows the direction, not a like-for-like index.
The “who leads at what” cards inherit their sources’ limits. Three of the six come from Google’s own table, and the long-document card is a test Google ran itself; each card says so.
Sources
LMArena (arena.ai). Text leaderboard, overall view, updated October 8, 2026. arena.ai/leaderboard/text