Independent guide. Not affiliated with, endorsed by or sponsored by Google, Google DeepMind or Gemini. About this site
gemini4argon.comUnofficial guide

Benchmarks

Gemini 4 Argon benchmarks: every score, and where it loses

Updated · Independent coverage, not affiliated with Google

Google published 19 benchmark results for Gemini 4 Argon, comparing it with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. Argon comes first on 13 of them, ties on one, and trails on five. Here is the whole table in plain English, including the parts the headlines skip.

These are Google's numbers. Every score on this page comes from Google's own announcement and its evaluation methodology. No independent leaderboard (such as LMArena or Artificial Analysis) had scored Argon when we published this, and we'll add those results when they appear.

The short version

First place13 of 19Including DeepSWE, LVBench and every knowledge-work test
Tied1CWE-bench v1, with GPT-6 Astra at 68.0%
Behind5Terminal work, FrontierSWE, ML engineering, science, OSWorld

In one sentence: Argon is strongest at long knowledge work, reading huge inputs and understanding video, is at or near the top at software engineering, and is not the best model for terminal-heavy agent work, where Claude Opus 5.5 leads.

Gemini 4 Argon coding benchmarks

DeepSWE v1.1

Real-world software engineering tasks. Higher is better.

Gemini 4 Argon77.9%
Claude Opus 5.574.2%
GPT-6 Astra74.1%
Claude Fable 5.167.4%
Agentic codingArgonGPT-6 AstraFable 5.1Opus 5.5
DeepSWE v1.177.9%74.1%67.4%74.2%
FrontierSWE v255.0%65.5%56.3%62.3%
Vibe Code Bench91.9%89.6%90.3%90.3%
Terminal-bench 4.057.4%58.2%57.9%66.4%
PostTrainBench (ML engineering)45.3%44.3%40.2%49.3%

Argon wins the headline coding test, DeepSWE, by 3.7 points, and edges out everyone on Vibe Code Bench. But it is last of the four on FrontierSWE v2, and nine points behind Claude Opus 5.5 on Terminal-bench 4.0, which measures getting real work done in a command line. If your work is mostly agents running in a terminal, Opus 5.5 still has the edge on Google's own chart.

Knowledge work: Argon's biggest lead

Knowledge workArgonGPT-6 AstraFable 5.1Opus 5.5
Vals Index68.9%63.1%65.8%67.0%
AutomationBench (score)51.3%41.4%31.4%42.5%
Vals Finance Agent v265.4%53.5%58.9%58.6%
Harvey's Legal Agent Benchmark19.6%5.4%6.7%3.8%

This is where Argon pulls away. It beats the next-best model on AutomationBench by almost 9 points and on Vals Finance Agent by 6.5. On Harvey's legal agent test the absolute scores are low for everyone, but Argon's 19.6% is roughly three times the next model. For finance, legal and business-automation work, these are the numbers to watch once independent testing starts.

Long context and video understanding

TestArgonGPT-6 AstraFable 5.1Opus 5.5
GraphWalks, up to 128K (BFS, F1)99.7%98.7%91.4%90.6%
GraphWalks, 256K to 1M (BFS, F1)84.2%71.8%65.0%66.8%
Chartography71.6%71.0%46.2%66.3%
LVBench (long video)91.7%87.5%79.7%83.7%

The GraphWalks results matter because they test whether a model can still find its way around a truly enormous input. Between 256K and 1M tokens, Argon keeps 84.2% while the others fall to between 65% and 72%. Add the LVBench lead on long videos and chart reading, and Argon looks built for feeding in a lot at once. Curious what that means for video? Read making videos with Gemini 4 Argon.

Science and math

Science and mathArgonGPT-6 AstraFable 5.1Opus 5.5
Terminal-Bench Science 0.157.6%68.1%52.6%63.3%
LABBench 288.8%85.4%68.6%73.1%
RiemannBench76.0%72.0%65.6%69.6%

A split result: Argon leads on lab-science questions (LABBench 2) and advanced math (RiemannBench), but GPT-6 Astra wins clearly when science work has to happen in a terminal.

Computer use

Computer useArgonGPT-6 AstraFable 5.1Opus 5.5
Agent's Last Exam (pass rate)39.5%34.2%not reported38.2%
OSWorld-2.0 (offline subset, partial score)69.2%72.6%not reportednot reported

Close race. Argon narrowly leads Agent's Last Exam and trails GPT-6 Astra on OSWorld-2.0. Google hasn't published how Argon's computer-use features work in the API yet, so treat these as a preview.

Cybersecurity: why Argon ships to defenders first

Security testGemini 4 ArgonCompared with
CWE-bench v168.0%GPT-6 Astra 68.0%, Opus 5.5 67.0%, Fable 5.1 58.0%
Real-world vulnerability discovery (20 languages)85.8%Gemini 3.8 Flash Cyber 71.0%
Wiz penetration test benchmark70.9%Gemini 3.8 Flash Cyber 58.2%
Gray Swan prompt injection, attack success (lower is better)0.7%Opus 5.5 1.0%, Fable 5.1 1.0%, GPT-6 Astra 8.5%

Argon is very good at finding security holes, and Google compares it here with its own security-tuned Gemini 3.8 Flash Cyber model. That is exactly why the first people to get it are vetted defenders in Google's Fairwind Program. It is also the hardest model on Google's chart to trick with hidden instructions: just 0.7% of prompt-injection attacks succeeded after 15 tries.

Want the models you can actually use today?

Veo 3.1, Gemini Omni, Nano Banana 2, Seedance and Kling are all live on Lumeta, in one account.

See the lineup

How to read these numbers

  • Vendors pick their tests. Every lab leads with the benchmarks it wins. To Google's credit, its table also shows five losses.
  • Small gaps are noise. A difference of one or two points (like Vibe Code Bench, 91.9% vs 90.3%) may not show up in your daily work.
  • Your task is the real test. Run the same five real prompts from your own work through each model when Argon opens up. That tells you more than any table.
  • Not everyone is convinced yet. Press coverage has reported some skepticism inside Google about Argon's real-world coding. Independent results will settle it.

For the full list of test descriptions and settings, see Google DeepMind's evaluation methodology page linked below.

Benchmark questions

Is Gemini 4 Argon better than Claude Opus 5.5?

On Google's chart, Argon beats Claude Opus 5.5 on 14 of 18 shared tests, including DeepSWE v1.1 (77.9% vs 74.2%). Opus 5.5 leads on Terminal-bench 4.0, FrontierSWE v2, PostTrainBench and Terminal-Bench Science. These are Google's numbers, not independent tests.

Is Gemini 4 Argon better than GPT-6 Astra?

Google's table puts Argon ahead of GPT-6 Astra on 14 of 19 tests, tied on CWE-bench v1, and behind on FrontierSWE v2, Terminal-bench 4.0, Terminal-Bench Science and OSWorld-2.0.

What is Gemini 4 Argon's DeepSWE score?

77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 (74.2%), GPT-6 Astra (74.1%) and Claude Fable 5.1 (67.4%), according to Google.

Where does Gemini 4 Argon lose?

On Google's own table it trails on FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0. Most of those involve agents working in a terminal or operating a computer.

Are there independent Gemini 4 Argon benchmarks?

Not yet. Argon only went to a small group of testers on September 30, 2026. Expect leaderboards like LMArena and Artificial Analysis once the API opens to paid customers.

Sources: Google's announcement and benchmark table, Google DeepMind evaluation methodology, The Next Web, 9to5Google.