Google published 19 benchmark results for Gemini 4 Argon, comparing it with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. Argon comes first on 13 of them, ties on one, and trails on five. Here is the whole table in plain English, including the parts the headlines skip.
These are Google's numbers. Every score on this page comes from Google's own announcement and its evaluation methodology. No independent leaderboard (such as LMArena or Artificial Analysis) had scored Argon when we published this, and we'll add those results when they appear.
The short version
In one sentence: Argon is strongest at long knowledge work, reading huge inputs and understanding video, is at or near the top at software engineering, and is not the best model for terminal-heavy agent work, where Claude Opus 5.5 leads.
Gemini 4 Argon coding benchmarks
DeepSWE v1.1
Real-world software engineering tasks. Higher is better.
| Agentic coding | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 |
|---|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% |
| Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| PostTrainBench (ML engineering) | 45.3% | 44.3% | 40.2% | 49.3% |
Argon wins the headline coding test, DeepSWE, by 3.7 points, and edges out everyone on Vibe Code Bench. But it is last of the four on FrontierSWE v2, and nine points behind Claude Opus 5.5 on Terminal-bench 4.0, which measures getting real work done in a command line. If your work is mostly agents running in a terminal, Opus 5.5 still has the edge on Google's own chart.
Knowledge work: Argon's biggest lead
| Knowledge work | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 |
|---|---|---|---|---|
| Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| AutomationBench (score) | 51.3% | 41.4% | 31.4% | 42.5% |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% |
| Harvey's Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% |
This is where Argon pulls away. It beats the next-best model on AutomationBench by almost 9 points and on Vals Finance Agent by 6.5. On Harvey's legal agent test the absolute scores are low for everyone, but Argon's 19.6% is roughly three times the next model. For finance, legal and business-automation work, these are the numbers to watch once independent testing starts.
Long context and video understanding
| Test | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 |
|---|---|---|---|---|
| GraphWalks, up to 128K (BFS, F1) | 99.7% | 98.7% | 91.4% | 90.6% |
| GraphWalks, 256K to 1M (BFS, F1) | 84.2% | 71.8% | 65.0% | 66.8% |
| Chartography | 71.6% | 71.0% | 46.2% | 66.3% |
| LVBench (long video) | 91.7% | 87.5% | 79.7% | 83.7% |
The GraphWalks results matter because they test whether a model can still find its way around a truly enormous input. Between 256K and 1M tokens, Argon keeps 84.2% while the others fall to between 65% and 72%. Add the LVBench lead on long videos and chart reading, and Argon looks built for feeding in a lot at once. Curious what that means for video? Read making videos with Gemini 4 Argon.
Science and math
| Science and math | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 |
|---|---|---|---|---|
| Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% |
| LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% |
| RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% |
A split result: Argon leads on lab-science questions (LABBench 2) and advanced math (RiemannBench), but GPT-6 Astra wins clearly when science work has to happen in a terminal.
Computer use
| Computer use | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 |
|---|---|---|---|---|
| Agent's Last Exam (pass rate) | 39.5% | 34.2% | not reported | 38.2% |
| OSWorld-2.0 (offline subset, partial score) | 69.2% | 72.6% | not reported | not reported |
Close race. Argon narrowly leads Agent's Last Exam and trails GPT-6 Astra on OSWorld-2.0. Google hasn't published how Argon's computer-use features work in the API yet, so treat these as a preview.
Cybersecurity: why Argon ships to defenders first
| Security test | Gemini 4 Argon | Compared with |
|---|---|---|
| CWE-bench v1 | 68.0% | GPT-6 Astra 68.0%, Opus 5.5 67.0%, Fable 5.1 58.0% |
| Real-world vulnerability discovery (20 languages) | 85.8% | Gemini 3.8 Flash Cyber 71.0% |
| Wiz penetration test benchmark | 70.9% | Gemini 3.8 Flash Cyber 58.2% |
| Gray Swan prompt injection, attack success (lower is better) | 0.7% | Opus 5.5 1.0%, Fable 5.1 1.0%, GPT-6 Astra 8.5% |
Argon is very good at finding security holes, and Google compares it here with its own security-tuned Gemini 3.8 Flash Cyber model. That is exactly why the first people to get it are vetted defenders in Google's Fairwind Program. It is also the hardest model on Google's chart to trick with hidden instructions: just 0.7% of prompt-injection attacks succeeded after 15 tries.
Want the models you can actually use today?
Veo 3.1, Gemini Omni, Nano Banana 2, Seedance and Kling are all live on Lumeta, in one account.
How to read these numbers
- Vendors pick their tests. Every lab leads with the benchmarks it wins. To Google's credit, its table also shows five losses.
- Small gaps are noise. A difference of one or two points (like Vibe Code Bench, 91.9% vs 90.3%) may not show up in your daily work.
- Your task is the real test. Run the same five real prompts from your own work through each model when Argon opens up. That tells you more than any table.
- Not everyone is convinced yet. Press coverage has reported some skepticism inside Google about Argon's real-world coding. Independent results will settle it.
For the full list of test descriptions and settings, see Google DeepMind's evaluation methodology page linked below.
Benchmark questions
Is Gemini 4 Argon better than Claude Opus 5.5?
On Google's chart, Argon beats Claude Opus 5.5 on 14 of 18 shared tests, including DeepSWE v1.1 (77.9% vs 74.2%). Opus 5.5 leads on Terminal-bench 4.0, FrontierSWE v2, PostTrainBench and Terminal-Bench Science. These are Google's numbers, not independent tests.
Is Gemini 4 Argon better than GPT-6 Astra?
Google's table puts Argon ahead of GPT-6 Astra on 14 of 19 tests, tied on CWE-bench v1, and behind on FrontierSWE v2, Terminal-bench 4.0, Terminal-Bench Science and OSWorld-2.0.
What is Gemini 4 Argon's DeepSWE score?
77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 (74.2%), GPT-6 Astra (74.1%) and Claude Fable 5.1 (67.4%), according to Google.
Where does Gemini 4 Argon lose?
On Google's own table it trails on FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0. Most of those involve agents working in a terminal or operating a computer.
Are there independent Gemini 4 Argon benchmarks?
Not yet. Argon only went to a small group of testers on September 30, 2026. Expect leaderboards like LMArena and Artificial Analysis once the API opens to paid customers.
Sources: Google's announcement and benchmark table, Google DeepMind evaluation methodology, The Next Web, 9to5Google.