Gemini 4 Argon · Research

Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5: Is Google's New Model Actually Better?

Data graphic: Gemini 4 Argon scores 52.6 on the Artificial Analysis Intelligence Index, level with GPT-6 Astra 52.7 and behind Claude Opus 5.5 57.6 . Cost per task: Argon $1.99 at intro price, about $3.98 at full price, Astra $3.26, Opus 5.5 $5.98. The leaked 88% DeepSWE score sits beside the real 77.9%, and the $2/$10 intro price beside the later $4/$20.
PM

Supply chain security reporter · Updated Oct 1, 2026, 2:02 AM EDT

The leaked 'Gemini 4 Pro' was billed as an Astra and Opus killer. The real Gemini 4 Argon wins Google's tests, ties Astra on neutral ones, and is cheap for now.

For two weeks the story was "Gemini 4 Pro beats GPT-6 Astra and Claude." A model believed to be Google's next flagship appeared on Arena on 17 September under the name gemini-3.8-flash, and testers posted numbers that left everything else behind. On 30 September Google announced the real thing. It is called Gemini 4 Argon, not Pro. It is good, but not as good as the leak suggested, it is only cheap for now, and most people cannot use it yet.

The leak versus the launch

Arena leak ("Gemini 4 Pro")Gemini 4 Argon, as announced
DeepSWE v1.188%77.9%
Output limit256K tokens1M tokens (up from 64K)
Input context"10 million tokens"not specified
Price per 1M tokens, in / out$2.25 / $11.25$2 / $10 introductory, then $4 / $20
AvailabilityArena, anonymouslyTrusted cyber defenders only

The leaked coding score was about ten points too high, and the 10-million-token context does not appear anywhere in Google's announcement. The leaked price was close. Unverified Arena screenshots are a poor guide to a release, and this one was off on the numbers that drew the most attention.

What Google claims

Google published 18 benchmarks. By VentureBeat's count, Argon leads outright on 12, ties on one, and trails GPT-6 Astra on three and Claude Opus 5.5 on two. The headline results:

BenchmarkGemini 4 ArgonGPT-6 AstraClaude Opus 5.5
DeepSWE v1.1 (software engineering)77.9%74.1%74.2%
AutomationBench (business tasks)51.3%41.4%42.5%
Vals Finance Agent v265.4%53.5%58.6%
Harvey Legal Agent Benchmark19.6%5.4%3.8%
GraphWalks (long-context reasoning)84.2%71.8%66.8%
LVBench (long video)91.7%87.5%83.7%
CWE-bench v1 (vulnerability fixing)68% (tie)68% (tie)67%

These are Google's chosen benchmarks, run on Google's settings. The coding lead is real but small, about 3.7 points on DeepSWE. The gaps in legal, finance and long-context work are much larger.

What independent testing says

Artificial Analysis ran Argon through its own index on launch day, and the picture changes.

Artificial AnalysisGemini 4 Argon (High)GPT-6 Astra (Max)Claude Opus 5.5 (Max)
Intelligence Index52.652.757.6
Terminal-Bench (terminal coding)57.1%59.1%59.6%
Humanity's Last Exam57.1%54.7%61.4%
GDPval-AA (work tasks, 44 jobs; Elo)1,6111,5421,846
SciCode61.8%56.5%66.9%
Cost to run one task$1.99$3.26$5.98

On a neutral harness, Argon ties Astra and trails Opus 5.5 by about five points. Its advantage is cost: Artificial Analysis says Astra-level performance costs about 39% less per task on Argon, and Opus 5.5's extra five points cost three times as much.

The price comes with a deadline

Argon's introductory rate is $2 input and $10 output per million tokens. That is a fifth of Astra's $10 / $50 list price and half of Opus 5.5's $4 / $20. Cached input is 95% off.

The regular price, once the introductory period ends, is $4 / $20, exactly what Opus 5.5 costs. Artificial Analysis's $1.99 per task was measured at the introductory rate. Every token price doubles, so by our arithmetic the same workload would cost about $3.98 per task at full price. That is more than Astra's $3.26 and still below Opus 5.5's $5.98. Google has not said when the introductory period ends.

You probably cannot use it yet

Argon is rolling out first to "trusted cyber defenders" through Google's Fairwind Program, and Google is taking part in the US government's voluntary pre-release testing process. Paid API customers and Google AI Ultra subscribers are next, with no date beyond "as soon as possible." Unless you are in Fairwind, any decision about Argon is a decision about a model you cannot call.

Why the security launch matters

Google is launching Argon as a cyber-defence model first, and the security claims are the most concrete part of the announcement:

  • Vulnerability repair: 68% on CWE-bench v1, tied with GPT-6 Astra and one point ahead of Opus 5.5. That is level with the best, not ahead of it.
  • Prompt injection: Gray Swan's indirect prompt-injection test measured a 0.7% attack success rate for Argon, against 1.0% for Opus 5.5 and Fable 5.1. Google calls Argon its "most resilient model yet" against indirect injection.
  • Field results: Google says Wiz used Argon to find a critical flaw exposing personal data in healthcare software used by hospitals worldwide, and that Argon mapped an attack surface better than 3.8 Flash Cyber.
  • Guardrails: monitors watch Argon's reasoning and actions and can stop execution. Its safeguards specifically target cyber and CBRN misuse.

For a security team the useful takeaway is narrower than the marketing. Argon matches the best on fixing vulnerabilities and slightly improves on injection resistance. A 0.7% attack success rate still means injection gets through sometimes. Agents that read untrusted content still need sandboxing and least-privilege tools whichever model runs them.

So, is it good?

Yes, with conditions.

  • Coding agents: Argon is roughly level with Astra and behind Opus 5.5 on independent tests. It is cheaper than both while the introductory price lasts. At full price it costs the same per token as Opus 5.5 and scores lower.
  • Legal, finance and long-context work: this is where Google's own lead is largest. It is worth testing as soon as you get access, but confirm the results on your own documents first.
  • Security teams: if you can get into Fairwind, it is a strong vulnerability-repair model with the best published injection score of the three. If not, Astra matches it on CWE-bench today.
  • "Gemini 4 Pro" as the leaks described it does not exist. Benchmark Argon on your own workload before you switch.

Sources

  • Google, "Gemini 4 Argon: our next era of frontier intelligence", 30 Sep 2026
  • VentureBeat, "Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic — but in limited release"
  • Trending Topics, Artificial Analysis results: "Gemini 4 Matches GPT-6 Astra but Trails Opus 5.5"
  • Help Net Security, "Google says Gemini 4 Argon can find and patch critical software flaws", 1 Oct 2026
  • TechBriefly, "Leaked Gemini 4 Pro benchmarks show it beating GPT-6 Astra and Claude", 21 Sep 2026