Mastodon Skip to content
LIVE - NYSE/-/- CRYPTO/OPEN/24/7
BTC$84,333▼ 0.07%ETH$2,685▼ 0.21%SOL$117.44▼ 1.97%TOTAL CRYPTO$2.89T▼ 2.81%S&P 5007,634.35▲ 0.04%NASDAQ26,791.32▲ 2.65%DOW50,733.64▼ 3.85%GOLD4,201.20▼ 4.44%WTI92.30▲ 2.31%BRENT101.24▲ 6.96%EUR/USD1.1241▼ 3.24%USD/JPY157.55▼ 1.38%DXY102.05▲ 2.39%
Technology

Google Rolls Out Gemini 4 Argon to Cyber Defenders First

Google's new flagship model outperforms rivals on most benchmarks but remains limited to the Fairwind security program, with wider release yet to come.

Pexels – Markus Winkler

Google has launched Gemini 4 Argon, its new flagship AI model, with access restricted for now to a select group of security organizations through the Fairwind Program.The company says Argon outperforms OpenAI’s and Anthropic’s top models on most benchmarks, sometimes by wide margins, though the results are less consistent in agentic coding, where it trails on two prominent tests. Argon is the successor to Gemini 3.5 Pro, announced at Google’s I/O conference in May but never actually shipped; the launch slot was instead filled by a series of lighter Flash models over the summer.

Cybersecurity first

Argon was trained specifically for defensive cyber work. Google says the model can autonomously find, validate and patch critical software vulnerabilities across complex codebases spanning 20 programming languages. In an early demonstration, Wiz, the cloud security company Google acquired for $32 billion in March, used Argon through its Scan for Good initiative and surfaced a critical vulnerability in healthcare software used by hospitals worldwide, a flaw earlier frontier models had missed.Rollout began through Fairwind, Google’s controlled-access program for trusted defenders, which includes government agencies, critical infrastructure operators and security partners such as CrowdStrike, Palo Alto Networks and Wiz. The program launched on Sept. 2 and already counts more than 650 organizations across government, healthcare, telecom, energy and finance. For Fairwind participants and Google’s internal teams, releases come without cyber guardrails so the model’s full detection and patching capabilities can be used unrestricted.”Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google,” the company said in a blog post Wednesday.

Where the benchmarks land

In Google’s own comparisons, Argon takes top or tied place in 13 of 18 benchmark categories against Anthropic’s Claude Opus 5.5 and Claude Fable 5.1, as well as OpenAI’s GPT-6 Astra. Several of those leads are narrow, under two points, but a few stand out: Argon scored 51.3% on Zapier’s AutomationBench, almost nine points ahead of Opus 5.5, and 84.2% on the GraphWalks long-context test for inputs between 256,000 and 1 million tokens, more than 12 points ahead of GPT-6 Astra.Coding results are less flattering. Argon leads on DeepSWE v1.1, at 77.9%, and on Vibe Code Bench, but comes last among the comparison models on FrontierSWE v2 and Terminal-Bench 4.0, where the competition’s strongest results beat it by roughly 10 and 9 points respectively.

Category Benchmark Gemini 4 Argon GPT-6 Astra Claude Opus 5.5
Knowledge work Vals Index 68.9% 63.1% 67.0%
Knowledge work AutomationBench 51.3% 41.4% 42.5%
Long context GraphWalks 256k-1M 84.2% 71.8% 66.8%
Agentic coding DeepSWE v1.1 77.9% 74.1% 74.2%
Agentic coding FrontierSWE v2 55.0% 65.5% 62.3%
Cybersecurity CWE-bench v1 68.0% 68.0% 67.0%

One caveat with these numbers: the CWE-bench v1 leaderboard measures each model together with its own agent tooling, so rivals ran through OpenAI’s Codex and Anthropic’s Claude Code harnesses. Comparisons of a model and its scaffolding are not purely about the model itself, a distinction that analysts flagged in coverage of the results.

Longer output, higher price

One notable capability is output length. One million input tokens is now standard among frontier models, but Argon can also generate up to 1 million output tokens, up from the 64,000 available in earlier Gemini releases. Google says this matters for long code migrations, extended documents and multi-hour agent runs that previously had to be split or restarted.Google set an introductory price of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward. That later output rate matches what Anthropic charges for Claude Opus 5.5, so Argon enters the frontier price tier rather than undercutting it.

“Argon can autonomously find, validate, and patch critical software vulnerabilities.”

Who can actually use it

Nobody outside Fairwind yet. Google says it will gather feedback from early testers and iterate on Argon’s guardrails before opening the model more broadly. Paid API customers and Google AI Ultra subscribers get access next, followed by developers, enterprises and general consumers. No date has been given for any of those steps.The rollout reflects a pattern across the industry, where models with strong offensive or dual-use cyber capability are released first to security partners in gated programs, the same approach rivals apply with their own access-limited cyber releases and predecessor models built for the same purpose.

Why the staged release matters

Limiting a powerful cyber-capable model to a vetted set of defenders has a specific rationale. A model that can find and patch vulnerabilities autonomously is also a model that, pointed the other way, could find them for exploitation. Gated releases give labs a window to observe real-world use and adjust safety boundaries before general availability. Wiz already reported practical gains, using Argon via its Scan for Good program to scan healthcare software used by hospitals worldwide and surface a critical flaw that peers had missed.For the broader public, the message is that Argon is real but not yet reachable without a subscription tier or a program slot. Its pricing, above Anthropic’s cheapest tiers and aimed at professional workloads, suggests Google is treating Argon as a model for complex work rather than a mass-market release.At stake is competitive position. Google wants Argon to be the model of choice for enterprise coding, legal drafting, financial research and cybersecurity, and the benchmark table backs that up, at least on its own measurements. Independent testing will be the next check, since Google’s own harnesses and benchmark selections can favor models built and trained in the same shop, and third parties will run their own evaluations before enterprise buyers commit.For now, Argon has set a strong opening benchmark record while most of the world waits on a release date, and Google’s product roadmap puts paid tier access sometime after Fairwind feedback, a sequence that has left some developers watching competitor releases instead.Sources: TechCrunch, Sept. 30; The New Stack, Sept. 30; Google DeepMind blog

Share: X