Google finally shipped a new flagship this week. Gemini 4 Argon was announced on September 30 — the first new Gemini generation in nearly a year, after the June model that never arrived and a summer of uncomfortable press about DeepMind morale. I have learned to read launch-week benchmark charts the way I read vendor load-test reports: politely, and then I check who ran the test. So forget the crown for a minute. What actually matters to me, as someone who ships AI systems inside a bank, is the shape of this release: defenders first, everyone else later, and a cyber-capable version going out the door with the guardrails off to a vetted few.
That pattern should look familiar. It is the same two-track playbook every frontier lab is converging on.
What’s new
Argon is Google’s new flagship, built for what the company calls long-horizon work: real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. The headline specs from Google’s announcement:
- 1M-token output limit, up from 64K on prior Gemini models — the model can think and generate hundreds of thousands of tokens in a single trajectory instead of chunking hard problems into pieces
- Frontier positioning in coding, business workflows, long-video understanding, and cyber defense rather than chatbot chat
- Introductory API pricing of $2 per million input tokens / $10 per million output, with cached input at 95% off
- Rolling out first to trusted cyber defenders through the Fairwind Program, with the US government in a voluntary pre-release access process before wider release
The most convincing part of the announcement is not a benchmark — it is Google’s own dogfooding. Argon agents are reportedly doing genuinely unglamorous work inside the company: finding 300+ TiB of memory savings from fleet telemetry, beating a published quantum-optimization baseline by 40% in minutes, and grinding through C/C++ to Rust migrations — including 800K+ lines of the Fuchsia Zircon kernel under rigorous audit, and a libgav1 video-decoder port where the agents’ safe Rust runs 2.7x faster than the previous Rust port with identical output. Anyone who has nursed a large migration knows the rewrite is the easy part and the auditing is the work.
On safety, Google lists four tracks: misuse/CBRN refusal under its Frontier Safety Framework with internal-activation monitoring, best-yet resistance to indirect prompt injection (leading on Gray Swan’s IPI benchmark, per Google), chain-of-thought and action monitors that can halt execution mid-task, and hardened sandboxes. One detail I liked: the monitors that watched Argon’s training runs reported to a dedicated incident team, and the findings were deliberately not fed back into training — so the model never got a gradient signal teaching it how to evade its own overseers. That is a small decision with large implications, and more labs should say out loud whether they do the same.
Benchmarks
Google’s published numbers (vendor-reported, via press charts):
| Benchmark | Gemini 4 Argon | Predecessor / rival |
|---|---|---|
| DeepSWE v1.1 (long-horizon SWE) | 77.9% | 74.2% (Claude Opus 5.5), 74.1% (GPT-6 Astra) |
| AutomationBench (business workflows) | 51.3% | 42.5% (Claude Opus 5.5) |
| LVBench (long video) | 91.7% | — (claimed SOTA) |
| CWE-bench v1 (vuln remediation) | 68% | tie (GPT-6 Astra) |
| Terminal-Bench 4.0 | — (Opus 5.5 keeps it) | rival lead |
| FrontierSWE v2 | — (Astra keeps it) | rival lead |
Sources: Google’s announcement charts, reported via SiliconANGLE, VentureBeat, and The News (Oct 1). VentureBeat counts 12 of 18 charted benchmarks led outright by Argon.
Now the honest part. These are launch-week vendor numbers with no independent replication yet, and the early third-party signals are mixed: Futurum’s analysts put Argon roughly at OpenAI’s level and behind Anthropic, and Bloomberg reported — anonymously, and denied by Google — that some Googlers found the model underwhelming in internal testing. The DeepSWE gap over Opus 5.5 and Astra is a few points, not a generation. Treat the crown as provisional until Artificial Analysis and the eval community finish their runs. The 1M output limit is the spec I would test first — if long single-trajectory rollouts actually hold coherence, that changes agent design more than three points on any chart.
Pricing
Official pricing is published, so here is the table ($ per million tokens, input / output):
| Model | Input | Output |
|---|---|---|
| Gemini 4 Argon (intro) | $2.00 ($0.10 cached) | $10.00 |
| Gemini 4 Argon (post-intro) | $4.00 | $20.00 |
| Claude Opus 5.5 | $4.00 | $20.00 |
| Claude Sonnet 5.5 | $2.00 | $10.00 |
| GPT-6 Sol | $2.00 | $10.00 |
| GPT-6 Astra (Standard) | $10.00 | $50.00 |
Sources: Google’s announcement for Argon; VentureBeat’s September 30 pricing table and SiliconANGLE for the comparisons; Astra per my September notes.
Two things stand out. First, the intro price parks Argon exactly at the Sonnet/Sol tier while claiming Opus/Astra-class capability — classic penetration pricing, and Google is open about it expiring to Opus-level rates. Second, that 95% cached-input discount is the most aggressive on the board. If you run agents with big pinned contexts, do the math on cached input before you compare list prices.
Availability
The rollout is deliberately staged, and the staging is the story:
- Now: vetted cyber defenders in the Fairwind Program (opened September 3 with Gemini 3.8 Flash Cyber; 650+ organizations including CrowdStrike and Palo Alto Networks, per SiliconANGLE) plus the US government’s voluntary pre-release process
- Next, no date: paid API customers and Google AI Ultra subscribers
- The cyber-defense version goes to trusted defenders and Google’s own teams without cyber guardrails, so they get full defensive capability
- No free tier, no open weights mentioned
The Wiz anecdote is doing a lot of marketing work here — Google’s own security firm says Argon found a critical vulnerability exposing personal information in hospital-used healthcare software that earlier frontier models missed — but even discounted for venue, it is the right kind of evidence: a named deployment finding a named bug class, not a chart.
What it means
Three things stand out to me.
First, the gated two-track release is now the industry default. OpenAI did it with Astra’s Daybreak program and a restricted cyber track; Google is doing it with Fairwind and an unguarded defender build. Everyone converging on the same design independently tells you it is load-bearing, not PR: the models are now good enough at offense that broad release on day one is off the table. Those of us in regulated industries will recognize the shape — it is our own change-advisory process, applied to model weights. Expect your vendor’s “trusted tester” form to become as routine as a firewall exception request.
Second, Google is running the opposite experiment from OpenAI on monitorability, and I am here for the contrast. A month ago OpenAI’s chief scientist told the Astra launch call that monitorability gets harder as models get more capable, because better models use fewer language tokens. Now Google is publishing essays urging the industry to preserve reasoning transparency, deploying CoT monitors with kill switches, and refusing to train against its own overseers. One lab says the audit trail is going dark by necessity; the other says keeping it legible is a design choice. In my world that is not philosophy — my model risk board signs off on systems we can explain. I know which vendor’s homework I would rather grade.
Third, price it per completed task, not per token. At intro pricing Argon sits in the middle of my routing cascade, not the top: Sonnet/Sol money for claimed Opus/Astra-class long-horizon work, with the cheapest cached reads anywhere. If the 1M-output coherence claim survives independent testing, the model that finishes a migration slice in one trajectory instead of five stitch-together runs wins on total cost even after the price doubles. Until then it is a wait-and-see tier — and note the context: an October industry roundup citing Anthropic claims open-weight GLM-5.3 is uncomfortably close to frontier on autonomous exploit benchmarks, which is exactly why every lab gates cyber first.
Worth noting from the same week: OpenAI shipped the incremental GPT-6.1 Sol on September 29 — a point release, not an answer. The real story is Google rejoining a race some had counted it out of after the cancelled 3.5 Pro and the delayed summer.
My verdict: do not re-route anything yet. Wait for independent evals, then test Argon on your longest agentic workflow and measure cost per completed task with caching on. A year-long wait buys Google exactly one chance at a first impression with practitioners — and first impressions, like model outputs, should be verified.