Google announced Gemini 4 on September 30, 2026. Its full name is Gemini 4 Argon, and it is Google's new frontier model, meaning the most capable one it makes. You can't use it yet. And the most useful thing Google published alongside it is a scorecard that shows where it wins, and also where it loses.
What changed, and who it applies to. Google DeepMind's Koray Kavukcuoglu announced Argon in a post on Google's blog. For now, Google is giving it to vetted cybersecurity partners and testers. Everyone else, including teams that use Gemini at work, waits until Google opens it “as soon as possible”, with no date. Below are 7 facts from Google's own pages, in the same order as our carousel. Where the reasoning is ours, it says "Our read".


1. Cyber defenders first. Everyone else waits.
Date and source: September 30, 2026. Google's announcement and its Fairwind Program page.
What Google says. “Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.” The Fairwind page describes who those defenders are: “The Fairwind Program gives high-priority defenders (like governments, healthcare providers, and telecommunications services) early access to advanced models”. Google adds: “We conduct background checks on organizations that apply, to verify security history and analyze their record of ethical operations.” And inside those partners, the use is narrow: “Organizations may only grant Gemini 4 Argon access to internal cybersecurity, incident response, or penetration testing teams, and must track employee access and use.” For everyone else, Google says it will make Argon available “as soon as possible”.
What it means for you. Our read: there is nothing to switch to this week. Argon is not in the Gemini app or the public API for ordinary accounts yet.
Try this. Don't move tools, budgets or workflows for Argon until you can actually use it.
2. Next in line: paid API and AI Ultra
Date and source: September 30, 2026. Google's announcement.
What Google says. When access widens, Google names 2 groups together: “We’re grateful for the initial cohort of cyber defenders and trusted testers whose real-world evaluations and feedback will help us strengthen our systems before we release to developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers.” The announcement doesn't mention the free or Pro Gemini plans.
What it means for you. Our read: if your team uses Gemini on a free or Pro plan, expect to wait longer than developers on paid API plans and Google AI Ultra subscribers.
Try this. Check which Gemini plan and which API tier your team actually pays for.
3. $2 in, $10 out, at an introductory price
Date and source: September 30, 2026. Google's announcement and its footnote.
What Google says. “Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.” The footnote adds: “After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.” Google doesn't say how long the introductory period lasts. A token is a small chunk of text.
What it means for you. Our read: these are API prices, charged per token, not a subscription. Budget at the higher price. A tool or workflow that only makes sense at $2 and $10 may not make sense once the introductory period ends and the price doubles.
Try this. If your team uses the API, price 1 real job at $4 in and $20 out per 1M tokens before you plan around Argon.
4. Answers up to 1M tokens, up from 64K
Date and source: September 30, 2026. Google's announcement.
What Google says. “we are significantly expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens.” The output limit is how much the model can write back in one go. Google says the room lets the model “think deeply and generate hundreds of thousands of tokens in a single trajectory”.
What it means for you. Our read: a long report, a full rewrite or a big analysis could come back in 1 pass instead of being stitched together from several prompts. At the API price, a full 1M-token answer would cost $10 at the introductory price and $20 after it, plus input tokens.
Try this. List 1 long job that takes your team several prompts today, and save it for testing.
5. First on all 4 knowledge-work tests
Date and source: September 30, 2026. Google DeepMind's evaluation table.
What Google says. Google's table compares Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. Argon leads all 4 of its knowledge-work tests: the Vals Index (68.9%), Zapier's AutomationBench (51.3%), Vals Finance Agent v2 (65.4%) and Harvey's Legal Agent Benchmark (19.6%). In the announcement: “On AutomationBench, Zapier’s benchmark measuring end-to-end execution across core business functions, Argon ranks #1 with a score of 51.3%.” The legal lead is the biggest, but every model scores under 20% on that test.
All 4 of those scores come from Vals AI and Zapier, not from Google's own runs, per its methods note.
What it means for you. Our read: research, finance and drafting work is where to test it first.
Try this. Pick 3 real knowledge tasks, such as a research brief, a financial summary and a long report, to try on day 1.
6. First on 2 coding tests, last on the other 2
Date and source: September 30, 2026. Google DeepMind's evaluation table.
What Google says. The same table puts Argon 4th of 4 on 2 coding tests: FrontierSWE v2 (55.0%, where GPT-6 Astra leads with 65.5%) and Terminal-bench 4.0 (57.4%, where Claude Opus 5.5 leads with 66.4%). It leads the other 2 coding tests, DeepSWE v1.1 (77.9% in Google's own run, which Google calls a new state of the art) and Vibe Code Bench (91.9%). It also trails on 3 more: PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0.

What it means for you. Our read: no model wins every job, even on its maker's own scorecard.
Try this. Keep your current coding model until you have compared both on your own work.
7. Google picked the tests. Pick your own.
Date and source: September 30, 2026. Google DeepMind's evaluation table and its methods note.
What Google says. Google chose the 19 scores in its table. Its methods note, at deepmind.google/models/evals-methodology/gemini-4-argon, says Google ran 10 of Argon's scores itself. The other 9 come from outside sources, including all 4 knowledge-work tests: “Vals Index results are sourced from Vals AI.” AutomationBench results come from Zapier’s official public leaderboard. For the rivals, the note names a source test by test. Most rival scores come from the same outside leaderboards, Google ran 5 of the tests on all 4 models itself, and only 3 rival scores are the rivals' own reported numbers.
What it means for you. Our read: treat the table as a reason to test, not a reason to switch.
Try this. Save your 3 test tasks, and run them the week Argon opens to you.
Google's full scorecard
Google's table, from deepmind.google/models/gemini/, read September 30, 2026. Google chose these tests; its methods note says which scores it ran itself and which come from outside leaderboards or the rivals' own reports. The methods note says Fable 5.1 is missing from the Agent's Last Exam leaderboard, and Anthropic reports OSWorld-2.0 only for both subsets combined, so Argon's OSWorld comparison is with 1 rival only. "No score" means Google listed none. Argon is first, or tied for first, on 14 of the 19 scores.
| Benchmark | Area | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|---|
| Vals Index | Knowledge work | 68.9% | 63.1% | 65.8% | 67.0% |
| AutomationBench | Knowledge work | 51.3% | 41.4% | 31.4% | 42.5% |
| Vals Finance Agent v2 | Knowledge work | 65.4% | 53.5% | 58.9% | 58.6% |
| Harvey's Legal Agent Benchmark | Knowledge work | 19.6% | 5.4% | 6.7% | 3.8% |
| DeepSWE v1.1 | Agentic coding | 77.9% | 74.1% | 67.4% | 74.2% |
| FrontierSWE v2 | Agentic coding | 55.0% | 65.5% | 56.3% | 62.3% |
| Vibe Code Bench | Agentic coding | 91.9% | 89.6% | 90.3% | 90.3% |
| Terminal-bench 4.0 | Agentic coding | 57.4% | 58.2% | 57.9% | 66.4% |
| PostTrainBench | ML engineering | 45.3% | 44.3% | 40.2% | 49.3% |
| Terminal-Bench Science 0.1 | Science and math | 57.6% | 68.1% | 52.6% | 63.3% |
| LABBench 2 | Science and math | 88.8% | 85.4% | 68.6% | 73.1% |
| RiemannBench | Science and math | 76.0% | 72.0% | 65.6% | 69.6% |
| GraphWalks, up to 128k | Long context | 99.7% | 98.7% | 91.4% | 90.6% |
| GraphWalks, 256k to 1M | Long context | 84.2% | 71.8% | 65.0% | 66.8% |
| Agent's Last Exam | Computer use | 39.5% | 34.2% | No score | 38.2% |
| OSWorld-2.0, offline subset, partial score | Computer use | 69.2% | 72.6% | No score | No score |
| Chartography | Multimodal understanding | 71.6% | 71.0% | 46.2% | 66.3% |
| LVBench | Multimodal understanding | 91.7% | 87.5% | 79.7% | 83.7% |
| CWE-bench v1 | Cybersecurity | 68.0% | 68.0% | 58.0% | 67.0% |
Before Argon reaches you: 3 steps
- Pick your test tasks now. Choose 3 real jobs your team does every week, such as a research brief, a financial summary and a long report, and write down what a good result looks like.
- Budget at the full price. If your team uses the API, plan on $4 per 1M input tokens and $20 per 1M output tokens, not the introductory $2 and $10.
- Compare before you switch. Run the same 3 tasks on Argon and on the tool you use now, and keep whichever does your work better.
What Google has not said
Checked in the announcement, the Gemini model page with its table, the methods note and the Fairwind Program page, all read September 30, 2026.
A release date. Google says “as soon as possible” and gives no date for developers, enterprises or consumers.
How long the introductory price lasts. The footnote says the $4 and $20 price applies “After the introductory period expires”, with no length given.
Anything about the free or Pro Gemini plans. The announcement names paid API customers and Google AI Ultra subscribers first, and doesn't mention the other plans.
That the 1M figure is a context window. It is the output limit: how much Argon can write back, up from 64K. The GraphWalks row for 256k to 1M tokens tests input length, a separate thing: the methods note says it uses problems “with context lengths between 256k and 1M tokens”. None of the 4 pages states Argon's context window.
That it has no guardrails for everyone. Google says it is releasing Argon without cyber guardrails only “For trusted defenders and our own internal teams at Google”. Before a broad release, Google says the model “is designed to refuse harmful requests”.
Why these 19 tests. Google chose the 19 scores it shows. Its methods note explains where each one came from, but not why these tests were picked.
Sources: verified September 30, 2026 against Google's announcement “Gemini 4 Argon: our next era of frontier intelligence” by Koray Kavukcuoglu, SVP, Google DeepMind and Chief AI Architect, Google; the Gemini model page on deepmind.google with its evaluation table; Google DeepMind's evaluation methods note for Gemini 4 Argon; and the Fairwind Program page on deepmind.google. Quotations are exact. Lines marked Our read are MichaelBennett.co's reasoning, not Google's.
Working out where AI actually fits in your marketing?
I write these while building the systems behind them: measurement, creative pipelines, and agents that do real work. Connect on LinkedIn and tell me what you are working on. That is where these conversations start.
Connect on LinkedIn