Gemini 4 Argon: What It Is and Who Can Use It
Gemini 4 Argon leads most of Google's own benchmarks, but almost nobody can use it yet. What it does, what it costs, and what to do this week.

Google finally shipped its next flagship, and you can't use it.
That's the whole story of Gemini 4 so far, in one line.
Gemini 4 is real. Its first model is called Gemini 4 Argon, Google DeepMind announced it on September 30, 2026, and on Google's own benchmark sheet it beats the best models from OpenAI and Anthropic more often than it loses. It is also, for now, rolling out only to a small group of trusted cybersecurity teams.
So you have a model everyone is talking about and almost nobody is allowed to touch.
That gap is where people make bad decisions. They cancel a subscription, rebuild a workflow, or promise a client something based on a press release.
Don't.
Gemini 4 Argon is worth preparing for this week and not worth betting your work on until you've run it on your own tasks.
Here's what it is, what it costs, where it wins, where it still loses, and the exact test to have ready for the day the door opens.
Gemini 4 is here, it's called Argon, and it's behind a locked door
Start with the plain facts, because the name alone has confused plenty of people.
Gemini 4 is the generation. Argon is the first model in it. Google calls it "our new frontier model" in the official announcement on the Google blog, written by Koray Kavukcuoglu, who runs Google DeepMind as SVP and serves as Google's Chief AI Architect.
The pitch in that post is specific. Argon is "built to sustain deep reasoning across complex, long-horizon workflows." Google names the areas where it says the model leads: real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
It's also bigger than what came before it.
A Google spokesperson told Reuters that Argon is larger than Google's previous line of advanced Pro models. That matters, because it tells you Google didn't just tune the old model and give it a new number.
Now the part that decides what you do next.
Access is limited to a set of trusted cyber defenders through what Google calls its Fairwind Program. There is no general API access announced in the post, no date for the Gemini app, and no date for developers. Google told Axios it will expand access after more testing, with paying subscribers first in line.
That last detail is the one to remember.
If you pay for Gemini today, you're probably closer to the front of the line than a developer with an API key. If you don't, you're waiting behind both.
Here's what that means for you this week, laid out by who you are:
| If you are... | What you can do with Gemini 4 right now | What to do instead |
|---|---|---|
| A cybersecurity team in the Fairwind Program | Use it | Run it on real incident work and keep notes |
| A paying Gemini subscriber | Nothing yet | Watch for the rollout, build your test set now |
| A developer on the Gemini API | Nothing yet | Price your workloads at both rates (more below) |
| Everyone else | Read about it | Keep working with what you have and prepare to compare |
Nobody outside that first group should be changing anything today.
You should be getting ready to judge it fast, because the first week after a model opens is when you learn the most and spend the least.
Why Google gave cyber defenders the keys first
A flagship model launching to security teams before developers looks strange. It isn't, once you see the year Google has had.
Back in May, Google CEO Sundar Pichai said the company's next major model would arrive in June. That model was widely expected to be Gemini 3.5 Pro. It never came out. Axios reported in July that poor morale inside Google DeepMind was part of the delay, and a Google spokesperson has now told Reuters the company no longer plans to release Gemini 3.5 Pro at all.
So Google skipped a version.
In the meantime, it kept shipping smaller, cheaper Flash models while OpenAI and Anthropic released bigger ones. By the time Argon arrived, Google needed a win that looked deliberate rather than late.
Leading with cyber defense does two things at once.
First, it puts the model where the risk is highest and the users are most careful. A model that's very good at finding and fixing software vulnerabilities is, by definition, very good at understanding how to exploit them. Giving it to defenders first is a reasonable way to find out what it does in practice before anyone with worse intentions gets a turn.
Second, it fits a political moment. Reuters and Axios both report that Google is taking part in the U.S. government's voluntary process for pre-release access to frontier models. Launching through vetted partners lets Google say it shipped responsibly, and that's a story worth more to Google right now than a few weeks of developer goodwill.

You might be thinking this is just marketing dressed up as caution.
Some of it is. The timing landed one day after OpenAI's DevDay, and no lab launches a flagship by accident the day after its rival's biggest event.
But the safety numbers are real enough to take seriously. On the Gray Swan indirect prompt injection test, which measures how often an attacker can hijack a model through content it reads, VentureBeat reports that Argon allowed a 0.7% attack success rate. Claude Opus 5.5 came in at 1.0%. GPT-6 Astra was 8.5%. Grok 4.8 was 51.8%.
If you plan to let a model read your inbox or act on documents someone else wrote, that number matters more than any coding score. Prompt injection is how an agent that reads a malicious email ends up doing what the email says instead of what you said.
So the slow rollout costs you time. It may also be the reason the model is safe to hand real work when you finally get it.
The benchmark lead is real, and narrower than the headline
Every lab publishes the benchmarks it wins. Read them as a company's best case, then look for someone who isn't selling the model.
Google's own comparison came first. Per VentureBeat's count of the 18 benchmarks Google disclosed, Argon leads outright on 12 and ties for first on one. GPT-6 Astra leads outright on three. Claude Opus 5.5 leads outright on two.
That's a clear lead by count. It's also Google grading its own homework.
The independent numbers arrived within a day, and they tell a tighter story.
Artificial Analysis scored Argon at 53 on its Intelligence Index with reasoning set to high. That matches GPT-6 Astra at 53 and sits one point ahead of GPT-6.1 Sol at 52. It's also 23 points above Google's previous non-Flash model, Gemini 3.1 Pro Preview, which scored 30.
Vals AI ranked Argon first of 41 models on its Vals Index at 68.90%. Second place, Claude Sonnet 5.5, scored 67.04%. Third, Claude Opus 5.5, scored 66.97%.
Here are the numbers side by side, with who measured them:
| Test | What it measures | Gemini 4 Argon | Closest rival | Source |
|---|---|---|---|---|
| Artificial Analysis Intelligence Index | Overall reasoning across many tests | 53 | GPT-6 Astra, 53 | Artificial Analysis |
| Vals Index | Mixed professional tasks | 68.90% | Claude Sonnet 5.5, 67.04% | Vals AI |
| Finance Agent v2 | Multi-step financial research | 65.4% | Claude Opus 5.5, 58.6% | Google, via VentureBeat |
| DeepSWE v1.1 | Long software engineering jobs | 77.9% | Claude Opus 5.5, 74.2% | Google, via VentureBeat |
| LVBench | Understanding long videos | 91.7% | GPT-6 Astra, 87.5% | Google, via VentureBeat |
| AutomationBench-AA | Agentic automation | 78% | Claude Sonnet 5.5, 71% | Artificial Analysis |
Read that table for what it says.
Argon has a big lead in a few places. Finance research by almost seven points. Agentic automation by seven. Long video by four.
On the broad indexes, it's a tie at the top or a lead of about two points.

Two points on an index is not a gap you'll feel writing an email. It's inside the range where your prompt and your files matter more than whose name is on the model.
Think of it like sprinters crossing the line within a stride of each other. The photo picks a winner. Your race, on your track, might not.
The honest read is this. Google is back at the top table, after a year of being treated like it had fallen off. That's real news. It isn't a reason to believe Argon is better at your job until you've watched it do your job.
Where Gemini 4 still loses
The places a new model loses tell you more than the places it wins, because nobody puts those in the headline.
Reuters noted that Argon remained behind on two of the four coding benchmarks Google included in its own release. That's Google's sheet, so these are losses Google chose to publish.
The independent data lines up. On Terminal Bench 4, which tests whether a model can get real work done at a command line, Artificial Analysis scored Argon at 57%. Claude Sonnet 5.5 scored 64%. Claude Opus 5.5 scored 60%. GPT-6 Astra scored 59%.
So Argon came fourth on that one.
That's still a 53 point jump over Gemini 3.1 Pro Preview on the same test, which tells you how far behind Google had been in agentic work. But fourth is fourth.
Here's how to read the pattern.
Argon is strongest when the job is long, document heavy and analytical. Reading a pile of filings and finding the problem. Watching an hour of video and pulling out the moment that matters. Planning a multi-step change across a large codebase.
It's weaker, for now, at the tight loop of running commands, reading the output and trying again. That's the work a coding agent does in a terminal hundreds of times an hour.
If you write code for a living, that split should shape your first test. Don't ask Argon to plan a refactor and call it a win. Plan the refactor with it, then hand it the terminal and watch whether it finishes.
There's one more gap worth knowing about, and it's speed.
Vals reports an average latency of 46 minutes and 33 seconds across its index tests with reasoning set to high. Those are long, hard test suites, not a single chat reply, so don't read it as "every answer takes 46 minutes." Read it as a warning that high reasoning on a big job is slow.
If your work is quick back and forth, a lighter model will feel better in your hands even when Argon would score higher on paper. We've written before about when a reasoning model is worth the wait and when a standard model is the better call. Argon makes that choice sharper, not easier.
Here's the failure mode to avoid.
Someone reads "most powerful model yet," hands Argon a five second task, waits, and decides it's overrated. The model didn't fail. The task was the wrong shape for it.
The price is a teaser rate, so budget for the second number
Google published two prices for Argon, and only one of them is the real one.
Per the Google announcement, Argon launches at an introductory price of $2 per million input tokens and $10 per million output tokens. Cached input tokens are 95% off the input price. A footnote in the same post says that after the introductory period ends, the price becomes $4 per million input tokens and $20 per million output tokens.
The post doesn't say how long the introductory period lasts.
So the honest price is double the launch price, starting on a date nobody has told you.
The launch rate isn't an accident either. Yahoo Finance reported that OpenAI launched GPT-6.1 Sol the day before at exactly the same $2 and $10 rates. Two labs, one day apart, same price. That's a price war, and you're the one it's aimed at.

Here's the math on a real job, so you can see what doubling does.
Say you run a contract review that sends 200,000 tokens in (the contract plus your instructions) and gets 20,000 tokens back (the review). That's a long document but not an unusual one.
| Input cost | Output cost | Per review | 500 reviews a month | |
|---|---|---|---|---|
| Introductory rate | $0.40 | $0.20 | $0.60 | $300 |
| Standard rate | $0.80 | $0.40 | $1.20 | $600 |
Same work. Twice the bill.
Now add caching. If the same 150,000 tokens of instructions and reference material go into every review, and only 50,000 tokens change each time, the cached part costs 95% less. That's the lever that matters at scale, and most people never pull it because they paste a fresh prompt every time.
Do this before you build anything on Argon:
- Write down the token count of one typical job, in and out. Your current provider's usage page shows it.
- Multiply by the standard rate, $4 and $20, never the launch rate.
- Multiply by how many times a month you run it.
- Mark which part of the input repeats every time, and price that part at the cached rate.
If the number at step three makes you flinch, build for it now. The launch price is a coupon, and coupons expire.
One more thing to put in the spreadsheet. Vals measured Argon at $15.68 per test across its index, against $21.34 for Claude Sonnet 5.5 and $32.14 for Claude Opus 5.5. So on the work Vals tested, Argon was both the highest scoring model and the cheaper one to run. That's a real advantage, and it's the one most likely to survive the price change.
Long-horizon work is the pitch, and it's a specific kind of work
Google uses the phrase "long-horizon" all over its announcement. It's easy to skim past. Don't, because it's the best clue to what Argon is for.
A long-horizon task is one where the model has to hold a goal across many steps without losing the thread. Not "summarize this PDF." More like: read these forty documents, figure out which ones contradict each other, check the numbers against this spreadsheet, draft the memo, then update the tracker.
Each step is easy on its own. The hard part is step thirty, when the model has to remember what you asked for in step one.
Two numbers tell you Argon was built for this. Vals lists a 1 million token context window and up to 262,000 output tokens. The context window is how much the model can read at once. The output limit is how much it can write back in one go.
A million tokens is several long books. In practice, it's a full contract data room, a quarter of meeting transcripts, or a large codebase in one pass.

Google also says it uses Argon internally, describing it as "fundamentally changing the way we work and build at Google." That's a claim you can't check from outside, so give it the weight you'd give any company praising its own product.
What you can check is whether your work has this shape.
Here's a quick test. Look at the last week and find the task that took you the longest to hand off. Not the hardest one. The one where explaining it to someone else would have taken nearly as long as doing it yourself.
That's your long-horizon task. It probably involves several files, a few rules that aren't written down anywhere, and a result that has to be right in more than one place at once.
If you found one, Argon is aimed squarely at you. If every task you have is short and self-contained, Argon's biggest advantage won't show up in your work, and a faster, cheaper model will serve you better.
The failure mode here is the context dump. A million tokens of room tempts you to throw everything in and let the model sort it out. Models still get worse at finding the one fact that matters when it's buried in a pile of things that don't. Give Argon the forty documents that matter, not the four hundred you happen to have.
How Gemini 4 compares to GPT-6 and Claude for your actual work
Nobody needs a winner. You need to know which model to try first for the thing in front of you.
Based on the published numbers above, here's where each model looks strongest today. Treat it as a starting order for your own tests, not a verdict:
| Your task | Try first | Why | Watch for |
|---|---|---|---|
| Financial research across many documents | Gemini 4 Argon | Leads Finance Agent v2 by about seven points | Introductory pricing ending |
| Long video review | Gemini 4 Argon | Leads LVBench at 91.7% | Slow responses at high reasoning |
| Agents that read untrusted content | Gemini 4 Argon or Claude Opus 5.5 | Lowest prompt injection rates on Gray Swan | Limited access to Argon for now |
| Command line coding agents | Claude Sonnet 5.5 | Leads Terminal Bench 4 at 64% | Higher cost per task on Vals |
| Large multi-step code changes | Gemini 4 Argon | Leads DeepSWE v1.1 at 77.9% | Weaker at the terminal loop |
| Everyday writing and quick questions | Whatever you already use | The top models are within a couple of points | Paying for power you won't feel |
Notice the last row. It covers most of what most people do with AI on most days.
For that work, the differences between the top models are temperament, not intelligence. One fills in gaps on its own, another does exactly what you said. We broke down that split in our Claude vs ChatGPT comparison, and it holds for Gemini too. Argon doesn't change which temperament suits you. It adds another option at the top.
You might be thinking: if the models are this close, why bother switching at all?
Fair. For a lot of daily work, you shouldn't.
But "close on average" hides "far apart on one task." A seven point lead on financial research is the difference between a draft you trust and a draft you check line by line. If that's your job, the average doesn't matter. Your task does.
That's the whole method. Find the one task where the gap is real for you, and test that.
If you want the wider view of how the big labs differ on features rather than scores, our ChatGPT vs Claude vs Gemini feature comparison covers the parts benchmarks don't measure, like file handling and memory.
Build your Gemini 4 test before you get access
The worst time to figure out how to judge a new model is the day you get it. You'll be excited, you'll try something easy, it'll do fine, and you'll decide it's great.
Then three weeks later it fails on the work you actually needed it for.
Build the test now, while you're waiting. It takes about an hour and you'll use it for every model release after this one.
- Pick five real tasks from the last month. Real ones, with the actual files. Not "write a poem about the ocean."
- Include one long-horizon task, the kind from the section above. That's where Argon should shine if Google's claims hold.
- Include one task where your current model failed or needed heavy fixing. That's where a new model has to prove itself.
- Include one fast, simple task you do every day. That's your check on speed and cost.
- Write down what a good answer looks like for each, before you run anything. One or two sentences. This stops you grading on vibes.
- Run all five through your current model today and save the outputs.
- When Argon opens to you, run the same five, with the same prompts and files, and compare side by side.
Step five is the one people skip, and it's the one that keeps the test honest.
Once you've seen an answer, you grade against it. Writing down what "good" means first is how you catch the model that sounds confident and gets the number wrong.
Score each task on whether it was right, how much you had to fix, and how long it took. Leave out "did I like it." Liking it is how a model with a nice tone beats a model with the correct answer.
Here's the payoff. When the next flagship drops, and at this pace one will, you already have the test. You run it in an afternoon and you know.
Everyone else will be reading launch posts and guessing.
Don't marry a model, keep your work portable
Flagships keep landing within days of each other. GPT-6.1 Sol on September 29, Gemini 4 Argon on September 30, and a fresh round of Claude releases in the same stretch.
The model at the top in October may not be the one at the top in January. That's the pattern of the last three years, and nothing about this week suggests it's slowing down.
So the expensive mistake isn't picking the wrong model. It's building your work so you can't leave.
That happens slowly. Your best prompts live in one app's history. Your custom instructions are set up in one product. Your files are uploaded in one place. Every month, switching costs a little more, until the better model comes out and you don't move because moving means rebuilding.

The fix is to keep the things that are yours, your prompts, your instructions, your files and your conversation history, separate from the model doing the work this month.
This is the problem we built Magai to solve, so weigh what I say next knowing I'm the founder.
In Magai, Gemini, ChatGPT, Claude, Grok and dozens of other models sit in one workspace. You can switch models in the middle of a conversation and the thread's context comes with you. The model changes. The memory of what you were doing doesn't reset.
That's what makes the test from the last section cheap to run. Same chat, same files, same instructions, a different model each time. You're comparing models instead of comparing setups.
Argon is in limited release, so I won't promise you a date for it anywhere. What I can tell you is that the habit of keeping your work portable pays off every time a release like this lands, whichever lab ships it. If you want the longer version of why that setup matters, here's what an AI aggregator is and what to check before you trust one.
You can see the plans on our pricing page. Every plan carries a 30 day, 100% money back guarantee.
What Gemini 4 means for the next few months
Step back from the scores for a second and look at what changed.
A year ago, the question was whether Google could keep up. As of this week, Google has a model that ties or leads the best from OpenAI and Anthropic on the broad independent indexes, leads on several specialist tests, and has one of the lowest prompt injection rates anyone has published.
That changes the market, even before you can use it.
It puts pressure on price. Argon and GPT-6.1 Sol launched at the same rate a day apart. When two labs match each other that precisely, the others tend to follow, and the people paying per token win.
It puts pressure on safety claims. Google led with security partners and a published attack success rate. Expect rivals to publish more of their own numbers on the same tests, which is good for anyone choosing a model for agent work.
And it puts pressure on loyalty. If the top labs are within a couple of points of each other on average, the reason to stay with one is habit, not capability. Habits are easy to break when switching is cheap and expensive when it isn't.
Here's what to watch for over the next few weeks:
- The date Google opens Argon to paying Gemini subscribers, since Axios reports they're first in line.
- The date the introductory price ends, which turns every cost estimate you made into the real one.
- Independent coding results from people using it outside Google, especially on terminal work, where it's behind today.
- Whether Google ships smaller Gemini 4 models, since the Flash line is what most people actually run every day.
You don't need to act on any of these today. You need to notice them when they happen.
For more on how the major models stack up as this race keeps moving, the rest of our artificial intelligence coverage tracks each release as it lands.
Gemini 4 is the best model most of you can't use yet.
Build your test. Keep your work portable. Then let the door open on your schedule, not Google's.
More in Artificial Intelligence

Artificial Intelligence
All AI Tools in One Website: What to Check First
Not every "all AI tools in one website" claim means the same thing. This is the real test, plus what running separate AI tools actually costs.

Artificial Intelligence
Claude vs ChatGPT: Which One Should You Pay For?
Claude and ChatGPT fail in opposite directions by design. Here's how to pick the right one for each task, without paying twice for tools that don't talk.

Artificial Intelligence
What Is an AI Aggregator? A Straight Answer
An AI aggregator puts GPT, Claude, and Gemini behind one login. What actually matters is whether context survives when you switch models mid-thread.