Google debuts Gemini 4 Argon, first for trusted cyber defenders

Google announced Gemini 4 Argon, its new frontier model, which is rolling out first to a set of trusted cyber defenders through the company's Fairwind Program. Google describes it as built to sustain deep reasoning across complex, long-horizon workflows, with frontier performance in real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.
Access will widen in phases. Google says it is actively engaged in the U.S. government's voluntary process for pre-release model access while it gradually expands access, and that it will keep gathering feedback from early testers and iterating on guardrails before making Argon available to developers, enterprises and consumers "as soon as possible". No date is given. The broader release is to start with paid API customers and Google AI Ultra subscribers. When it launches, Argon will carry an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off the input price. The output token limit rises to 1M tokens, which Google calls industry-leading, up from 64K.
Google says Argon is already powering its internal workflows, with thousands of Googlers praising its specialized coding, deeper research and writing quality. Three internal examples are given. In quantum algorithm optimization, Argon helped researchers optimize the spacetime resources (qubits x gates) of subroutines that bottleneck important applications; in one example it beat the published baseline by 40% in minutes. In memory efficiency, a team of Argon agents analyzed fleet-wide profiling telemetry and autonomously identified and applied optimizations across Google's data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings. In codebase migration, Argon agents are moving C/C++ code to Rust, from tens of thousands of lines in core libraries such as re2 and libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Because many of these systems are critical, Google says the rewrites are undergoing automated and manual auditing, emulation testing and review before production. For libgav1, Google's open source video decoder, agents took an existing Rust port and replaced 32K lines of SIMD code through many rounds of profile-guided experiments, producing safe Rust the compiler can vectorize automatically. The result is a memory-safe decoder that runs 2.7x faster than the Rust port, with identical video output.
On benchmarks, Google reports a new state of the art on DeepSWE v1.1 (77.9%), which measures real-world long-horizon software engineering. Argon ranks #1 on Zapier's AutomationBench with 51.3% and is state of the art on LVBench, a long video understanding test, with 91.7%. Google also says Argon leads the Vals Index (economic impact across finance, coding, legal and tax work) and gives similarly leading results on Vals Finance Agent v2 and Harvey's Legal Agent Benchmark, though no scores are listed for those. On CWE-bench v1, which tests remediation of security vulnerabilities, Argon ties for first place with a top score of 68%.
Cybersecurity is the headline use. Google says it trained Argon to be highly capable at defense and that it can autonomously find, validate and patch critical software vulnerabilities. Trusted defenders and Google's internal teams will get Argon without cyber guardrails. Wiz is already using it through its Scan for Good initiative, a free program protecting critical public infrastructure; in an early demonstration the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, a risk Google says previous frontier models had missed. Google also says Argon beats 3.8 Flash Cyber on its internal vulnerability benchmark (codebases spanning 20 programming languages) and on Wiz's internal black-box penetration testing benchmark, without giving numbers.
Google is strengthening safeguards in four areas before broad release. Against misuse, the model is designed to refuse harmful cyber and CBRN requests while preserving legitimate dual-use science, under the Frontier Safety Framework, with improved monitoring of internal activations and testing by internal and external red teams. Against prompt injection, Argon is called Google's most resilient model yet and is said to lead Gray Swan's Indirect Prompt Injection benchmark. For misalignment, Google is deploying mitigations that monitor Argon's chain-of-thought and actions and stop execution when necessary; a similar system watched training runs and alerted an incident response team, and Google took care not to feed those findings back into training so as not to shape reasoning to evade monitoring. It urges the rest of the industry to preserve reasoning transparency. For hardening, sandboxed environments are being isolated and sealed before high-risk training or evaluations begin, in line with Google's agent control roadmap.
Key facts
- Gemini 4 Argon is rolling out first to trusted cyber defenders via Google's Fairwind Program; developers, enterprises and consumers get it later, starting with paid API customers and Google AI Ultra subscribers, with no date given.
- Introductory price at launch: $2 per million input tokens and $10 per million output tokens, cached input 95% off; output limit rises from 64K to 1M tokens.
- Google reports 77.9% on DeepSWE v1.1 (new state of the art), 51.3% on AutomationBench (#1), 91.7% on LVBench (state of the art) and a 68% first-place tie on CWE-bench v1.
- Internal examples: a 40% beat of a published quantum-subroutine baseline in one case, over 300 TiB of data center memory freed once rolled out, and a libgav1 decoder that runs 2.7x faster than the earlier Rust port.
- Four safeguard areas are being strengthened: misuse, prompt injection, misalignment monitoring of chain-of-thought and actions, and hardened sandboxes.
Why it matters
This is Google's next frontier model, and its first release is aimed at defenders rather than the general public. Google says Argon can autonomously find, validate and patch critical software vulnerabilities, and that trusted defenders will get it without cyber guardrails. The post also shows the model doing long agentic work inside Google, from fleet-wide memory tuning to C/C++ to Rust migration, and pairs that with a 1M-token output limit, up from 64K, to support long reasoning trajectories. Google's call for the industry to preserve reasoning transparency is a notable position on how such models should be monitored.
Who it affects
First, the trusted cyber defenders in the Fairwind Program; Wiz is the only named participant, using Argon through its Scan for Good initiative. Later, developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers. Google's own engineers already use it daily. Teams doing software engineering, legal and finance work, or long-video analysis are the areas where Google reports leading results.
How to use it
There is no general access yet. For now Argon goes to a set of trusted cyber defenders through the Fairwind Program. Google plans to open it to developers, enterprises and consumers "as soon as possible", beginning with paid API customers and Google AI Ultra subscribers. At launch the introductory price is $2 per million input tokens and $10 per million output tokens, with cached input tokens at 95% off the input price. The output limit is 1M tokens.
How solid is it
This is Google's own announcement, so every result is self-reported. Several figures are specific: 77.9% on DeepSWE v1.1, 51.3% on AutomationBench, 91.7% on LVBench and 68% on CWE-bench v1, where Argon ties for first rather than leading outright. The post does not give competing models' scores on those benchmarks, does not name who Argon ties with on CWE-bench v1, and gives no scores for the Vals Index, Vals Finance Agent v2 or Harvey's Legal Agent Benchmark. The vulnerability claims on Google's internal and Wiz's internal benchmarks carry no numbers. The 40% quantum result is a single example against a published baseline.
Risks and caveats
The memory savings are stated as what will be freed once rolled out, not as already achieved. The large Rust rewrites, including the Fuchsia Zircon kernel, are still being audited, tested and reviewed before production. The introductory price has no stated end date, and no general-availability date is given. Releasing a model without cyber guardrails to trusted defenders and internal teams is a deliberate choice Google pairs with phased access. Google itself says prompt injection requires constant vigilance and multiple layers of defense, and that it is still strengthening safeguards across four areas before broad release.
“Argon can autonomously find, validate, and patch critical software vulnerabilities.”
— Google, Gemini 4 Argon announcement