Google Unveils Gemini 4 Argon With 1M-Token Output Limit
-
- by THEFLGHT,
- October 01, 2026
- in Artificial-Intelligence
Gemini 4 Argon is Google’s new frontier model, launching first with a small group of trusted cybersecurity defenders and through a voluntary U.S. government pre-release process. Google has not announced a general-public release date.
The September 30 announcement pairs a sharply larger output window with stronger coding, research and cyber capabilities. Google says Argon can generate as many as one million output tokens in a single response, compared with 64,000 for its previous flagship generation.
The initial Gemini 4 Argon release has four defining limits and capabilities:
- Access begins with selected cyber defenders.
- General availability has no announced date.
- Maximum output expands to one million tokens.
- Introductory API pricing starts at $2 per million input tokens.
Gemini 4 Argon Starts With Trusted Cyber Defenders
Google is distributing the model through its Fairwind Program, which gives vetted defenders access to advanced cyber capabilities. The company says participants will receive Argon without cyber guardrails that could otherwise restrict defensive investigations, while its broader safeguards against harmful requests remain in place.
Wiz is among the partners using Argon through its Scan for Good initiative, according to Google’s official announcement. The program scans open-source projects for critical vulnerabilities and helps maintainers validate and patch flaws before attackers can exploit them.
The model is also entering the U.S. government’s voluntary pre-release access process. Reuters reported that Argon anchors Google’s Gemini 4 generation and is larger than the company’s previous top-tier Pro models, citing a Google spokesperson.
This is therefore a controlled launch rather than a conventional product release. Developers cannot yet assume that Argon is available in Google AI Studio or through the public Gemini API, and consumers do not have an announced date for using it in the Gemini app.
One Million Output Tokens Target Long-Horizon Work
The one-million-token output ceiling is one of Argon’s most unusual specifications. Input context describes how much information a model can read; output context determines how much it can produce before a response ends. Google says the larger limit supports long-running software migrations, research and other workflows that produce extensive intermediate work.
Google reports that internal teams used Argon to migrate large C and C++ codebases to Rust, including projects exceeding 800,000 lines. In another test, the model produced a 32,000-line Rust SIMD implementation for the libgav1 video decoder that ran 2.7 times faster than an earlier Rust port while preserving identical video output.
The company also describes infrastructure work beyond code generation. An Argon-generated optimization for a quantum-computing subroutine reportedly beat a published baseline by 40%, while a memory-efficiency project freed about 300 tebibytes of capacity and identified between 500 tebibytes and one pebibyte of potential savings.
Those results are company-reported case studies, not independent proof that every organization will receive similar gains. Large migrations still require tests, code review and security checks, while infrastructure savings depend on the architecture, workload and quality of the model’s access to internal tools.
Google’s Benchmarks Show Argon’s Strengths and Gaps
Google’s published evaluations emphasize software engineering and autonomous computer work. Argon scored 77.9% on DeepSWE v1.1, 68.9% on Vals Index and 51.3% on AutomationBench. The company also reports 91.7% on LVBench for long-video understanding and 68% on CWE-bench v1 for vulnerability repair.
The benchmark table does not show universal leadership. Competing models remained ahead on several tests, including some terminal, operating-system and frontier software-engineering evaluations. That mixed picture matters because model selection increasingly depends on the exact task rather than a single composite score.
Cyber performance also creates a deployment tension. A model that can find, validate and patch serious vulnerabilities may help defenders reduce exposure faster, but the same technical competence can be misused. Google says it tested Argon against prompt injection, harmful cyber requests and chemical or biological misuse before this limited release.
The company says tool actions can be watched through chain-of-thought and action monitoring, with execution stopped when a system detects dangerous behavior. Argon is also designed to operate inside hardened, sealed sandboxes when it runs code, limiting the damage that a compromised prompt or faulty action could cause.
Pricing and Safety Shape the Wider Rollout
Google plans an introductory API price of $2 per million input tokens and $10 per million output tokens. After the introductory period, those rates are scheduled to double to $4 and $20 respectively. Cached input will receive a 95% discount, which could materially reduce costs for repeated context.
The company says paid API customers and Google AI Ultra subscribers will be first in line when access expands. It has not stated when that expansion begins, how long introductory pricing will last or whether the full one-million-token output limit will be available in every product surface.
Price comparisons also need caution. A long response can create substantial output charges even at the introductory rate, and agentic tasks can consume additional compute through retries, tools and verification. The practical cost will depend on how efficiently developers constrain the model and validate its work.
Argon arrives as Google tries to convert frontier-model research into more durable software and security workflows. Its controlled release gives the company time to study high-capability cyber use before broad distribution, while early partners test whether the model’s long-horizon performance survives production conditions.
The next concrete milestones are wider API access, consumer availability and independent testing at scale. Until those arrive, Gemini 4 Argon should be understood as a significant new model with unusually large output capacity, but one whose strongest claims and deployment economics still need external validation.
Background Reading
0 Comments:
Leave a Reply