Drooid Logo
Back to story perspectives

Full Breakdown

Google Unveils Gemini 4 Argon, Its New Frontier AI Model

By Drooid · · How we work

Core Announcement: Gemini 4 Argon Details

On September 30, Alphabet’s Google announced Gemini 4 Argon, the latest flagship model in its Gemini 4 series. The model supports up to 1 million output tokens (roughly 750 000 words) and is marketed for complex enterprise tasks such as software engineering, legal reasoning, finance, and cybersecurity. Google is rolling the model out in phases: an initial release to a select group of cybersecurity defenders through its Fairwind program, participation in the U.S. government’s voluntary pre-release safety-testing process, and a later expansion to developers, enterprise customers, and paid-API users. Introductory pricing is $2 per million input tokens and $10 per million output tokens, with rates potentially doubling later.

Background & Context

Google’s AI roadmap saw Gemini 3 launched late 2025, followed by a cancelled Gemini 3.5 Pro. During that period, rivals OpenAI and Anthropic released frontier models (e.g., OpenAI’s Astra and GPT-6.1 Sol; Anthropic’s Opus and Fable 5). Leadership changes at DeepMind—founder Demis Hassabis stepped aside and Koray Kavukcuoglu assumed the head role—have been framed as a response to the accelerating competition.

Data & Statistics

  • Benchmark performance: Artificial Analysis’s Intelligence Index ranks Gemini 4 behind only Claude Opus 5.5 and Claude Sonnet 5.5; the model ties OpenAI’s Astra on a key cybersecurity benchmark and leads on software-engineering tests.
  • Hallucination rate: 15 % hallucination rate, the lowest among leading models, versus 54 % for Astra and GPT-6.1 Sol.
  • Memory optimization: Internal use has freed more than 300 TiB of data-center memory, with projected savings between 500 TiB and 1 PiB.
  • Specialized scores: 77.9 % on DeepSWE v1.1, 51.3 % on AutomationBench, and 91.7 % on LVBench.
  • Token limits: Output capacity increased from the previous 64 000-token ceiling to 1 million tokens; OpenAI’s Astra caps at 128 000 tokens.

Official Statements & Responses

  • Tim Law, director of research for AI at IDC, noted that Gemini 4 Argon shows advanced reasoning on critical tasks such as legal and financial knowledge work.
  • Google has rebutted reports of internal doubt, stating that employees have been testing versions of Gemini 4 for weeks, with some receiving unlimited access.

Criticism & Opposition

  • Bloomberg reported that unnamed employees claim the model’s strong benchmark scores do not translate to real-world coding performance, citing uneven results on front-end design tasks and high operational costs.
  • Edwin Chen, founder of AI startup Surge AI, warned that the industry’s focus on “benchmaxxing” can produce models that excel on test suites but fall short in practical applications.
  • Internal sources highlighted concerns that the model’s size could pressure Google’s margins if deployed at scale in continuous-use agents.

Conflicting Reports & Gaps

  • Performance vs. Benchmarks: Google emphasizes leading scores, while Bloomberg-sourced insiders describe coding struggles in actual use. Independent testing outside the limited partner program has not yet occurred.
  • Pricing Outlook: Introductory rates are set, but future price increases lack a timeline.
  • Release Schedule: The model is unavailable to the general public, and Google has not disclosed a definitive public launch date.

What’s Next

Google will continue its participation in the U.S. voluntary safety-testing process while expanding access from cybersecurity partners to paid-API customers and Google AI Ultra subscribers. Internal evaluations will focus on mitigating prompt-injection attacks, monitoring for cyber-misuse, and refining cost efficiency before a full public rollout.