Google Gemini 4 Argon matches GPT-6 Astra on intelligence index, leads 13 of 19 benchmarks
Google DeepMind released Gemini 4 Argon, its first proprietary model above the Flash class in over seven months. The model scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra, and leads or ties on 13 of 19 benchmarks against GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1. It achieved a 77.9% score on DeepSWE v1.1 and a 15% hallucination rate, the lowest among leading models. Access is limited to selected users, paid API customers, and Google AI Ultra subscribers.
Reference imageEditorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Reporting timeline
Gemini 4 Argon matches GPT-6 Astra on Artificial Analysis Intelligence Index with 53 points
A post on X by user kimmonismus reports that Google's Gemini 4 Argon model has achieved a score of 53 points on the Artificial Analysis Intelligence (AAI) Index, matching the score of OpenAI's GPT-6 Astra and surpassing GPT-6.1 Sol's 52 points. The post claims Gemini 4 Argon also takes the number one spot on the AutomationBench-AA benchmark. Additionally, on the AA-Omniscience evaluation, the author states that Argon hallucinates significantly less at 15% compared to Astra's 51%, and is more willing to admit uncertainty, though its answer accuracy is lower at 50% versus Astra's 63%. The author concludes that this appears to be a very good release and extends congratulations to the Gemini team.
Read sourceGoogle's Gemini 4 Argon outscores GPT-6 Astra and Claude Opus 5.5 on most benchmarks
A post from analyst rohanpaul_ai reports a massive reveal from Google: its new flagship AI model, Gemini 4 Argon, outscores OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5 on most industry benchmarks. The model's widest lead is in legal work, scoring 19.6% on Harvey's Legal Agent Benchmark compared to 6.7% for Anthropic's Claude Fable 5.1. The output limit has been increased from 64K to 1M tokens, described as an industry-leading ceiling. Access is currently limited to three groups: Google's own staff, vetted cyber defenders including government agencies and security companies, and trusted testers providing feedback. Internally, Google reports that Argon agents freed over 300 TiB of data-center memory, with estimated total savings of 500 TiB to 1 PiB, and made a Rust port of the libgav1 video decoder 2.7x faster by replacing 32,000 lines of SIMD code.
Read sourceGoogle DeepMind's Gemini 4 Argon matches GPT-6 Astra on intelligence index at 60% cost
Google DeepMind has released Gemini 4 Argon, its first proprietary model above the Flash class in over seven months. According to the Artificial Analysis Intelligence Index, Gemini 4 Argon (high reasoning) scores 53, matching GPT-6 Astra (max) and surpassing GPT-6.1 Sol (max) by one point. The model is currently available at a 50% promotional discount, costing $1.99 per Intelligence Index task, which is 60% of GPT-6 Astra's cost. After the promotion ends, the price will rise to $3.98. Gemini 4 Argon demonstrates strong agentic performance, ranking first on AutomationBench-AA at 77.5%, and achieves the lowest hallucination rate (15%) among leading models. However, its accuracy is slightly lower than GPT-6 Astra. The model supports a 1M token context window, multimodal input (text, image, video, speech), and text output. It is currently being rolled out to selected users and is not publicly available. Google has not confirmed the promotion end date.
Read sourceShow 3 older updatesHide older updates
Google releases Gemini 4 Argon, topping 14 of 19 benchmarks against GPT-6 and Claude models
Google has released Gemini 4 Argon, a new AI model that achieves top or tied scores on 14 of 19 benchmarks in Google's internal comparison table. The model was tested against GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1. Specific results include a 77.9% score on DeepSWE v1.1 (versus Opus 5.5 at 74.2% and Astra at 74.1%), 51.3% on AutomationBench (versus 42.5%), 84.2% on long context tasks at 256K–1M tokens (versus 71.8%), and 91.7% on LVBench (versus 87.5%). The source, DanDr1s, notes that while the benchmarks appear impressive, real-world outputs remain to be evaluated, expressing caution that the model may be optimized for benchmarks rather than practical performance.
Read sourceGoogle releases stable Gemini 4 Argon model, leading 13 of 19 benchmarks
According to a post by analyst ChrisGPT, Google has released a stable pro model, Gemini 4 Argon, after over a year and a half. The model leads outright on 13 of 19 benchmark results shown in a chart and ties with Astra on cybersecurity. It achieves 84.2% on GraphWalks at 256K-1M tokens for long-context reasoning, compared to 71.8% for Astra and 66.8% for Opus 5.5. On Harvey's legal benchmark, Argon scores 19.6%, nearly three times the next-best model's 6.7%. It also leads all four knowledge-work evaluations, including 51.3% on AutomationBench against 42.5% for Opus and 41.4% for Astra. Astra still leads on FrontierSWE and Terminal-Bench Science, while Opus 5.5 leads Terminal-Bench 4.0 and PostTrainBench. The post argues Argon's strongest case is its breadth of results across knowledge work and long-context reasoning.
Google announces Gemini 4 Argon model with 77.9% DeepSWE score, new SOTA
Google has announced Gemini 4 Argon, a new frontier model designed for complex workflows in software engineering, enterprise knowledge work such as legal and finance, and cyber defense. The model achieved a 77.9% score on the DeepSWE v1.1 benchmark, establishing a new state-of-the-art (SOTA) result. According to the announcement, Gemini 4 Argon performs better than GPT-6 Astra, Opus 5.5, and Fable 5.1 across many benchmarks. The model will be rolling out soon, starting with paid API customers and Google AI Ultra subscribers.
Read source