Wire flash
Analyst: Google releases stable Gemini 4 Argon model, leads 13 of 19 benchmarks
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
According to a post by analyst ChrisGPT, Google has released a stable pro model, Gemini 4 Argon, after over a year and a half. The model leads outright on 13 of 19 benchmark results shown in a chart and ties with Astra on cybersecurity. It achieves 84.2% on GraphWalks at 256K-1M tokens for long-context reasoning, compared to 71.8% for Astra and 66.8% for Opus 5.5. On Harvey's legal benchmark, Argon scores 19.6%, nearly three times the next-best model's 6.7%. It also leads all four knowledge-work evaluations, including 51.3% on AutomationBench against 42.5% for Opus and 41.4% for Astra. Astra still leads on FrontierSWE and Terminal-Bench Science, while Opus 5.5 leads Terminal-Bench 4.0 and PostTrainBench. The post argues Argon's strongest case is its breadth of results across knowledge work and long-context reasoning.
Source report
After more than a year and a half of anticipation, Google has finally released a stable pro model: Gemini 4 Argon.
Benchmark Performance
Gemini 4 Argon leads outright on 13 of 19 results in the reported chart and ties Astra on cybersecurity.
Long-Context Reasoning
The model demonstrates strong long-context performance, though it still trails Muse 1.3 (which was not reported in the same chart). Key results on GraphWalks at 256K–1M tokens:
- Gemini 4 Argon: 84.2%
- Astra: 71.8%
- Opus 5.5: 66.8%
Legal and Knowledge-Work Benchmarks
On Harvey's legal benchmark, Gemini 4 Argon scores 19.6% — nearly 3x the next-best model's score of 6.7%.
It also leads all four knowledge-work evaluations in the table, including:
- AutomationBench: 51.3% (vs. 42.5% for Opus and 41.4% for Astra)
Areas Where Competitors Still Lead
- Astra leads on FrontierSWE and Terminal-Bench Science
- Opus 5.5 leads on Terminal-Bench 4.0 and PostTrainBench
Key Takeaway
The strongest argument for Gemini 4 Argon is the breadth of its results, particularly across knowledge work and long-context reasoning.
Source
ChrisGPTNeutral / independent
Part of this Story
Google Gemini 4 Argon matches GPT-6 Astra on intelligence index, leads 13 of 19 benchmarks