GPT-6 Astra tops multiple benchmarks at lower cost but faces price hike and mixed results
OpenAI's GPT-6 Astra achieves top scores on Terminal Bench 4.0 at half the cost of the second-place model and scores 0.682 on WANDR, 13.5% above Fable 5.1 at 6.1% lower cost. It matches Fable 5 in coding at less than half the cost and reduces hallucination rate from 92% to 51%. However, pricing is set at $10/$50 per million tokens, a 2.5x increase over GPT-5.6 Sol, making it 75% more expensive per task, with regressions on SciCode and AA-LCR.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page itself is projected from evidence records.
- Current automated evidence projection
Cross-source coverage
Wire timeline
GPT-6 Astra tops Terminal Bench 4.0 at half the cost of second-place model
A post on X by user thsottiaux announces that GPT-6 Astra has achieved the number one ranking on Terminal Bench 4.0 using the Codex harness. The post highlights that GPT-6 Astra accomplishes this at 50% of the cost of the second-place model. Terminal Bench 4.0 is a benchmark for evaluating AI model performance, and the Codex harness refers to a specific testing framework. The claim positions GPT-6 Astra as both the top performer and a cost-efficient option in the competitive AI landscape. The post includes a link to further details, though the full context of the benchmark and the identity of the second-place model are not provided in the post itself. This development underscores ongoing advancements in AI model efficiency and performance.
GPT-6 Astra achieves Perplexity's strongest WANDR result, 13.5% above Fable 5.1
A post from rohanpaul_ai reports that GPT-6 Astra has delivered Perplexity's strongest result yet on the WANDR benchmark. The model scored 13.5% above Fable 5.1 while costing 6.1% less per task. WANDR is a benchmark that tests wide-and-deep research capability, requiring models to find large sets of qualifying entities and back every requested fact with checkable evidence. It contains 500 public tasks requiring 170,495 source-backed records, penalizing incomplete research even when found facts are correct. Scoring tracks both precision and completion, with stricter hard scores requiring entire requested branches to be correct for full credit. The 0.682 result indicates a substantial gain on long, evidence-heavy research work where an agent must continuously find, check, and organize information at scale.
GPT-6 Astra scores 0.682 on WANDR benchmark at $11.98 per task, highest of tested models
An evaluation of GPT-6 Astra on the WANDR benchmark has been published, reporting a score of 0.682 at a cost of $11.98 per task, the highest score among all models tested. The post compares GPT-6 Astra's performance to two other models: it scored 13.5% higher than Fable 5.1 at a 6.1% lower cost, and 27.0% higher than Opus 5 at a 3.3% higher cost. The source is an X post from an account identified as perplexity_ai, suggesting the evaluation was conducted by Perplexity AI. The results highlight GPT-6 Astra's superior performance and cost efficiency relative to competing models on this specific benchmark.
Show 2 older updatesHide older updates
GPT-6 Astra achieves new Pareto frontier with 10% fewer output tokens than GPT-5.6 Sol
A post on X announces that GPT-6 Astra defines a new Pareto frontier for the Intelligence Index versus Output Tokens per Task. The model achieves approximately a 10% reduction in output tokens at maximum effort compared to its predecessor, GPT-5.6 Sol. This indicates a significant improvement in efficiency, delivering higher intelligence per token used. The post includes a link to further details, suggesting a formal announcement or benchmark results. The comparison highlights a direct performance gain in the ongoing development of large language models, with GPT-6 Astra setting a new standard for the trade-off between computational output and intelligence capability.
GPT-6 Astra matches Fable 5 in coding at lower cost, but price hike offsets token gains
Artificial Analysis released an evaluation of OpenAI's GPT-6 Astra, showing significant gains in the Coding Agent Index where it scores equal to Fable 5 at less than half the cost, driven by substantial token efficiency improvements. In the Intelligence Index, GPT-6 Astra scores 61, equal to GPT-5.6 Sol, using approximately 10% fewer output tokens but at a 2.5x price increase, making it 75% more expensive per task. The model shows a dramatic reduction in hallucination rate from 92% to 51% at max effort, with a 4-point accuracy increase. It also gained 80 points in the AA-Briefcase long-horizon knowledge work evaluation. However, it saw mixed results on other benchmarks, including a drop of 80 Elo points in GDPval-AA v2 and regressions on SciCode and AA-LCR. Pricing is set at $10/$50 per million input/output tokens, with a 90% discount for cache reads and 25% premium for cache writes.