Wire flash
TechGPT-6 Astra achieves Perplexity's strongest WANDR result, 13.5% above Fable 5.1
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
A post from rohanpaul_ai reports that GPT-6 Astra has delivered Perplexity's strongest result yet on the WANDR benchmark. The model scored 13.5% above Fable 5.1 while costing 6.1% less per task. WANDR is a benchmark that tests wide-and-deep research capability, requiring models to find large sets of qualifying entities and back every requested fact with checkable evidence. It contains 500 public tasks requiring 170,495 source-backed records, penalizing incomplete research even when found facts are correct. Scoring tracks both precision and completion, with stricter hard scores requiring entire requested branches to be correct for full credit. The 0.682 result indicates a substantial gain on long, evidence-heavy research work where an agent must continuously find, check, and organize information at scale.
Source report
GPT-6 Astra has delivered Perplexity's best performance yet on the WANDR benchmark, scoring 13.5% above Fable 5.1 while costing 6.1% less per task.
About the WANDR Benchmark
WANDR is an unusual benchmark designed to test models' wide-and-deep research capabilities. It requires models to:
- Find large sets of qualifying entities
- Back every requested fact with checkable evidence
The benchmark consists of 500 public tasks that require 170,495 source-backed records. Incomplete research is directly penalized, even when the facts an agent did find are correct.
Scoring Methodology
WANDR scoring tracks both precision and completion. Stricter hard scores require an entire requested branch to be correct before receiving full credit.
Key Result
The 0.682 score indicates a substantial gain on long, evidence-heavy research work, where an agent must continuously find, check, and organize information at scale.
Source
rohanpaul_aiNeutral / independent
Part of this Story
GPT-6 Astra tops multiple benchmarks at lower cost but faces price hike and mixed results