OpenAI’s GPT-6 Astra scores 99% on ARC-AGI-3, surpassing humans on 96% of levels
OpenAI’s GPT-6 Astra achieved state-of-the-art performance on the ARC-AGI-3 benchmark, scoring 99% with a provider-adapter harness and surpassing human performance on 96% of levels. Under a standard harness, it scored 63%. The model demonstrated efficient symbolic world modeling, developing its own shorthand DSL. The benchmark is now nearing saturation, marking a major milestone in AI reasoning.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page itself is projected from evidence records.
- Current automated evidence projection
Cross-source coverage
Wire timeline
GPT-6 Astra scores 99% on ARC-AGI-3 with provider-adapter harness, beating humans on 96% of levels
A post on X reports that OpenAI's GPT-6 Astra has achieved 99% accuracy on the ARC-AGI-3 benchmark using a new 'Provider Adapter harness,' surpassing human performance on 96% of the benchmark's levels. The ARC-AGI-3 benchmark was designed to expose failures in novel reasoning and is now described as nearly saturated for the strongest system. The post highlights a critical nuance: the same model scores only 62.7% under ARC's standard, provider-neutral harness, which provides a minimal common interface. The dramatic performance jump to 99.9% occurs when using OpenAI's native context-management and reasoning-state preservation features. This distinction, according to ARC, answers two different questions: how models compare under a neutral setup versus how well a model performs with its own provider's context features.
OpenAI's GPT-6 Astra Achieves State-of-the-Art on ARC-AGI-3, Benchmark Nearing Saturation
Greg Brockman, co-founder of OpenAI, reposted an evaluation from ARC Prize showing that OpenAI's GPT-6 Astra has achieved state-of-the-art (SOTA) performance on the ARC-AGI-3 benchmark, which he claims is now approaching saturation. Astra scored 63% on the standard evaluation harness and 99% on the new Provider Adapter harness, surpassing human performance on 96% of ARC-AGI-3 levels. The leaderboard data also indicates that higher reasoning levels generally cost less, as Astra completes levels with fewer actions, reducing model calls and token usage. The significant score gap between the two harnesses highlights the impact of the evaluation methodology on scoring. This development marks a major milestone in AI reasoning capabilities, though the benchmark's saturation suggests the need for new, more challenging tests.
OpenAI's GPT-6 Astra scores 63% on ARC-AGI-3, surpassing human baseline in action efficiency
A post from testingcatalog on X reports that OpenAI's GPT-6 Astra achieved a 63% score on the ARC-AGI-3 benchmark using a standard harness. The model surpassed the human baseline in action efficiency, using fewer actions than the median tested human on 96% of levels. A key observed behavior was GPT-6 Astra's ability to convert unfamiliar environments into compact symbolic world models, representing game mechanics as logical rules and developing its own domain-specific language shorthand to track state and plan actions. This result highlights a significant advancement in AI reasoning and efficiency on a challenging general intelligence benchmark.
Show 2 older updatesHide older updates
GPT-6 Astra scores 66% on ARC-AGI-3, nearly 100% with custom harness at $360 per game
François Chollet announced GPT-6 Astra, a new AI model that achieves a step-function improvement in interactive reasoning. On the ARC-AGI-3 benchmark, it scores 66% using a standard harness and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. The continuous harness version outperforms the human baseline in action efficiency across almost all levels. Analysis of the model's reasoning chains reveals it performs efficient, on-the-fly symbolic world modeling for each game, even developing its own shorthand DSL to represent in-game situations. Chollet notes that Astra exhibits symbolic modeling behaviors previously only seen with sophisticated harnesses, indicating that harness capabilities are increasingly shifting into the model itself. He describes Astra as a major breakthrough in model intelligence.
OpenAI's GPT-6 Astra scores 98.6% on ARC-AGI-3 and tops multiple benchmarks
A post from testingcatalog on X reports that OpenAI's GPT-6 Astra achieved a score of 98.6% on the ARC-AGI-3 benchmark, topping most other benchmarks. Additional results include a 74% score on DeepSWE 1.1, and significant improvements on FrontierMath Tier 4 and Terminal-Bench Science 0.1. The post includes a link to further details. These results indicate substantial progress in AI reasoning and problem-solving capabilities.