Wire flash
TechOpenAI's GPT-6 Astra scores 63% on ARC-AGI-3, surpassing human baseline in action efficiency
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
A post from testingcatalog on X reports that OpenAI's GPT-6 Astra achieved a 63% score on the ARC-AGI-3 benchmark using a standard harness. The model surpassed the human baseline in action efficiency, using fewer actions than the median tested human on 96% of levels. A key observed behavior was GPT-6 Astra's ability to convert unfamiliar environments into compact symbolic world models, representing game mechanics as logical rules and developing its own domain-specific language shorthand to track state and plan actions. This result highlights a significant advancement in AI reasoning and efficiency on a challenging general intelligence benchmark.
Source report
OpenAI's GPT-6 Astra has scored 63% on the ARC-AGI-3 benchmark using a standard harness.
Key Highlights
- Surpasses human baseline: GPT-6 Astra outperforms the median human baseline in action efficiency on ARC-AGI-3.
- Efficient action usage: The model used fewer actions than the median tested human on 96% of levels.
Observed Behavior
A notable capability observed in GPT-6 Astra was its ability to transform unfamiliar environments into compact symbolic world models. Specifically, the model:
- Represented game mechanics as logical rules
- Developed its own domain-specific language shorthand to track state
- Used this shorthand to plan actions
Source
testingcatalogNeutral / independent
Part of this Story
OpenAI’s GPT-6 Astra scores 99% on ARC-AGI-3, surpassing humans on 96% of levels