ARC Prize reported three GPT-6 Astra scores on ARC-AGI-3 Semi-Private: 62.7% with the Standard harness at maximum reasoning effort, 98.6% with the Provider Adapter at maximum effort, and 99.9% with the Provider Adapter at high effort. The results, published on September 3, 2026, depend on both the harness and the reasoning setting. ARC Prize says they do not establish that Astra is artificial general intelligence (AGI).
For background on the model’s launch, see NeoTeo’s earlier coverage of GPT-6 Astra. The ARC-AGI-3 results address a narrower question: how well Astra performed in a particular family of interactive tasks.
GPT-6 Astra’s ARC-AGI-3 scores, side by side
A harness is the setup that connects a model to a test and determines how it interacts with that test. Here are the three results ARC Prize reported for the Semi-Private set:
| Evaluation condition | Score | Reasoning effort | Context handling |
| Standard harness | 62.7% | Maximum | Provider-neutral interface |
| Provider Adapter | 98.6% | Maximum | Preserves opaque reasoning state between requests and uses compaction |
| Provider Adapter | 99.9% | High | Preserves opaque reasoning state between requests and uses compaction |
The 99.9% result is not the Provider Adapter’s maximum-effort score: that result was 98.6%. And the 62.7% Standard result used a different harness. Keeping those conditions together gives each figure its proper meaning.
What the two harnesses change
ARC-AGI-3 tests agents in novel, abstract, turn-based environments. Instead of spelling out every instruction, it requires a model to explore, infer goals, build a working understanding of the environment, plan and act.
The Standard harness is provider-neutral, giving models a common interface for comparison across providers. The Provider Adapter uses provider-specific context-management features: it can preserve opaque reasoning state between requests and use compaction to manage long conversations. That makes the setups meaningfully different; the score is a result of a model operating under a specified test configuration, not a setting-free measure of intelligence.
What Astra’s human action comparison measures
ARC Prize also compared action counts with a human baseline. The foundation tested about 500 members of the general public; for each level, the baseline was the median number of actions taken by the people who completed it.
In the Provider Adapter run at maximum reasoning effort, Astra used fewer actions than that median baseline on 96.0% of levels. Across levels, it used 51.7% fewer actions per level on average. Those figures describe action efficiency on ARC-AGI-3 under that configuration. They do not measure performance across the full range of human tasks.
Why these results do not establish AGI
ARC Prize defines its AGI objective in terms of acquiring any skill a human can, as efficiently as a human. ARC-AGI-3 offers a way to test aspects of learning and acting toward that objective, but its environments are bounded, deterministic and closed-ended. The foundation says they do not capture the complexity of the real world.
A high score therefore reflects performance on the benchmark’s task family. ARC Prize describes Astra’s results as progress toward generalization, while explicitly saying it is not claiming that GPT-6 Astra is AGI.