On September 14, a reported retest of GPT-5.6 Luna produced different answers after small changes to closely related prompts. That result matters because artificial general intelligence (AGI) implies broad, transferable ability—not just fluent answers on familiar-looking tasks. GPT-5.6’s published ARC-AGI-3 results add another wrinkle: the score changes substantially with the evaluation setup.
The evidence points to impressive but uneven performance. It does not settle whether GPT-5.6 meets any universal AGI threshold, partly because no universally accepted threshold exists.
A small wording change produced a different answer
The Luna retest used a Winograd schema, a type of language problem in which the correct interpretation depends on resolving an ambiguous pronoun. GPT-5.6 Luna reportedly answered one version correctly, identifying the brown suitcase as too small. After a minimal wording change in a parallel example, it reportedly answered that the table was too small, even though the intended antecedent was the car.
That is a narrow observation from a reported retest, not a failure rate for the entire GPT-5.6 family. Its significance is more practical: changing a few words can expose whether a model has transferred the underlying relationship or matched the surface pattern of the first prompt.
The same model also vacillated over a joke and a game
The retest covered three kinds of reasoning rather than one isolated puzzle. GPT-5.6 Luna reportedly gave inconsistent explanations of a Will Rogers joke about average intelligence in Oklahoma and California. It also alternated between treating a one-time 90-degree rotation of a tic-tac-toe board as strategically meaningful and treating it as a relabeling of an ordinary 3×3 board.
These examples test whether the model can preserve the relevant relationships when the presentation changes. They do not measure every form of reasoning, but they complicate any claim that polished language alone demonstrates general understanding.
Why GPT-5.6 Sol’s ARC-AGI-3 score changes so much
ARC-AGI-3 places agents in unfamiliar 2D game environments and asks them to infer how those environments work while acting efficiently. The published results for GPT-5.6 Sol vary according to the test split and harness configuration:
| Evaluation | Conditions | Result | What it measures |
| ARC-AGI-3 public set | ARC Prize results page’s listed evaluation | 13.33% | Performance on the public benchmark set |
| ARC-AGI-3 semi-private set | ARC Prize results page’s listed evaluation | 7.78% | Performance on the semi-private benchmark set |
| ARC-AGI-3 public set | OpenAI harness with retained reasoning and compaction | 38.3% | Performance under OpenAI’s reported configuration |
OpenAI reported that retaining reasoning and enabling compaction raised GPT-5.6 Sol’s public-set score from 13.3% to 38.3% in its own harness. Retained reasoning preserves reasoning across turns, while compaction replaces rolling truncation during long tasks. The benchmark result therefore belongs to the model and the surrounding evaluation system; separating those two is essential.
A later model must also remain a separate model. GPT-6 Astra’s reported ARC-AGI-3 figures—99.9% in OpenAI’s harness and 62.7% in the standard setup—describe GPT-6 Astra, not GPT-5.6. They are useful only as another example of how sharply benchmark results can depend on the harness.
What a specialized mathematics demonstration can—and cannot—tell us
A separate demonstration shows GPT-5.6 Sol Pro working through advanced mathematics, including Riemann–Hilbert analysis, oscillatory integrals and determinant representations. The model generates LaTeX-formatted responses, while the presenter discusses verbosity, mistakes and the role of prompt engineering and custom agent harnesses.
Specialist performance is valuable evidence of capability. It is not the same as reliable transfer across unfamiliar tasks, changing instructions and messy real-world goals.
ARC-AGI-3 is not a complete AGI test
ARC-AGI-3 probes a constrained skill: learning the rules of unfamiliar game environments and acting within them. François Chollet has pointed out that real-world intelligence involves much longer learning horizons, more complex world models, greater goal ambiguity and larger exploration spaces than these games.
That distinction answers the central question directly: GPT-5.6’s ARC-AGI-3 scores do not, by themselves, demonstrate AGI. A benchmark score describes performance under defined conditions. AGI is generally discussed as the ability to match or exceed human abilities across virtually every cognitive task, but the field has no universal threshold for declaring that milestone reached.
The useful test is transfer, not just fluency
For GPT-5.6, the most revealing evidence sits in the gap between different kinds of success: a fluent answer can coexist with a different answer after a small prompt change; a benchmark score can rise when the harness retains more reasoning and compacts long context; and a strong mathematics demonstration can still require expert oversight and careful prompting.
The practical question is therefore not whether GPT-5.6 can produce an impressive answer. It is whether the model can preserve the right relationships, recover when conditions change and deliver reliable results across unfamiliar tasks. The current tests make that question sharper; they do not close it.