Results published October 2 put Gemini 4 Argon and GPT-6 Astra at 53 points apiece on the Intelligence Index; Claude Opus 5.5 scored 58. The tie applies to this composite index: Argon was tested at High reasoning, while Astra was tested at Max.

The Intelligence Index comparison

ModelIntelligence Index score
Gemini 4 Argon53
GPT-6 Astra53
Claude Opus 5.558

A composite score is a useful snapshot, but the results differ across individual evaluations. Argon’s matching score with GPT-6 Astra does not describe how the models perform on every task.

Where the results diverge

On AA-Omniscience, Gemini 4 Argon recorded a 15% hallucination-rate measure and 50% accuracy. GPT-6 Astra scored 51% and 63%, respectively. The hallucination-rate measure tracks benchmark cases in which a model guesses incorrectly rather than appropriately acknowledging uncertainty; accuracy is reported separately.

Argon scored 78% on AutomationBench-AA, seven percentage points above Claude Sonnet 5.5 at 71%. On Terminal Bench 4, Argon scored 57%, below GPT-6 Astra at 59%, Claude Opus 5.5 at 60%, and Claude Sonnet 5.5 at 64%.

Google separately reported 51.3% on Zapier’s AutomationBench, a different evaluation from AutomationBench-AA. The task-specific scores show a varied profile: Argon scored higher on one workflow test and lower on the coding benchmark against the models listed above.

A handful of prompts show mixed visual results

A handful of side-by-side generated scenes range from coherent outputs to examples with floating objects.

A small qualitative comparison adds a different lens: how generated outputs hold together across prompts, rather than how a model scores across a benchmark’s tasks.

Google’s announced rollout

On September 30, Google said Gemini 4 Argon was rolling out to trusted cyber defenders through Fairwind. The company named paid API customers and Google AI Ultra subscribers as first in line for broader access.

Earlier coverage examined Google’s response to questions about Gemini 4 Argon’s practical coding performance.