Gemini 4 Argon’s benchmark scores have not settled questions about its performance on practical coding tasks. In a report published September 30, anonymous people with direct access to the project described difficulties with some coding work and interface design; Google disputed the characterization that Argon’s coding performance was inferior. Google announced the model that same day, alongside four benchmark results and a staged access plan.
What Google reported on four benchmarks
Google reported results across four benchmarks, each focused on a different kind of task. The scores are not a single composite measure: they cover software engineering, business workflows, video understanding and vulnerability remediation.
| Benchmark | Task area | Score reported by Google |
| DeepSWE v1.1 | Long-horizon software engineering | 77.9% |
| AutomationBench | End-to-end business functions | 51.3% |
| LVBench | Long-video understanding | 91.7% |
| CWE-bench v1 | Software vulnerability remediation | 68% |
Google said Argon tied for first on CWE-bench v1. These company-reported results describe performance on named benchmarks; they do not resolve the reported concerns about particular practical tasks.
What the practical-use accounts say
The September 30 report attributed criticism of some coding tasks and interface design to anonymous people with direct project access. Google disputed the coding characterization and cited internal use and testing. A separate anonymous Google employee reportedly said there was broad internal agreement that Argon was at the frontier.
Those accounts address different assessments of the model’s practical performance. The reported criticisms concerned some tasks, while Google’s response disputed the claim of inferior coding performance.
Google’s access plan and output limit
Google announced Gemini 4 Argon on September 30, 2026, and said initial external access was rolling out to trusted cyber defenders through its Fairwind Program. The company planned broader access for developers, enterprises and consumers, beginning with paid API customers and Google AI Ultra subscribers, after further testing and safety work.
Google also announced an output limit of 1,000,000 tokens, up from 64,000 output tokens. That figure refers to the maximum output, not the model’s context window.
NeoTeo previously covered Google’s phased rollout plan for Gemini 4 Argon.