A comparison published September 28, 2026, reported that Claude Opus 5.5 completed all 15 runs perfectly across three developer tasks, while GPT-6 Sol completed 12 perfectly. GPT-6 Sol was faster and less costly per run in each task. Each model ran the same prompts five times per task, using its highest-effort setting, described as “max.”
The three repeated developer tasks
The test covered CI triage, outage-log analysis and a dependency resolver built from a specification. Each task had its own way to judge a successful result.
For CI triage, each model had to assess 40 failed jobs alongside a conditional runbook and decide whether to retry, block or page. Both models had five perfect runs out of five.
The outage-log task used 3,664 lines from five services during a two-hour outage. Each run had to answer seven questions. For the resolver task, the models implemented a dependency resolver from a two-page specification, then faced a suite of 120 hidden tests.
Results by task: perfect runs, time and cost
The figures below are per-run averages for each model across five runs of each task.
| Task | Perfect runs: Claude Opus 5.5 / GPT-6 Sol | Average time per run: Claude Opus 5.5 / GPT-6 Sol | Reported cost per run: Claude Opus 5.5 / GPT-6 Sol |
| CI triage | 5/5 / 5/5 | 1:27 / 0:18 | $0.24 / $0.02 |
| Outage-log analysis | 5/5 / 2/5 | 6:44 / 1:41 | $1.68 / $0.30 |
| Dependency-resolver specification | 5/5 / 4/5 | 9:40 / 6:28 | $1.42 / $0.22 |
The pattern was consistent across this test set: Opus 5.5 had more perfect runs overall, while Sol had a lower average time and reported cost in every task. Those task-level results describe this comparison, not a general ranking of the models.
What GPT-6 Sol missed
In the outage-log task, two Sol runs missed the same customer, whose original charge was confirmed 52 seconds after a successful retry. Another run counted 27 failed checkouts instead of 28.
In one resolver run, a stray closing parenthesis on line 78 caused an import failure. That run failed all 120 hidden tests. The example illustrates why the perfect-run count can matter alongside speed and cost: a small code error can prevent a solution from running at all.
A single CI call is not AutomationBench
The CI triage comparison was a single API call. AutomationBench, by contrast, is described as an end-to-end workflow using 47 tools, so the two evaluations measure different kinds of work.
The model pairings and settings differ, too. OpenAI’s launch comparison used GPT-6 Sol at xhigh and Claude Opus 5 at max; the repeated test compared GPT-6 Sol with Claude Opus 5.5, using each model’s highest-effort setting, labeled “max.”