A comparison published September 28, 2026, reported that Claude Opus 5.5 completed all 15 runs perfectly across three developer tasks, while GPT-6 Sol completed 12 perfectly. GPT-6 Sol was faster and less costly per run in each task. Each model ran the same prompts five times per task, using its highest-effort setting, described as “max.”

The three repeated developer tasks

The test covered CI triage, outage-log analysis and a dependency resolver built from a specification. Each task had its own way to judge a successful result.

For CI triage, each model had to assess 40 failed jobs alongside a conditional runbook and decide whether to retry, block or page. Both models had five perfect runs out of five.

The outage-log task used 3,664 lines from five services during a two-hour outage. Each run had to answer seven questions. For the resolver task, the models implemented a dependency resolver from a two-page specification, then faced a suite of 120 hidden tests.

Results by task: perfect runs, time and cost

The figures below are per-run averages for each model across five runs of each task.

TaskPerfect runs: Claude Opus 5.5 / GPT-6 SolAverage time per run: Claude Opus 5.5 / GPT-6 SolReported cost per run: Claude Opus 5.5 / GPT-6 Sol
CI triage5/5 / 5/51:27 / 0:18$0.24 / $0.02
Outage-log analysis5/5 / 2/56:44 / 1:41$1.68 / $0.30
Dependency-resolver specification5/5 / 4/59:40 / 6:28$1.42 / $0.22

The pattern was consistent across this test set: Opus 5.5 had more perfect runs overall, while Sol had a lower average time and reported cost in every task. Those task-level results describe this comparison, not a general ranking of the models.

What GPT-6 Sol missed

In the outage-log task, two Sol runs missed the same customer, whose original charge was confirmed 52 seconds after a successful retry. Another run counted 27 failed checkouts instead of 28.

In one resolver run, a stray closing parenthesis on line 78 caused an import failure. That run failed all 120 hidden tests. The example illustrates why the perfect-run count can matter alongside speed and cost: a small code error can prevent a solution from running at all.

A single CI call is not AutomationBench

The CI triage comparison was a single API call. AutomationBench, by contrast, is described as an end-to-end workflow using 47 tools, so the two evaluations measure different kinds of work.

The model pairings and settings differ, too. OpenAI’s launch comparison used GPT-6 Sol at xhigh and Claude Opus 5 at max; the repeated test compared GPT-6 Sol with Claude Opus 5.5, using each model’s highest-effort setting, labeled “max.”