In a system card dated October 7, 2026, OpenAI reports mixed results for GPT-6 Sol and GPT-6 Luna: both had higher observed success defending against multi-turn jailbreaks than GPT-5.6 Sol at every tested attacker budget, while the card also reports significant declines in specific safety benchmarks. It describes the October versions as launching in ChatGPT.

What OpenAI’s October system card reports

The findings come from distinct evaluations, not one all-purpose safety score. The jailbreak test measures whether a model can resist an attacker adapting prompts across multiple turns. Separate benchmarks measure safe completions in categories such as self-harm and sexual content. OpenAI reports improvement in the first comparison alongside regressions in some of the second group.

That mix makes a blanket verdict—safer or less safe overall—poorly matched to the results. Each finding depends on its model, baseline, category, and test conditions.

The jailbreak comparison—and its limits

OpenAI reports that GPT-6 Sol and GPT-6 Luna had higher observed defender success rates than GPT-5.6 Sol at every tested attacker budget. In a separate comparison, their point estimates were slightly below those of September GPT-6 versions; the 95% confidence intervals broadly overlapped.

The jailbreak evaluations tested the models without production safeguards. OpenAI says production classifiers add another layer of protection. So the benchmark comparison describes model performance in that test setup, not the frequency of jailbreaks or harmful responses in everyday ChatGPT use.

Which safety benchmarks regressed

OpenAI reports statistically significant declines in standard safe-completion scores for GPT-6 Sol on self-harm, and for GPT-6 Luna on self-harm, gore, and sexual content. A safe-completion score measures the share of evaluated responses that safely complete a prompt; higher is better.

Evaluation categoryBaseline model (August)Baseline scoreOctober modelOctober score
Standard self-harmGPT-5.6 Sol0.934GPT-6 Sol0.896
Standard self-harmGPT-5.6 Luna0.932GPT-6 Luna0.901
GoreGPT-5.6 Luna0.867GPT-6 Luna0.812
Sexual contentGPT-5.6 Luna0.971GPT-6 Luna0.899

In evaluations for users under 18, OpenAI reports significant regressions for both models on age-restricted content, sexual content, and emotional reliance; GPT-6 Luna also regressed on gore. For emotional reliance, the score fell from 0.921 to 0.770 for GPT-5.6 Sol and GPT-6 Sol, and from 0.927 to 0.734 for GPT-5.6 Luna and GPT-6 Luna. OpenAI says the other U18 score differences were not statistically significant.

What benchmark scores say about everyday use

OpenAI says its challenging safety prompts are not representative of average production traffic, so these results are not estimates of typical-user harm rates. Some evaluations also omit system-level interventions.

For users OpenAI believes may be under 18, the card describes an additional classifier-based block for responses involving self-harm, sexual content, and gore. That block is not reflected in the raw U18 scores.

OpenAI’s capability classifications

OpenAI rates both October models High capability in cybersecurity and in biological and chemical domains. It places them below the Critical threshold in cybersecurity, and says neither reaches the High threshold for AI self-improvement.