In the October 6, 2026, Artificial Analysis Intelligence Index v4.3.2 comparison, Mistral Large 4 Preview scored 38. Five listed Chinese open-weight models scored higher, from 39 to 46. Its cybersecurity results told a more varied story: Mistral scored 50 on the Cyber Index and 81.7% on one separate component test. NeoTeo previously covered Mistral Large 4’s public-preview announcement.
How Mistral Large 4 Preview scored against open-weight models
The comparison placed Mistral Large 4 Preview below five listed Chinese open-weight models. Xiaomi MiMo-V2.6-Pro led these six entries at 46, eight points above Mistral’s 38.
| Model | Intelligence Index score |
| Xiaomi MiMo-V2.6-Pro | 46 |
| Z.ai GLM-5.3 (max) | 45 |
| Moonshot Kimi K3 (max) | 44 |
| Z.ai GLM-5.3-Flash | 42 |
| DeepSeek V4.1 Flash (max) | 39 |
| Mistral Large 4 Preview | 38 |
These are scores from one named index and one dated comparison, not a single verdict on performance across every kind of task. The cybersecurity measures, for example, produce a different ordering.
Cybersecurity results vary by test
Mistral Large 4 Preview scored 50 on the Cyber Index, matching Z.ai GLM-5.3-Flash in the results below. The Cyber Index is the aggregate measure; CWE-Bench-AA, DeepsecBench-AA and CyberGym-E2E-AA are separate component tests.
| Model and variant | Cyber Index | CWE-Bench-AA | DeepsecBench-AA | CyberGym-E2E-AA |
| Grok 4.7 (xhigh) | 56 | 68% | 27% | 74% |
| Xiaomi MiMo-V2.6-Pro | 56 | 63% | 26% | 79% |
| OpenAI GPT-6 Luna (max) | 53 | 57% | 24% | 78% |
| Z.ai GLM-5.3-Flash | 50 | 56% | 20% | 74% |
| Mistral Large 4 Preview | 50 | 51% | 16% | 81.7% |
The 81.7% CyberGym-E2E-AA result is Mistral’s strongest figure in these component tests, but it does not replace the aggregate Cyber Index score. The other two component results—51% on CWE-Bench-AA and 16% on DeepsecBench-AA—show why the individual test matters when interpreting the profile.
What creator-run coding demonstrations show
Separate coding demonstrations tested tasks that differ from the index comparison. In one, a Wikipedia redesign placed 13th on the creator’s leaderboard, while a LEGO South Park game world placed 9th. The demonstrations showed a narrow, mobile-style redesign with animation issues, and gameplay problems involving movement and physics.
A separate eight-task KingBench 3 run used OpenCode through OpenRouter and reported a permission-corrected score of 42/80 (52.5%). An elevator simulation and a Three.js folding table worked in the demonstration; default runs of a contact-lens case and a counting task produced no file or answer. Retry and repair runs were excluded from the score.
These task-specific results offer a practical view of generated prototypes, but they are separate creator-run demonstrations, not standardized substitutes for the index scores.
The preview and Mistral’s planned weight release
The benchmark comparison evaluated Mistral Large 4 Preview through Mistral’s API. Mistral planned to release the model’s weights by the end of October 2026; that was the company’s stated plan when it announced the preview.