A reported Bloomberg Intelligence estimate put leading Chinese AI models 3% behind U.S. rivals on benchmark scores after DeepSeek released V4.1 Flash in September 2026. The same comparison had put the gap at 15% earlier in the year and about 9% in May.
The reported estimate after DeepSeek V4.1 Flash
The 3% figure describes an estimated difference in benchmark scores between leading models from China and the United States. DeepSeek V4.1 Flash’s release in September preceded that estimate; the sequence does not, by itself, describe performance on every AI task.
One estimate, three points in time
The reported Bloomberg Intelligence comparison gives this sequence:
| Comparison point | Estimated gap in benchmark scores |
| Earlier in 2026 | 15% |
| May 2026 | About 9% |
| After DeepSeek V4.1 Flash’s September 2026 release | 3% |
Stanford’s separate March 2026 comparison
Stanford HAI’s 2026 AI Index gives a distinct result: as of March 2026, Anthropic’s top model led by 2.7% in its U.S.–China model-performance comparison. The Index also says models from the two countries traded the lead multiple times beginning in early 2025.
For a different angle on Chinese AI, NeoTeo previously covered Mozilla’s assessment of Chinese open-weight models.
What the gap means for a specific task
The 3% estimate concerns leading models’ benchmark scores. For production work, the useful comparison is how the specific models perform on the tasks a team needs to run; a country-level estimate does not answer that question for every workflow.