A reported Bloomberg Intelligence estimate put leading Chinese AI models 3% behind U.S. rivals on benchmark scores after DeepSeek released V4.1 Flash in September 2026. The same comparison had put the gap at 15% earlier in the year and about 9% in May.

The reported estimate after DeepSeek V4.1 Flash

The 3% figure describes an estimated difference in benchmark scores between leading models from China and the United States. DeepSeek V4.1 Flash’s release in September preceded that estimate; the sequence does not, by itself, describe performance on every AI task.

One estimate, three points in time

The reported Bloomberg Intelligence comparison gives this sequence:

Comparison pointEstimated gap in benchmark scores
Earlier in 202615%
May 2026About 9%
After DeepSeek V4.1 Flash’s September 2026 release3%

Stanford’s separate March 2026 comparison

Stanford HAI’s 2026 AI Index gives a distinct result: as of March 2026, Anthropic’s top model led by 2.7% in its U.S.–China model-performance comparison. The Index also says models from the two countries traded the lead multiple times beginning in early 2025.

For a different angle on Chinese AI, NeoTeo previously covered Mozilla’s assessment of Chinese open-weight models.

What the gap means for a specific task

The 3% estimate concerns leading models’ benchmark scores. For production work, the useful comparison is how the specific models perform on the tasks a team needs to run; a country-level estimate does not answer that question for every workflow.