A scoping review published in PLOS Digital Health on October 9, 2026, reports that GPT-4o scored 552 points (92.00%) and DeepSeek-R1 scored 523 on evaluations of China’s 2024 National Medical Licensing Examination. The review summarizes a 600-question written evaluation: every question was text-only and multiple choice, so the result was a test of written answers, not clinical care.

The review says both models exceeded 60% accuracy, the passing criterion it standardized as the most commonly used across the Chinese medical exams in its analysis. The scores are research results; they do not establish that either model is clinically competent or ready to practice medicine.

What the 2024 scores measured

The National Medical Licensing Examination is an official route to physician certification in China. The review describes the GPT-4o and DeepSeek-R1 evaluation as using 600 questions from the examination’s written portion. It reports these raw scores:

ModelPoints reported in the review
GPT-4o552
DeepSeek-R1523

Those results come from an evaluation summarized by the review, not a new exam administered by the review’s authors. A multiple-choice score can show how a model handled that question set under those evaluation conditions. It cannot, on its own, show how the model would assess a patient, make clinical decisions, or provide care.

What the wider review adds

The review brought together 14 peer-reviewed studies, 51 evaluation records and eight types of Chinese medical examinations. It covered nine language models, while the 2024 CNMLE scores are one specific comparison within that wider collection.

A separate exploratory analysis compared evaluation records for GPT-3.5 and GPT-4 that reported pass-or-fail outcomes. None of 24 GPT-3.5 records met the review’s passing threshold; nine of 10 GPT-4 records did. The review reported P < 0.001, but characterized the comparison as descriptive rather than definitive evidence that one model is generally superior: the records varied in exams, datasets, prompts and evaluation settings.

The review also rated 10 of the 14 studies as high quality and four as moderate using a custom 12-item checklist. The authors noted that the checklist had not been formally validated, so those ratings are indicative. In one included study, GPT-4 scored 84% on the Chinese original of the exam and 86% on an English translation—a result from that study, not a general rule about model performance across languages.

Differences in exam years, question selection, prompts, model versions, access methods and scoring make results across studies harder to compare. The review also says possible exposure of questions through public online material means training-data contamination or memorization cannot be ruled out.