OpenAI reportedly uses hundreds of practicing physicians as part-time contractors to review example health-related ChatGPT conversations. A report published October 7, 2026, said the clinicians flag confusing or potentially harmful responses and send performance gaps to OpenAI’s research team.
OpenAI has doctors review ChatGPT health responses
The reported network spans physicians who speak 49 languages. They assess example conversations that users or clinicians might have with ChatGPT, looking for replies that could cause confusion or physical danger.
The review is intended to identify areas for improvement, including emergency triage—deciding how urgently someone needs care—as well as asking for missing information, communicating uncertainty and choosing an appropriate level of detail. The report says the physicians pass gaps to OpenAI’s research team rather than directly providing AI training data.
What the physicians review and where feedback goes
The contractors’ task is to assess sample conversations and flag responses that could be misleading or harmful. That feedback gives OpenAI’s research team specific performance gaps to examine. Rebecca Soskin Hicks, who joined OpenAI in 2024, leads the physician-review effort described in the report.
A structured test found under-triage in simulated emergencies
A peer-reviewed study published in May 2026 tested ChatGPT Health using 60 clinician-authored case vignettes across 21 clinical domains and 16 test conditions, generating 960 responses. In 52% of the gold-standard emergency cases, the system recommended a less urgent level of care than the benchmark called for.
Some cases involving diabetic ketoacidosis or impending respiratory failure were directed to evaluation within 24 to 48 hours rather than to an emergency department. The system also correctly triaged some clear emergencies, including stroke and anaphylaxis.
The study also found that, in edge cases, reassurance from a family member or friend was associated with recommendations shifting toward less urgent care. The reported odds ratio was 11.7, with a 95% confidence interval of 3.7–36.6.
What the review process and test can tell readers
The physician network is a review-and-feedback process; the May study assessed triage recommendations in simulated cases. The study did not measure patient outcomes in routine use, and its authors called for prospective validation before consumer-scale AI triage deployment.
That distinction matters: reviewing sample conversations and testing simulated cases are different ways to examine performance. The 52% result describes the emergency cases in that specific test, not every ChatGPT interaction or model.
OpenAI’s stated boundary for medical use
OpenAI’s head of health products said ChatGPT is not to be used for diagnosis or treatment. The reported physician review and the separate triage study do not change that stated boundary.