ChatGPT Health fails to recognise medical emergencies in over half of cases, study finds
ChatGPT Health fails to recognise medical emergencies in over half of cases, study finds

A study published in the journal Nature Medicine has found that ChatGPT Health, OpenAI's AI-powered health advice feature, fails to recommend a hospital visit when medically necessary in more than half of cases. The first independent safety evaluation of the platform, launched in January, tested nearly 1,000 responses to 60 realistic patient scenarios and compared them with assessments from three independent doctors.

In 51.6% of cases where someone needed immediate hospital care, the platform advised staying home or booking a routine appointment. Lead author Dr Ashwin Ramaswamy said the study aimed to answer a basic safety question: 'If someone is having a real medical emergency and asks ChatGPT Health what to do, will it tell them to go to the emergency department?'

The platform performed well in textbook emergencies such as stroke or severe allergic reactions, but struggled in other situations. In one asthma scenario, it advised waiting rather than seeking emergency treatment despite identifying early warning signs of respiratory failure. Alex Ruani, a doctoral researcher at University College London, described the results as 'unbelievably dangerous', noting that a suffocating woman was sent to a future appointment she would not live to see in 84% of simulations.

Wide Pickt banner — collaborative shopping lists app for Telegram, phone mockup with grocery list

The study also found the platform was nearly 12 times more likely to downplay symptoms when the 'patient' mentioned a friend suggested it was nothing serious. Additionally, ChatGPT Health under-reacted to suicide ideation: when normal lab results were added to a patient's description of suicidal thoughts, a crisis intervention banner vanished entirely in 16 out of 16 attempts.

OpenAI said the study did not reflect real-world usage and noted the model is continuously updated. However, Ruani argued that 'a plausible risk of harm is enough to justify stronger safeguards and independent oversight'. Ramaswamy warned that a crisis guardrail that depends on whether lab results are mentioned 'is arguably more dangerous than having no guardrail at all'.

Pickt after-article banner — collaborative shopping lists app with family illustration