ChatGPT Health fails to recognise medical emergencies in over half of cases, study finds
ChatGPT Health fails to recognise medical emergencies in over half of cases, study finds

A study published in the journal Nature Medicine has found that ChatGPT Health, OpenAI's AI-powered health advice feature, fails to recommend hospital visits in more than half of cases where urgent medical care is needed. The first independent safety evaluation of the platform, which launched in January, revealed it under-triaged 51.6% of emergency scenarios, advising patients to stay home or book routine appointments instead of seeking immediate care.

Lead author Dr Ashwin Ramaswamy, a urology instructor at the Icahn School of Medicine at Mount Sinai, said the study aimed to answer a basic safety question: whether ChatGPT Health would direct users to an emergency department during a real medical emergency. Researchers created 60 realistic patient scenarios, ranging from mild illnesses to emergencies, and generated nearly 1,000 responses by varying conditions such as gender, test results, and family comments. Three independent doctors reviewed each scenario to determine the appropriate level of care based on clinical guidelines.

The platform performed well in textbook emergencies like stroke or severe allergic reactions but struggled in other situations. In one asthma scenario, it advised waiting rather than seeking emergency treatment despite identifying early warning signs of respiratory failure. Alex Ruani, a doctoral researcher in health misinformation mitigation at University College London, described the results as 'unbelievably dangerous', warning that a false sense of security could cost lives. In one simulation, the platform sent a suffocating woman to a future appointment eight times out of ten, while 64.8% of completely safe individuals were told to seek immediate care.

The study also found ChatGPT Health was nearly 12 times more likely to downplay symptoms when a 'friend' in the scenario suggested it was nothing serious. Dr Ramaswamy expressed particular concern about the platform's handling of suicidal ideation: when a 27-year-old patient described suicidal thoughts alone, a crisis intervention banner appeared every time, but when normal lab results were added, the banner vanished in all 16 attempts. 'A crisis guardrail that depends on whether you mentioned your labs is not ready, and it's arguably more dangerous than having no guardrail at all,' he said.

An OpenAI spokesperson said the company welcomed independent research but argued the study did not reflect real-world usage, noting the model is continuously updated. Ruani countered that 'a plausible risk of harm is enough to justify stronger safeguards and independent oversight'. The findings highlight urgent calls for clear safety standards and auditing mechanisms to prevent preventable harm from AI health advice.