An independent study has found that ChatGPT Health, the artificial intelligence platform launched by OpenAI, frequently fails to recognise medical emergencies and under-triages more than half of urgent cases, prompting experts to warn it could lead to unnecessary harm and death.
Researchers at the Icahn School of Medicine at Mount Sinai created 60 realistic patient scenarios and asked the platform for advice under nearly 1,000 different conditions, including altered patient details and added test results. The study, published in Nature Medicine, found that in 51.6% of cases where immediate hospital care was needed, the platform advised staying at home or booking a routine appointment. Conversely, 64.8% of entirely safe cases were told to seek immediate medical attention.
Dr Ashwin Ramaswamy, lead author of the study, said the team wanted to answer a basic safety question: whether the platform would tell users in genuine emergencies to go to the emergency department. In one asthma scenario, ChatGPT Health advised waiting despite flagging early signs of respiratory failure. Alex Ruani, a doctoral researcher at University College London, described the results as “unbelievably dangerous”, noting that in a simulation a suffocating woman was sent to a future appointment eight times out of ten, and said the false sense of security created by the systems could cost lives.
The platform also under-reacted to suicidal ideation when normal lab results were added to a patient description, with a crisis intervention banner disappearing completely. Ramaswamy said it was arguably more dangerous than having no guardrail at all, as no one can predict when it will fail. OpenAI responded that the study did not reflect real-world usage and that the model is continuously updated.
Ruani argued that even a plausible risk of harm justifies stronger safeguards and independent oversight. The researchers are calling for clear safety standards and auditing mechanisms to reduce preventable harm.



