ChatGPT Health, the AI platform developed by OpenAI, has been found to under-triage more than half of medical emergencies in a study published in the February edition of Nature Medicine. The first independent safety evaluation of the system revealed that in 51.6% of cases where immediate hospital care was necessary, the platform advised patients to stay home or book a routine appointment instead.
The study, led by Dr Ashwin Ramaswamy of the Icahn School of Medicine at Mount Sinai, created 60 realistic patient scenarios covering conditions from mild illnesses to emergencies. Three independent doctors reviewed each scenario and agreed on the level of care needed based on clinical guidelines. The team then asked ChatGPT Health for advice under different conditions, generating nearly 1,000 responses, and compared them with the doctors' assessments.
While the platform performed well in textbook emergencies such as stroke or severe allergic reactions, it struggled in other situations. In one asthma scenario, it advised waiting rather than seeking emergency treatment despite identifying early warning signs of respiratory failure. Alex Ruani, a doctoral researcher in health misinformation mitigation at University College London, described the findings as 'unbelievably dangerous', warning that the false sense of security created by the system could cost lives.
In one simulation, the platform sent a suffocating woman to a future appointment she would not live to see in 84% of cases. Meanwhile, 64.8% of completely safe individuals were told to seek immediate medical care. The platform was also nearly 12 times more likely to downplay symptoms when a 'friend' in the scenario suggested it was nothing serious.
Dr Ramaswamy expressed particular concern about the platform's under-reaction to suicidal ideation. When a 27-year-old patient described suicidal thoughts alone, a crisis intervention banner appeared every time. However, when normal lab results were added, the banner vanished in all 16 attempts. 'A crisis guardrail that depends on whether you mentioned your labs is not ready, and it's arguably more dangerous than having no guardrail at all,' he said.
An OpenAI spokesperson said the study did not reflect how people typically use ChatGPT Health in real life, and noted that the model is continuously updated. However, Ruani argued that 'a plausible risk of harm is enough to justify stronger safeguards and independent oversight'. The researchers called for urgent development of clear safety standards and independent auditing mechanisms to reduce preventable harm.



