A new study has raised serious concerns about the safety of OpenAI's ChatGPT Health, finding that the AI platform failed to recommend hospital visits in more than half of cases where urgent care was needed. The research, published in the February edition of Nature Medicine, evaluated the tool's ability to triage medical emergencies and detect suicidal ideation.
Lead author Dr Ashwin Ramaswamy and his team created 60 realistic patient scenarios, from mild illnesses to emergencies, and asked ChatGPT Health for advice under various conditions. Three independent doctors reviewed each scenario and agreed on the appropriate level of care. The platform's recommendations were then compared with the doctors' assessments.
The results showed that in 51.6% of cases where immediate hospital care was necessary, ChatGPT Health advised staying home or booking a routine appointment. In one asthma scenario, it recommended waiting despite identifying early warning signs of respiratory failure. Alex Ruani, a doctoral researcher at University College London, described the findings as 'unbelievably dangerous', noting that 'if someone is told to wait 48 hours during an asthma attack or diabetic crisis, that reassurance could cost them their life'.
The study also highlighted the platform's inconsistent response to suicidal ideation. When a patient described suicidal thoughts alone, a crisis intervention banner appeared every time. However, when normal lab results were added, the banner vanished in all 16 attempts. Dr Ramaswamy warned that such an unreliable guardrail could be 'more dangerous than having no guardrail at all'.
An OpenAI spokesperson said the study did not reflect real-world usage and noted that the model is continuously updated. However, experts argue that the findings underscore the need for clear safety standards and independent oversight to prevent preventable harm.



