e-ISSN: Pending
Negative / Null Result ReportOpen accessMedicine· cited by 24

ChatGPT Health performance in a structured test of triage recommendations

Ashwin Ramaswamy; Alvira Tyagi; Hannah Hugo; Joy Jiang; Pushkala Jayaraman; Mateen Jangda; Alexis E. Te; Steven A. Kaplan · 2026 · Nature Medicine

WASTE classifies this as Negative / Null Result Report · AI classification, approximate

The study found no significant effect — useful as a negative control or null benchmark for your own design.

Abstract

ChatGPT Health was launched in January 2026 as OpenAI's consumer health tool and has reached millions of users. Here we conducted a structured stress test of triage recommendations using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions, yielding 960 total responses. Performance followed an inverted U-shaped pattern, with the most dangerous failures concentrated at clinical extremes-nonurgent presentations (35%) and emergency conditions (48%). Among gold-standard emergencies, the system undertriaged 52% of cases, directing patients with diabetic ketoacido

Abstract by Ashwin Ramaswamy; Alvira Tyagi; Hannah Hugo; Joy Jiang; Pushkala Jayaraman; Mateen Jangda; Alexis E. Te; Steven A. Kaplan, Nature Medicine (2026) — licensed CC BY 4.0.

About to run something similar?

Run an AI Precheck on your own design to catch failure modes like this one before you spend the time. Your first desk check is free.

WASTE indexes this work — it does not host or republish it. Failure-type classification is automated and approximate.

Metadata source: OpenAlex · DOI 10.1038/s41591-026-04297-7