Key points
- OpenAI recently rolled out ChatGPT Health despite poor evidence of efficacy.
- Disclaimers and caveats are likely to fall on deaf ears.
- Industry leaders continue to call for greater restraint; but it may be too late for public health.
More than 300 million people consult ChatGPT for health-related questions every week. To accommodate the demand, OpenAI quietly rolled out ChatGPT Health last month to a select group of US adults; a soft opening of sorts. For the world’s first patient-focused AI, there was surprisingly little fanfare, and I can’t help but think there was a reason.
ChatGPT Health is based on the same underlying architecture as regular ChatGPT, but with a key difference: with permission, the AI can link to your medical records and download health metrics recorded by Apple wearables, such as the Apple Watch. Though initially a separate tab in ChatGPT’s interface, it’s now seamlessly integrated; an operational decision by OpenAI after seeing users posing health questions to the chatbot organically during daily conversations. An OpenAI press release said ChatGPT Health can “help explain a visit note, understand lab results, dig into doctor-patient discussions, and help prepare questions for follow-up appointments…” The platform can optimize the user’s experience with more “personalized conversations.”
But this initiative has tension at its heart. On the one hand, OpenAI qualifies the rollout by saying it’s not for diagnosis or treatment and is “not designed to replace the care and judgment of qualified medical professionals.” On the other hand, that’s exactly what the public will use it for, and OpenAI knows it. “The reality is, we’re living in a time when not everyone has access to quality care or care in a timely manner,” said one executive to reporters last month. “The average doctor appointment in the United States is less than 15 minutes… You’re often left on your own to sort through a lot of fragmented information.” For many, the black box that is ChatGPT Health will be used in the absence of a physician; after all, if users had access to their doctor, what would they need from AI?
A Limited Skillset
When ChatGPT launched in 2022, it operated exclusively on its “parametric” architecture; that is, it was entirely constrained by the reams of text-based data on which it had been trained, hence the name Large Language Model (LLM). A query such as “What’s the best source of vitamin C?” would prompt the AI to draw on the statistical model it developed during training to predict the most likely word sequence in a response. The model’s strength was its conversation mimicry: fluent, confident, and generally articulate. The weakness was that the training data were scattered, comprising everything from open-access research papers and mainstream articles to blog posts, Q&A forums (e.g., Reddit), and even social media—a cesspit of misinformation. A user asking about vitamin C would need to hope that “guava, kiwi, and citrus fruit” appeared more often in the training data than “sausages and caviar.” Worse still, if the query wasn’t covered in training, the AI would fabricate an answer, like a teenager trying to impress his friends. The hallucination was born.
But AI is evolving at a pace that makes our own evolutionary history appear positively glacial. In the spring of 2023, ChatGPT was upgraded with its RAG plugin, short for retrieval-augmented generation. This software bolt-on lets current AI generations perform online searches in real-time when the AI deems its training data insufficient for a full-throated reply. And while this hands AI the keys to the full wealth of human digital knowledge, there’s a wrinkle: finding data isn’t the same as evaluating it. You can give someone access to the entire PubMed catalog, but it doesn’t mean they’ll be able to make clinical decisions. So, for AI and humans alike, access to information isn’t the limiting factor; rather, it’s the ability to distinguish among quality sources, weigh evidence, and apply knowledge and experience. These are qualities AI doesn’t possess.
The Shallow Body of Evidence
Programming an AI with these high-level skills is certainly possible, but it’s proving complex. And published audits on health-related content suggest better AI reasoning will be crucial to the success of initiatives like ChatGPT Health.
One recent study found that chatbots given medical queries produced “problematic” responses one-fifth to one-half of the time, and “unsafe” responses varied from 5 to 13% of the total. Our own audit of five AI chatbots, published earlier this year in the British Medical Journal, recorded “problematic” answers in roughly 50% of chatbot responses to health questions in misinformation-prone fields; one-fifth were expert-rated as “highly problematic,” and could potentially cause harm if followed. These are two among dozens of studies on the efficacy of AI health-related outputs, showing them to be inconsistent and unreliable.
The limited data on ChatGPT Health, in particular, is no less disconcerting. In a structured test of triage recommendations, Ramaswamy and colleagues at Mount Sinai Health System, NY, reported that 52% of gold-standard emergencies were under-triaged, directing patients with diabetic ketoacidosis or impending respiratory failure to 24–48 h evaluation rather than the emergency department. A preprint awaiting peer review similarly indicates that ChatGPT Health agrees with nurse and physician triage decisions only 50% of the time.
A 50% under-triage rate for genuine emergencies is not a marginal performance defect; it has major implications with asymmetric consequences: over-triage a minor ailment risks stressing a medical system already near breaking point, but under-triaging a real emergency can be catastrophic. The evidence suggests AI is not a suitable tool for first-contact triage. Try telling that to the >40 million people who turn to ChatGPT with healthcare questions every day.
None of this means that lower-risk functions are unsafe. For example, the AI may accurately translate clinical terminology, organize medical records, plot and analyze lab results chronologically, and help users prepare questions for their clinician. The difficulty is that ChatGPT Health does not maintain a dependable boundary between those functions and clinical judgment; the moment it interprets symptoms and offers advice, it becomes a functioning triage system, irrespective of OpenAI’s disclaimer. There’s a mismatch between what the system is designed to do and how the end user will ultimately use it. And there’s no evidence that the AI itself understands that limitation.
Artificial Intelligence Essential Reads
Should You Trust the Algorithm or Your Gut?
Does AI Need a Psychologist?
The good news is that OpenAI continues to evaluate ChatGPT Health using an assessment matrix it developed called Health Bench. Created in collaboration with 262 physicians from 60 countries, across 26 medical specialties, Health Bench periodically audits thousands of conversations to identify and repair weaknesses in the platform. It’s a robust and necessary step, but no data have been published yet, and it’s clearly a work in progress rather than a finished product.
Such public-facing initiatives continue even as industry leaders call for greater restraint. In September 2026, OpenAI’s Sam Altman joined Anthropic’s Dario Amodei and Elon Musk in calling for a slowdown in advanced AI development, with stronger safeguards and independent oversight. The same principle applies to healthcare: safety checks must keep pace with capability. That isn’t happening.
Consider the consequences if Ford rolled out a new truck before proving it was safe to drive. We wouldn’t accept a half-baked product or the manufacturer’s reassurance that it was a “work in progress.” We should demand the same accountability from companies with so much potential influence on public health. Without that accountability, we are the last phase of the experiment.
