OpenAI announced today that ChatGPT Health will now be available to all users over 18, and on all plans, including free ones. My general skepticism of AI-based health advice aside, I think there are good reasons to avoid this iteration of ChatGPT Health entirely—including a major downgrade in privacy from the original ChatGPT Health setup. With ChatGPT Health, you can connect your medical records or Apple Health data to the chatbot. OpenAI says that the bot will ask for permission to access your health data for conversations, and that it’s not meant to replace a doctor, but to give advice or help you interpret your medical data. Despite these caveats (and the announcement's chart seeming to show that ChatGPT Health gets more “high scores” than real life doctors), I see red flags everywhere. Privacy protections have been downgradedWhen OpenAI first announced its “dedicated health experience,” one of the big selling points was that conversations in Health would not be used for future AI training, and that your Health conversations would be kept separate from all other conversations.
But the terms have now shifted. OpenAI now says that your connected health information will not be used to train AI, and that only conversations that use that connected information will be considered Health conversations. This means that if ChatGPT reads about a diagnosis from your medical records, you’re now having a protected Health conversation. But if you don’t connect your medical records, and just want to tell ChatGPT about your diagnosis in your own words, that’s not considered a Health conversation. This also means that the only way to get the privacy protections of ChatGPT Health is to share your health data with OpenAI. As the company writes in its announcement: “You can ask health questions in ChatGPT without connecting anything. To use Health and get responses grounded in your own information, you can choose to connect Apple Health, supported medical records from U.S. hospital systems, One Medical, or Function Health.” The former conversations would have privacy protections. The latter would not.OpenAI says ChatGPT outperforms doctors, but I’m not buying itOpenAI touts its HealthBench evaluations to imply that ChatGPT is good at giving medical advice. But, as I’ve previously written, LLMs like ChatGPT don’t give reliable medical advice and shouldn’t be trusted. Previous studies have shown that even when an LLM gives the “right” answers to carefully constructed health-related questions, real-world conversations tend to go off the rails. HealthBench seems to suffer from the same limitations. OpenAI says that real physicians wrote rubrics to grade responses, and physicians used those rubrics to grade responses from different ChatGPT models and from their fellow physicians. Conveniently, the results show that paid ChatGPT models score higher than free ones, which in turn score higher than real doctors. But that doesn’t mean a chatbot is actually trustworthy. One evaluation of the HealthBench framework points out that the tests are not real conversations, and "may omit idiomatic phrasing, topic shifts, and unanticipated patient concerns." Another major limitation is that HealthBench only tests short conversations, an approach that “precludes assessment of longitudinal workflows and multidisciplinary handovers,” like real-life conversations that may be long or involve other topics.












