Study Finds AI Chatbots Provide Inaccurate Medical Advice
Researchers from two universities found that AI chatbots frequently provide inaccurate and incomplete medical information, highlighting a need for regulatory oversight.
A study published in the medical journal BMJ Open reveals that AI chatbots frequently provide inaccurate and incomplete medical information. Researchers from the University of Alberta and Loughborough University tested five different chatbots using 50 medical questions, finding that half of the responses were problematic.
Results showed that Grok had the highest rate of problematic responses at 58%, followed by ChatGPT at 52% and Meta AI at 50%. The AI models struggled most with topics regarding athletic performance, nutrition, and stem cells, while showing higher accuracy when addressing vaccines and cancer.
Researchers attributed these failures to hallucinations caused by incomplete or biased training data. They also noted a tendency toward sycophancy, where chatbots prioritize a user's existing beliefs over factual truth. Because these models infer statistical patterns instead of weighing evidence or reasoning, the study concludes there is an urgent need for professional training and regulatory oversight to protect public health.