A new medical AI study found the same flaw in OpenEvidence, OpenAI, Anthropic, and Doximity

2 weeks ago 8

Want Your Business Featured Here?

Get instant exposure to our readers

Chat on WhatsApp

Medical AI Systems Fail to Deliver on Promises of Perfection

A recent study has exposed a concerning flaw in the performance of several leading medical AI systems, including OpenEvidence, OpenAI, Anthropic, and Doximity. The findings of this study should be a wake-up call for the medical community, as they highlight the limitations and potential dangers of relying on these systems for critical healthcare decisions.

Background & Context

The medical AI boom has been gaining momentum in recent years, with numerous startups and established companies investing heavily in the development of AI-powered healthcare tools. One such tool is OpenEvidence, a free, ad-supported AI search engine for doctors that pulls answers straight from peer-reviewed medical journals at the point of care and labels how strong the evidence is. This platform has become a poster child for the medical AI boom, with a valuation of $12 billion as of January this year.

Meanwhile, Doximity has taken a different approach, focusing on the enterprise market with its AI assistant, Ask. This platform helps doctors summarize patient notes, check drug interactions, and draft documentation, and has been sold to over 150 health systems. Doximity's revenue has been growing steadily, reaching $145.4 million in the spring, a 5% increase year-over-year.

Key Details

A recent independent benchmark called NOHARM put both Doximity and OpenEvidence's AI tools to the test, alongside OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5. The study ran 1,100 real clinical cases through each model and collected roughly 13,000 physician annotations to score for patient harm. The results were surprising, with Doximity's Ask coming out on top, but OpenEvidence contesting the accuracy of the score.

However, the finding of the study is not about who "won," but rather that even the best-performing models still miss things. Across every AI system tested, 76.6% of harmful errors were omissions, meaning the AI left something out, not that it stated something factually wrong. This distinction is significant, as experts warn that errors of omission need to be brought as close to zero as possible.

Eric Topol, a cardiologist and scientist at Scripps Research, has spent his career studying diagnostic error. He noted that today's models maintain an "illusion of readiness" that has followed medical AI even as it improves. Topol added that doctors equipped with AI still give better care compared to those without, but emphasized that this is not a justification for relying solely on AI.

What Experts Say

The regulatory backdrop makes NOHARM's timing pointed. The FDA loosened its stance on AI-powered clinical decision-support tools this January, giving them more room to operate as long as they are designed to provide "sufficient" information to support informed decision-making. However, the study highlights the need for stricter regulations to ensure that AI systems are held to high standards of accuracy and safety.

OpenEvidence's CEO, Daniel Nadler, responded to the study's findings by stating that the NOHARM study's methodology does not allow for "re-tests" and that the basic table stakes requirement in medicine for even the flimsiest medical conclusions is peer review. However, experts argue that this is not a valid excuse for the lack of accuracy in AI systems.

Key Takeaways

  • 76.6% of harmful errors in AI systems were omissions, not factual errors.
  • Even the best-performing models still miss things, highlighting the limitations of AI in healthcare.
  • The study's findings emphasize the need for stricter regulations to ensure the safety and accuracy of AI systems in healthcare.
  • AI systems are not a replacement for human judgment and expertise, but rather a tool to support healthcare decisions.

What This Means For You

The study's findings have significant implications for patients and healthcare providers alike. As AI systems become increasingly integrated into healthcare, it is essential to recognize the limitations and potential dangers of relying solely on these systems. Patients should be aware of the potential risks of relying on AI for critical healthcare decisions and should advocate for human oversight and judgment in these cases.

Healthcare providers should also be aware of the study's findings and take steps to ensure that AI systems are used in a way that complements human judgment and expertise. This includes implementing robust quality control measures, providing ongoing training and education for healthcare providers, and ensuring that AI systems are designed with safety and accuracy in mind.

In conclusion, the study's findings serve as a wake-up call for the medical community, highlighting the need for greater caution and oversight in the development and deployment of AI systems in healthcare. By recognizing the limitations and potential dangers of AI, we can work towards creating a safer and more effective healthcare system for all.

Read Entire Article
Chatroom