Vishwanath, K., et al. (2026).
Nature Medicine.
Abstract
Specialized clinical artificial intelligence (AI) tools are entering medical practice despite scarce independent evaluation. We quantitatively evaluate two clinical AI tools, OpenEvidence and UpToDate Expert AI, built on large language models (LLMs) against three frontier LLMs: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. Our evaluation has three stages: (1) 500 MedQA questions testing medical knowledge, (2) 500 HealthBench items measuring alignment with clinicians and (3) the real clinical queries (RCQ) benchmark, built from 100 de-identified queries from physicians to a general-purpose language model in a live clinical environment. For the RCQ benchmark, 12 US clinicians performed randomized, blinded review of model outputs, producing 1,800 model–question annotations. Frontier LLMs outperformed clinical AI tools in all three evaluations. Clinical AI tools performed comparably to auto-enabled Google Search AI Overview on the RCQ. These findings highlight the need for independent, real-world evaluation of AI tools before they enter clinical settings.
Here are some thoughts:
This 2026 Nature Medicine study asked a simple question: are the special AI tools being sold to doctors actually better than the regular AI chatbots anyone can use? The researchers tested two clinical tools (OpenEvidence and UpToDate Expert AI) against three general-purpose models (GPT-5.2, Gemini, and Claude) on medical exam questions, expert-alignment tests, and real questions that doctors had asked during patient care, with twelve doctors blindly scoring the answers.
The answer was clear: the general-purpose chatbots beat the specialized medical tools on every test. In fact, the medical tools did no better than the free AI summary that shows up at the top of a Google search. The specialized tools mostly struggled with being clear and complete rather than getting facts wrong, and none of the tools were notably more dangerous than the others.
The takeaway is that paying for a fancy, doctor-branded AI tool may not get you better results than a regular chatbot, which matters a lot given that one of these companies was recently valued at billions of dollars. A few caveats: the study was small, couldn't measure speed or quality of sources, and one author consults for Google, whose model won. The authors think the real future may be hospitals building their own AI on their own data, rather than buying these off-the-shelf medical tools.








