General-purpose large language models outperform specialized clinical AI tools on medical benchmarks.
- Open access
- 32 cites
Large language models like GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 outperform specialized clinical AI tools across multiple medical benchmarks, including real-world queries.
- Why it matters: The widespread adoption of clinical AI tools occurs without sufficient independent validation, risking ineffective or unsafe medical decision support.
- What they did: Researchers evaluated two clinical AI tools against three frontier LLMs using 1,500 questions from medical knowledge tests, clinician alignment assessments, and live clinical queries, with clinician review of outputs.
- The result: Frontier LLMs consistently outperformed clinical AI tools, emphasizing the importance of independent, real-world testing of AI systems before clinical deployment to ensure safety and effectiveness.