Errors, Hallucinations, and Clinical Impact of General-Purpose Multimodal Large Language Models in Histopathology
- Open access
General-purpose multimodal large language models in histopathology exhibit errors and hallucinations in over 80% of outputs, impacting diagnostic safety and reliability.
- Why it matters: Understanding how these models fail is crucial because their widespread use in diagnostic pathology could lead to clinical misjudgments, yet prior studies mainly focused on accuracy without examining failure modes.
- What they did: The study evaluated four LLMs across 153 cases from 20 organs, analyzing 612 outputs for diagnostic correctness, errors, hallucinations, and clinical impact without model training or fine-tuning.
- The result: Errors and hallucinations were prevalent regardless of correctness, with a significant link between hallucination burden and clinical impact, highlighting the need for comprehensive evaluation of model reliability.