Kaiser Permanente National Cross-Vendor Validation of Mammography Artificial Intelligence Computer-Aided Diagnosis Algorithms in a US-Representative Population
Algorithm D achieved the highest AUROC of 0.849 and sensitivity across all AI-positive rates in a large US screening cohort, outperforming other commercial and open-source models.
- Why it matters: Understanding the comparative performance of AI CAD algorithms is crucial for safe and effective integration into clinical mammography workflows, addressing the lack of large-scale independent validations.
- What they did: The study analyzed 786,124 mammograms from 2022-2023, applying three FDA-cleared commercial algorithms and one open-source model, comparing their AUROC, sensitivity, specificity, and false negatives at various AI-positive thresholds.
- The result: Findings show significant performance differences among algorithms, with the open-source model performing comparably to some commercial options, emphasizing the need for local validation and tailored threshold selection before clinical adoption.