Spurious model comparisons are widespread in biomedical artificial intelligence
- Open access
- 1 cites
Spurious model comparisons are prevalent in biomedical AI, with nearly two-thirds of studies using invalid tests that inflate false-positive rates.
- Why it matters: This widespread issue undermines the reliability of performance claims in biomedical research, risking false conclusions and impeding scientific progress.
- What they did: Analyzing 184 studies across 30 fields, the authors identified that 97% used invalid cross-validation tests, and simulations showed false-positive rates nearly reach 100% with repeated testing.
- The result: Introducing SHARP, a redesigned cross-validation method, offers a practical solution to improve the validity of model comparisons, enabling more trustworthy biomedical AI research.