Deep learning perturbation models can outperform baselines on calibrated metrics.
- Open access
Deep learning genetic perturbation models outperform uninformative baselines when evaluated with calibrated metrics across multiple datasets.
- Why it matters: Accurate benchmarking of these models is crucial for advancing genetic research, but current metrics often misrepresent true model performance, hindering progress.
- What they did: The authors developed a positive control baseline and a metric calibration framework, applying it to 14 datasets and 18 metrics to assess model performance.
- The result: They found that properly calibrated metrics reveal deep-learning models significantly outperform uninformative baselines, enabling more reliable evaluation and development of genetic perturbation models.