Interpretable distillation reveals that deep learning splicing models suffer from pervasive confounders and blind spots.
Deep learning splicing models rely on confounders and blind spots, resulting in systematic errors and poor performance on non-reference sequences.
- Why it matters: Understanding the predictive mechanisms of these models is essential for accurate gene regulation insights and genetic variation interpretation, but their interpretability is limited.
- What they did: A framework using interpretable distillation was developed to analyze model predictions, revealing that models recognize exons through simple motif combinations and exploit unrelated confounders.
- The result: Findings highlight fundamental limitations in current models, emphasizing the need to address confounders and incorporate RNA structural effects to improve prediction accuracy.