Related Experiment Videos
Interpretable distillation reveals that deep learning splicing models suffer from pervasive confounders and blind
Simon Liu1, Wenjing Zhang1, Oded Regev2
1New York University, New York, NY, 10012, USA.
Background:
Predicting RNA splicing from genomic sequence is a crucial task for understanding gene regulation and interpreting genetic variation. Recent deep learning advancements have led to splicing prediction algorithms that achieve state-of-the-art performance compared to earlier models. However, owing to the limited interpretability of deep learning models, the predictive mechanisms of current splicing models remain poorly understood.
Results:
Here we develop a framework to explain model prediction logic using interpretable distillation. Applying our framework, we find that RNA splicing prediction models suffer from pervasive confounders and blind spots, leading to poor performance on non-reference sequences. We find that splicing models recognize exons through surprisingly simple additive combinations of sequence motifs, including known splicing regulatory elements. Critically, our analysis also reveals that splicing models exploit genomic confounders unrelated to splicing and fail to adequately capture the effects of RNA structure, leading to systematic prediction errors.
Conclusions:
Our findings illuminate fundamental limitations of training models on genomic sequences and suggest ways to overcome them.