Aberrant splicing prediction across human tissues
Celik, M. H.; Wagner, N.; Hoelzlwimmer, F.; Yepez, V.; Mertes, C.; Prokisch, H.; Gagneur, J.
Show abstract
Aberrant splicing is a major cause of genetic disorders but its direct detection in transcriptomes is limited to clinically accessible tissues such as skin or body fluids. While DNA-based machine learning models allow prioritizing rare variants for affecting splicing, their performance on predicting tissue-specific aberrant splicing remains unassessed. Here, we generated the first aberrant splicing benchmark dataset, spanning over 8.8 million rare variants in 49 human tissues. At 20% recall, state-of-the-art DNA-based models cap at 10% precision. By mapping and quantifying tissue-specific splice site usage transcriptome-wide and modeling isoform competition, we increased precision by three-fold at the same recall. Integrating RNA-sequencing data of clinically accessible tissues brought precision to 60%. These results, replicated in two independent cohorts, substantially contribute to non-coding loss-of-function variant identification and to genetic diagnostics design and analytics.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- The SpliZ generalizes "Percent Spliced In" to reveal regulated splicing at single-cell resolution 97%
- Systematic assessment of long-read RNA-seq methods for transcript identification and quantification 95%
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 95%
Similar papers in this journal
- Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance 96%
- Systematic assessment of regulatory effects of human disease variants in pluripotent cells 96%
- Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.