MetaSplice: an ensemble pathogenicity predictor for intronic splice variants
Liu, X.
Show abstract
Splice-altering variants cause an estimated 15-30% of genetic diseases, yet computational tools lose accuracy outside the canonical GT-AG dinucleotides, leaving intronic variants of uncertain significance (VUS) hard to interpret. Here we present MetaSplice, a 53-feature gradient-boosted ensemble integrating deep-learning splice predictions (SpliceTransformer, Pangolin), evolutionary and gene-level constraint, and splicing-regulatory motifs to score single-nucleotide variants across seven non-exonic splice regions: canonical donor and acceptor sites, donor and acceptor regions, the polypyrimidine tract (PPT), branch point, and deep intronic positions. Trained on 381,226 intronic ClinVar SNVs (35,506 pathogenic), MetaSplice achieved an area under the precision-recall curve (auPRC) of 0.995 in five-fold gene-grouped cross-validation. On a temporally held-out ClinVar test set (107,933 variants), it reached auPRC 0.987 (95% CI 0.985-0.989), outperforming SpliceAI (0.950), CADD v1.7 (0.915), SPIDEX (0.767) and S-CAP (0.253), with the largest gains in the PPT, branch-point and deep intronic regions. As a pathogenicity predictor trained on clinical significance, MetaSplice complements mechanism-specific splice-effect tools. Applied to 47,272 ClinVar splice-region VUS, it nominated 22% for functional follow-up. MetaSplice is freely available as a Docker image.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Misexpression of inactive genes in whole blood is associated with nearby rare structural variants 95%
- Impact of genome build on RNA-seq interpretation and diagnostics 95%
- Advanced variant classification framework reduces the false positive rate of predicted loss of function (pLoF) variants in population sequencing data 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.