APPRIS principal isoforms and MANE Select transcripts in clinical variant interpretation
Pozo, F.; Rodriguez, J. M.; Vazquez, J.; Tress, M. L.
Show abstract
Most coding genes are able to generate multiple alternatively spliced transcripts. Determining which of these transcript variants produces the main protein isoform, and which of a genes multiple splice variants are functionally important, is crucial in comparative genomics and essential for clinical variant interpretation. Here we show that the principal isoforms chosen by APPRIS and the MANE Select variants provide the best approximations of the main cellular protein isoforms. Principal isoforms are predicted from conservation and from protein features, and MANE transcripts are chosen from the consensus between teams of expert manual curators. APPRIS principal isoforms coincide in over 94% of coding genes with MANE Select transcripts and the two methods are particularly discriminating when they agree on the main splice variant. Where the two methods agree, the splice variants coincide with the main isoform detected in proteomics experiments in 98.2% of genes with multiple protein isoforms. We also find that almost all ClinVar pathogenic mutations map to MANE Select or APPRIS principal isoforms. Where APPRIS and MANE agree on the main isoform, 99.93% of validated pathogenic variants map to principal rather than alternative exons. MANE Plus Clinical transcripts cover most validated pathogenic mutations in alternative coding exons. TRIFID functional importance scores are particularly useful for distinguishing clinically important alternative isoforms: the highest scoring TRIFID isoforms are more than 300 times more likely to have validated pathogenic mutations. We find that APPRIS, MANE and TRIFID are important for determining the biological relevance of splice isoforms and should be an essential part of clinical variant interpretation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- REVEL is better at predicting pathogenicity of loss-of-function than gain-of-function variants 94%
- Splicing impact of deep exonic missense variants in CAPN3 explored systematically by minigene functional assay 93%
- Using single molecule Molecular Inversion Probes as a cost-effective, high-throughput sequencing approach to target all genes and loci associated with macular diseases 92%
Similar papers in this journal
- Cataloging the potential functional diversity of Cacna1e splice variants using long-read sequencing 93%
- Mutational Constraint Analysis Workflow for Overlapping Short Open Reading Frames and Genomic Neighbours 92%
- Expanding the Potential Genes of Inborn Errors of Immunity through Protein Interactions 92%
Similar papers in this journal
- Curated Multiple Sequence Alignment for the Adenomatous Polyposis Coli (APC) Gene and Accuracy of In Silico Pathogenicity Predictions 95%
- Towards development of a statistical framework to evaluate myotonic dystrophy type 1 mRNA biomarkers in the context of a clinical trial 94%
- Proteome-scale prediction of molecular mechanisms underlying dominant genetic diseases 94%
Similar papers in this journal
- CHESS 3: an improved, comprehensive catalog of human genes and transcripts based on large-scale expression data, phylogenetic analysis, and protein structure 96%
- Functional enrichment of alternative splicing events with NEASE reveals insights into tissue identity and diseases 95%
- Variant effect predictor correlation with functional assays is reflective of clinical classification performance 94%
Similar papers in this journal
- Mutation severity spectrum of rare alleles in the human genome is predictive of disease type 95%
- eVIP2: Expression-based variant impact phenotyping to predict the function of gene variants 94%
- Large scale analyses of genotype-phenotype relationships of glycine decarboxylase mutations and neurological disease severity. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.