Back

Mutation Pathogenicity Prediction by a Biology Based Explainable AI Multi-Modal Algorithm

Kellerman, R.; Nayshool, O.; Barel, O.; Paz, S.; Amariglio, N.; Klang, E.; Rechavi, G.

2024-06-05 genetic and genomic medicine
10.1101/2024.06.05.24308476 medRxiv
Show abstract

Most known pathogenic mutations occur in protein-coding regions of DNA and change the way proteins are made. Deciphering the protein structure therefore provides great insight into the molecular mechanisms underlying biological functions in human disease. While there have recently been major advances in the artificial intelligence-based prediction of protein structure, the determination of the biological and clinical relevance of specific mutations is not yet up to clinical standards. This challenge is of utmost medical importance when decisions, as critical as suggesting termination of pregnancy or recommending cancer-directed rational drugs, depend on the accuracy of prediction of the effect of the specific mutation. Currently, available tools are aiming to characterize the effect of a mutation on the functionality of the protein according to biochemical criteria, independent of the biological context. A specific change in protein structure can result either in loss of function (LOF) or gain-of-function (GOF) and the ability to identify the directionality of effect needs to be taken into consideration when interpreting the biological outcome of the mutation. Here we describe Triple-modalities Variant Interpretation and Analysis (TriVIAI), a tool incorporating three complementing modalities for improved prediction of missense mutations pathogenicity: protein language model (pLM), graph neural network (GNN) and a tabular model incorporating physical properties from the protein structure. The TriVIAl ensembles predictions compare favorably with the existing tools across various metrics, achieving an AUC-ROC of 0.887, a precision-recall curve (PRC) score of 0.68, and a Brier score of 0.16. The TriVIAI ensemble is also endowed with two major advantages compared to other available tools. The first is the incorporation of biological insights which allow to differentiate between GOF mutations that tend to cluster in specific hotspots and affect structure in a specific functional way versus LOF mutations that are usually dispersed and can cripple the protein in a variety of different ways. Importantly, the advantage over other available tools is more noticeable with GOF mutations as their effect on the protein structure is less disruptive and can be misinterpreted by current variant prioritization strategies. Until now available AI-based pathogenicity predicting algorithms were a black box for the users. The second significant advantage of TriVIAI is the explainability of the ensemble which contrasts the other available AI-based pathogenicity predicting algorithms which constitute a black box for the users. This explainability feature is of major importance considering the clinical responsibility of the medical decision-makers using AI-based pathogenicity predictors.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.1%
12.6%
2
Scientific Reports
3612 papers in training set
Top 2%
12.4%
3
Bioinformatics
1204 papers in training set
Top 4%
6.1%
4
Nature Communications
5641 papers in training set
Top 28%
5.3%
5
Frontiers in Bioinformatics
49 papers in training set
Top 0.1%
5.3%
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.8%
4.7%
7
Bioinformatics Advances
203 papers in training set
Top 1%
4.2%
50% of probability mass above
8
Human Genetics and Genomics Advances
84 papers in training set
Top 0.5%
3.9%
9
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
3.9%
10
International Journal of Molecular Sciences
494 papers in training set
Top 4%
3.1%
11
PLOS Computational Biology
1863 papers in training set
Top 10%
3.1%
12
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.7%
13
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.7%
14
Communications Biology
993 papers in training set
Top 11%
2.1%
15
Nature Computational Science
55 papers in training set
Top 0.4%
2.1%
16
Frontiers in Molecular Biosciences
102 papers in training set
Top 0.5%
2.1%
17
iScience
1154 papers in training set
Top 18%
1.7%
18
Nature Machine Intelligence
70 papers in training set
Top 2%
1.7%
19
Molecular Systems Biology
162 papers in training set
Top 2%
1.6%
20
Protein Science
246 papers in training set
Top 2%
1.5%
21
Heliyon
152 papers in training set
Top 6%
1.0%
22
Genetics in Medicine
78 papers in training set
Top 0.9%
0.9%
23
Journal of Chemical Theory and Computation
140 papers in training set
Top 1%
0.8%
24
Computers in Biology and Medicine
128 papers in training set
Top 5%
0.8%
25
eLife
5828 papers in training set
Top 66%
0.8%
26
Patterns
78 papers in training set
Top 3%
0.8%
27
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 1%
0.8%
28
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 1%
0.6%
29
Genome Biology
637 papers in training set
Top 10%
0.6%