Back

Accuracy of a machine learning method based on structural and locational information from AlphaFold2 for predicting the pathogenicity of TARDBP and FUS gene variants in ALS

Hatano, Y.; Ishihara, T.; Onodera, O.

2022-07-07 bioinformatics
10.1101/2022.07.07.499092 bioRxiv
Show abstract

BackgroundIn the sporadic form of amyotrophic lateral sclerosis (ALS), the pathogenicity of rare variants in the causative genes characterizing the familial form remains largely unknown. To predict the pathogenicity of such variants, in silico analysis is commonly used. In some cases of ALS, the gene mutations are concentrated in specific regions, and the resulting alterations in protein structure are thought to significantly affect pathogenicity. However, existing methods have not taken this issue into account. To address this, we have developed a technique termed MOVA (method for evaluating the pathogenicity of missense variants using AlphaFold2), which applies positional information for structural variants predicted by AlphaFold2. Here we examined the utility of MOVA for analysis of several causative genes of ALS. MethodsWe analyzed variants of six ALS-related genes (TARDBP, FUS, SETX, TBK1, OPTN, and SOD1) and classified them as pathogenic or neutral. For each gene, the features of the variants, including their positions in the 3D structure predicted by AlphaFold2, were entered into a random forest algorithm and evaluated by leave-one-out cross-validation. We compared how accurately MOVA was able to classify the pathogenic and neutral mutation variants. ResultsMOVA yielded useful results (AUC [≥]0.70 for 3 (TARDBP 0.755, FUS 0.844, and SOD1 0.787) of the 6 genes) and was particularly useful for genes where pathogenic mutations were concentrated at specific sites (TARDBP, FUS). ConclusionsMOVA is useful for predicting the virulence of rare variants of ALS-causing genes in which mutations are concentrated at specific structural sites.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.