Back

Improved KCNQ2 gene missense variant interpretation with artificial intelligence

Saez-Matia, A.; Muguruza-Montero, A.; M-Alicante, S.; Nunez, E.; Gamis, R.; Ballesteros, O. R.; Garcia-Ibarlueza, M.; Fons, C.; Leonardo, A.; Bergara, A.; Villarroel, A.

2022-10-20 pathology
10.1101/2022.10.20.513007 bioRxiv
Show abstract

Advances in DNA sequencing technologies have revolutionized rare disease diagnosis, resulting in an increasing volume of available genomic data. Despite this wealth of information and improved procedures to combine data from various sources, identifying the pathogenic causal variants and distinguishing between severe and benign variants remains a key challenge. Mutations in the Kv7.2 voltage-gated potassium channel gene (KCNQ2) have been linked to different subtypes of epilepsies, such as benign familial neonatal epilepsy (BFNE) and epileptic encephalopathy (EE). To date, there is a wide variety of genome-wide computational tools aiming at predicting the pathogenicity of variants. However, previous reports suggest that these genome-wide tools have limited applicability to the KCNQ2 gene related diseases due to overestimation of deleterious mutations and failure to correctly identify benign variants, being, therefore, of limited use in clinical practice. In this work, we found that combining readily available features, such as AlphaFold structural information, Missense Tolerance Ratio (MTR) and other commonly used protein descriptors, provides foundations to build reliable gene-specific machine learning ensemble models. Here, we present a transferable methodology able to accurately predict the pathogenicity of KCNQ2 missense variants with unprecedented sensitivity and specificity scores above 90%.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.