A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences
Motta, J. A.; Motta, M. d. M.; Fernandez, C.
Show abstract
In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automatic Detection and Extraction of Key Resources from Tables in Biomedical Papers 93%
- Expanding a Database-derived Biomedical Knowledge Graph via Multi-relation Extraction from Biomedical Abstracts 92%
- A compact encoding of the genome suitable for machine learning prediction of traits and genetic risk scores. 92%
Similar papers in this journal
- Variomes: a high recall search engine to support the curation of genomic variants 94%
- 3Cnet: Pathogenicity prediction of human variants using knowledge transfer with deep recurrent neural networks 93%
- E-SNPs&GO: Embedding of protein sequence and function improves the annotation of pathogenic variants. 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.