Back

Identification of single nucleotide genetic polymorphism sites using machine learning methods

Yatskou, M. M.; Smolyakova, E. V.; Skakun, V. V.; Grinev, V. V.

2023-10-21 bioinformatics
10.1101/2023.10.19.563060 bioRxiv
Show abstract

The paper presents an algorithm for simulation modelling of nucleotide variations in the genomic DNA molecule. To identify single nucleotide genetic polymorphisms, it is proposed to use machine learning methods trained on simulated data. A comparative analysis of the effective classical and machine learning algorithms for identifying single nucleotide polymorphisms was performed on simulated data. The most optimal method for identifying single nucleotide genetic polymorphisms in DNA molecules at various experimental noise levels is the machine learning algorithm CART.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.