Cluster Buster: A Machine Learning Algorithm for Genotyping SNPs from Raw Data
Martin, J. L.; Kuznetsov, N.; Levine, K.; Koretsky, M.; Hong, S.; Vitale, D.; Nalls, M. A.
Show abstract
Genotyping single nucleotide polymorphisms (SNPs) is fundamental to disease research, as researchers seek to establish links between genetic variation and disease. Although significant advances in genome technology have been made with the development of bead-based SNP genotyping and Genome Studio software, some SNPs still fail to be genotyped, resulting in "no-calls" that impede downstream analyses. To recover these genotypes, we introduce Cluster Buster, a genotyping neural network and visual inspection system designed to improve the quality of neurodegenerative disease (NDD) research. Concordance analysis with whole genome sequencing (WGS) and imputed genotypes validated the reliability of predicted genotypes, with dozens of high-performing SNPs across LRRK2, APOE, and GBA loci achieving at least 90% concordance per SNP location. Further analysis of concordance between Genome Studio genotypes and imputed and WGS genotypes revealed discrepancies between the genotyping technologies, highlighting the need for selective application of Cluster Buster on SNP locations based on concordance rates. Cluster Busters implementation significantly reduces manual labor for recovering no-call SNPs, refining genotype quality for the Global Parkinsons Genetics Program (GP2). This system facilitates better imputation and GWAS outcomes, ultimately contributing to a deeper understanding of genetic factors in NDDs.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 93%
- Comparing low-pass sequencing and genotyping for trait mapping in pharmacogenetics 93%
- Flexible, Production-Scale, Human Whole Genome Sequencing On A Benchtop Sequencer 92%
Similar papers in this journal
- MADloy: Robust detection of mosaic loss of chromosome Y from genotype-array-intensity data 93%
- CNValidatron: Accurate And Efficient Validation of PennCNV Calls Using Computer Vision 92%
- Rare Copy Number Variant analysis in case-control studies using SNP Array Data: a scalable and automated data analysis pipeline 92%
Similar papers in this journal
- Impute.me: an open source, non-profit tool for using data from DTC genetic testing to calculate and interpret polygenic risk scores. 93%
- MetaPhat: Detecting and decomposing multivariate associations from univariate genome-wide association statistics 93%
- MicroHapDB: a portable and extensible database of all published microhaplotype marker and frequency data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.