Exploration of deep-learning based classification with human SNP image graphs
Chen, C.-H.; Tung, K.-F.; Lin, W.-c.
Show abstract
BackgroundWith the advancement of NGS platform, large numbers of human variations and SNPs are discovered in human genomes. It is essential to utilize these massive nucleotide variations for the discovery of disease genes and human phenotypic traits. There are new challenges in utilizing such large numbers of nucleotide variants for polygenic disease studies. In recent years, deep-learning based machine learning approaches have achieved great successes in many areas, especially image classifications. In this preliminary study, we are exploring the deep convolutional neural network algorithm in genome-wide SNP images for the classification of human populations. ResultsWe have processed the SNP information from more than 2,500 samples of 1000 genome project. Five major human races were used for classification categories. We first generated SNP image graphs of chromosome 22, which contained about one million SNPs. By using the residual network (ResNet 50) pipeline in CNN algorithm, we have successfully obtained classification models to classify the validation dataset. F1 scores of the trained CNN models are 95 to 99%, and validation with additional separate 150 samples indicates a 95.8% accuracy of the CNN model. Misclassification was often observed between the American and European categories, which could attribute to the ancestral origins. We further attempted to use SNP image graphs in reduced color representations or images generated by spiral shapes, which also provided good prediction accuracy. We then tried to use the SNP image graphs from chromosome 20, almost all CNN models failed to classify the human race category successfully, except the African samples. ConclusionsWe have developed a human race prediction model with deep convolutional neural network. It is feasible to use the SNP image graph for the classification of individual genomes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of 12 cancer types through genome deep learning 95%
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 94%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 94%
Similar papers in this journal
- Comparative analysis of novel MGISEQ-2000 sequencing platform vs Illumina HiSeq 2500 for whole-genome sequencing 94%
- Whole genome sequencing revealed genetic diversity and selection of Guangxi indigenous chickens 94%
- Classification of early and late stage Liver Hepatocellular Carcinoma patients from their genomics and epigenomics profiles. 94%
Similar papers in this journal
- GMQN: A reference-based method for correcting batch effects as well as probes bias in HumanMethylation BeadChip 94%
- Identification of Platform-Independent Diagnostic Biomarker Panel for Hepatocellular Carcinoma using Large-scale Transcriptomics Data 94%
- Analysis of Pan-Omics Data in Human Interactome Network (APODHIN) 94%
Similar papers in this journal
- Integrated ACMG approved genes and ICD codes for the translational research and precision medicine 95%
- CATA: a comprehensive chromatin accessibility database for cancer 94%
- MicroRNA childhood Cancer Catalog (M3Cs): A Resource for Translational Bioinformatics Toward Health Informatics in Pediatric Cancer. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.