Back

A Recurrent Neural Network Based Method for Genotype Imputation on Phased Genotype Data

Kojima, K.; Tadaka, S.; Katsuoka, F.; Tamiya, G.; Yamamoto, M.; Kinoshita, K.

2019-10-30 bioinformatics
10.1101/821504 bioRxiv
Show abstract

Genotype imputation estimates genotypes of unobserved variants from genotype data of other observed variants, and such estimation is enabled using haplotype data of a large number of other individuals. Although existing imputation methods explicitly use haplotype data, the accessibility of haplotype data is often limited because the agreement is necessary from donors of genome data. We propose a new imputation method that uses bidirectional recurrent neural network, and haplotype data of a large number of individuals are encoded as its model parameters through the training step, which can be shared publicly due to the difficulty in restoring genotype data at the individual-level. In the performance evaluation using the phased genotype data in the 1000 Genomes Project, the imputation accuracy of the proposed method in R2 is comparative with existing methods for variants with MAF [&ge;] 0.05 and is slightly worse than those of the existing methods for variants with MAF < 0.05. In a scenario of limited availability of haplotype data to the existing methods, the accuracy of the proposed method is higher than those of the existing methods at least for variants with MAF [&ge;] 0.005. Python code of our implementation for imputation is available at https://github.com/kanamekojima/rnnimp/.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.