Back

Whole Genome Sequencing-based Characterization of Human Genome Variation and Mutation Burden in Botswana

Thami, P. K.; Choga, W. T.; Mulisa, D. D.; Dandara, C.; Shevchenko, A. K.; Leteane, M. M.; Novitsky, V.; O'Brien, S. J.; Essex, M.; Gaseitsiwe, S.; Chimusa, E. R.

2020-12-15 genomics
10.1101/2020.12.15.422821 bioRxiv
Show abstract

The study of human genome variations can contribute towards understanding population diversity and the genetic aetiology of health-related traits. We sought to characterise human genomic variations of Botswana in order to assess diversity and elucidate mutation burden in the population using whole genome sequencing. Whole genome sequences of 390 unrelated individuals from Botswana were available for computational analysis. The sequences were mapped to the human reference genome GRCh38. Population joint variant calling was performed using Genome Analysis Tool Kit (GATK) and BCFTools. Variant characterisation was achieved by annotating the variants with a suite of databases in ANNOVAR and snpEFF. The genomic architecture of Botswana was delineated through principal component analysis, structure analysis and FST. We identified a total of 27.7 million unique variants. Variant prioritisation revealed 24 damaging variants with the most damaging variants being ACTRT2 rs3795263, HOXD12 rs200302685, ABCB5 rs111647033, ATP8B4 rs77004004 and ABCC12 rs113496237. We observed admixture of the Khoe-San, Niger-Congo and European ancestries in the population of Botswana, however population substructure was not observed. This exploration of whole genome sequences presents a comprehensive characterisation of human genomic variations in the population of Botswana and their potential in contributing to a deeper understanding of population diversity and health in Africa and the African diaspora.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.