An Atlas of Indian Genetic Diversity
Subramanian, K.; Bhattacharyya, C.; Machha, P.; Mukherjee, A.; Tripathi, D.; Chakraborty, S.; Majumdar, S. S.; Sengupta, S.; Singh, P.; More, V.; Bari, S.; MS, S.; Macwan, E.; Mondal, D.; Dey, A.; Afreen, H.; Jani, S. P.; Mukherjee, P.; Singh, N.; De, T.; Sharma, P.; Upilli, B.; Maitra, A.; Singh, K.; Sharma, P.; Sharma, N.; Raghav, S. K.; Prasad, P.; Soniya, E. V.; Jaleel, A.; Pillai, M. R.; Sathi, S. N.; Joshi, M.; Joshi, C.; Lahiri, M.; Dixit, S.; Shashidhara, L. S.; Kumar, N. S.; Lalhruaitluanga, H.; Nundanga, L.; Shivakumar, V.; Venkatasubramanian, G.; Rao, N. P.; Ganie, M. A.; Wani, I. A
Show abstract
India, the most populous country, remains significantly underrepresented in the global genomics landscape. Previous efforts to catalog Indian genetic diversity were limited in scale, scope, and representation. Here, we present the GenomeIndia dataset, comprising whole genome sequences of 9,768 healthy individuals from 83 populations spanning the ethnolinguistic and biogeographic spectrum of India. We identify 129.93 million high-confidence biallelic variants, 44.03 million of which are previously unreported in global databases. In contrast to large populations that show steady population growth and internal homogeneity, we observe low effective population sizes, significant genetic drift, and profound homozygosity in small tribal groups, likely shaped by antiquity, isolation, and endogamy. We report multiple population-specific pharmacogenomic and deleterious variants, necessitating the integration of local genetic architecture and the inclusion of underrepresented South Asian genomes in global reference resources. Finally, we highlight the limited transferability of Eurocentric polygenic scores to Indian populations, and present an imputation panel that outperforms existing resources for both rare and common variants. Together, our work fills a significant gap in the equity of global human genomics, and paves way for precision medicine strategies that will benefit a quarter of the world population.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genotyping sequence-resolved copy number variationusing pangenomes reveals paralog-specific global diversityand expression divergence of duplicated genes 97%
- Systematic assessment of regulatory effects of human disease variants in pluripotent cells 97%
- A combined polygenic score of 21,293 rare and 22 common variants significantly improves diabetes diagnosis based on hemoglobin A1C levels 96%
Similar papers in this journal
- Fine-scale population structure and demographic history of British Pakistanis 98%
- Whole-genome sequencing of 1,171 elderly admixed individuals from the largest Latin American metropolis (Sao Paulo, Brazil) 97%
- Quantifying the contribution of Neanderthal introgression to the heritability of complex traits 97%
Similar papers in this journal
- Characterizing substructure via mixture modeling in large-scale genetic summary statistics 97%
- Enrichment analyses identify shared associations for 25 quantitative traits in over 600,000 individuals from seven diverse ancestries 97%
- Brain eQTLs of European, African American, and Asian ancestry improve interpretation of schizophrenia GWAS 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.