Back

An Atlas of Indian Genetic Diversity

Subramanian, K.; Bhattacharyya, C.; Machha, P.; Mukherjee, A.; Tripathi, D.; Chakraborty, S.; Majumdar, S. S.; Sengupta, S.; Singh, P.; More, V.; Bari, S.; MS, S.; Macwan, E.; Mondal, D.; Dey, A.; Afreen, H.; Jani, S. P.; Mukherjee, P.; Singh, N.; De, T.; Sharma, P.; Upilli, B.; Maitra, A.; Singh, K.; Sharma, P.; Sharma, N.; Raghav, S. K.; Prasad, P.; Soniya, E. V.; Jaleel, A.; Pillai, M. R.; Sathi, S. N.; Joshi, M.; Joshi, C.; Lahiri, M.; Dixit, S.; Shashidhara, L. S.; Kumar, N. S.; Lalhruaitluanga, H.; Nundanga, L.; Shivakumar, V.; Venkatasubramanian, G.; Rao, N. P.; Ganie, M. A.; Wani, I. A

2026-03-24 genetic and genomic medicine
10.64898/2026.03.20.26348801 medRxiv
Show abstract

India, the most populous country, remains significantly underrepresented in the global genomics landscape. Previous efforts to catalog Indian genetic diversity were limited in scale, scope, and representation. Here, we present the GenomeIndia dataset, comprising whole genome sequences of 9,768 healthy individuals from 83 populations spanning the ethnolinguistic and biogeographic spectrum of India. We identify 129.93 million high-confidence biallelic variants, 44.03 million of which are previously unreported in global databases. In contrast to large populations that show steady population growth and internal homogeneity, we observe low effective population sizes, significant genetic drift, and profound homozygosity in small tribal groups, likely shaped by antiquity, isolation, and endogamy. We report multiple population-specific pharmacogenomic and deleterious variants, necessitating the integration of local genetic architecture and the inclusion of underrepresented South Asian genomes in global reference resources. Finally, we highlight the limited transferability of Eurocentric polygenic scores to Indian populations, and present an imputation panel that outperforms existing resources for both rare and common variants. Together, our work fills a significant gap in the equity of global human genomics, and paves way for precision medicine strategies that will benefit a quarter of the world population.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.