Back

Inland-coastal bifurcation of southern East Asians revealed by Hmong-Mien genomic history

Xia, Z.-Y.; Yan, S.; Wang, C.-C.; Zheng, H.-X.; Zhang, F.; Liu, Y.-C.; Yu, G.; Yu, B.-X.; Shu, L.-L.; Jin, L.

2019-08-09 genetics
10.1101/730903 bioRxiv
Show abstract

The early history of the Hmong-Mien language family and its speakers is elusive. A good variety of Hmong-Mien-speaking groups distribute in Central China. Here, we report 903 high-resolution Y-chromosomal, 624 full-sequencing mitochondrial, and 415 autosomal samples from 20 populations in Central China, mainly Hunan Province. We identify an autosomal component which is commonly seen in all the Hmong-Mien-speaking populations, with nearly unmixed composition in Pahng. In contrast, Hmong and Mien respectively demonstrate additional genomic affinity to Tibeto-Burman and Kra-Dai speakers. We also discover two prevalent uniparental lineages of Hmong-Mien speakers. Y-chromosomal haplogroup O2a2a1b1a1b-N5 diverged [~]2,330 years before present (BP), approximately coinciding with the estimated time of Proto-Hmong-Mien ([~]2,500 BP), whereas mitochondrial haplogroup B5a1c1a significantly correlates with Pahng and Mien. All the evidence indicates a founding population substantially contributing to present-day Hmong-Mien speakers. Consistent with the two distinct routes of agricultural expansion from southern China, this Hmong-Mien founding ancestry is phylogenetically closer to the founding ancestry of Neolithic Mainland Southeast Asians and present-day isolated Austroasiatic-speaking populations than Austronesians. The spatial and temporal distribution of the southern East Asian lineage is also compatible with the scenario of out-of-southern-China farming dispersal. Thus, our finding reveals an inland-coastal genetic discrepancy related to the farming pioneers in southern China and supports an inland southern China origin of an ancestral meta-population contributing to both Hmong-Mien and Austroasiatic speakers.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.