Back

Circumpolar peoples and their languages: lexical and genomic data suggest ancient Chukotko-Kamchatkan-Nivkh and Yukaghir-Samoyedic connections

Starostin, G.; Altınısık, N. E.; Zhivlov, M.; Changmai, P.; Flegontova, O.; Spirin, S. A.; Zavgorodnii, A.; Flegontov, P.; Kassian, A. S.

2021-02-28 genomics
10.1101/2021.02.27.433193 bioRxiv
Show abstract

Relationships between universally recognized language families represent a hotly debated topic in historical linguistics, and the same is true for correlation between signals of genetic and linguistic relatedness. We developed a weighted permutation test and applied it on basic vocabularies for 31 pairs of languages and reconstructed proto-languages to show that three groups of circumpolar language families in the Northern Hemisphere show evidence of relationship though borrowing in the basic vocabulary or common descent: [Chukotko-Kamchatkan and Nivkh]; [Yukaghir and Samoyedic]; [Yeniseian, Na-Dene, and Burushaski]. The former two pairs showed the most significant signals of language relationship, and the same pairs demonstrated parallel signals of genetic relationship implying common descent or substantial gene flows. For finding the genetic signals we used genome-wide genetic data for present-day groups and a bootstrapping model comparison approach for admixture graphs or, alternatively, haplotype sharing statistics. Our findings further support some hypotheses on long-distance language relationship put forward based on the linguistic methods but lacking universal acceptance. Significance statementIndigenous people inhabiting polar and sub-polar regions in the Northern Hemisphere speak diverse languages belonging to at least seven language families which are traditionally thought of as unrelated entities. We developed a weighted permutation test and applied it to basic vocabularies of a number of languages and reconstructed proto-languages to show that at least three groups of circumpolar language families show evidence of relationship though either borrowing in the basic vocabulary or common descent: Chukotko-Kamchatkan and Nivkh; Yukaghir and Samoyedic; Yeniseian, Na-Dene, and Burushaski. The former two pairs showed the most significant signals of language relationship, and the same pairs demonstrated parallel signals of genetic relationship implying common descent or substantial gene flows.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.