Back

A hybrid approach combining a phylogenetic method and Approximate Bayesian Computation Random Forest for phylogenetic network inference: application to the rice domestication process in Asia

Rabier, C.-E.; Berry, V.; Glaszmann, J.-C.

2026-07-17 evolutionary biology
10.64898/2026.07.13.738130 bioRxiv
Show abstract

Asian rice is one of the best documented crops in terms of genetic diversity. The domestication process, that probably started 9000 years ago in China, remains difficult to infer since the main vertical signal is blurred by horizontal signals related to gene flow among cultivars and wild relatives. Consequently, a large number of hypotheses on the domestication process of rice have been published. Besides, most of the methods used to infer these scenarios do not model all the known biological phenomena at stake. Here, we present a methodological study based on a rich stochastic model, that incorporates introgression events, incomplete lineage sorting, and mutations that happen over time. The global evolutionary scenario is represented by a phylogenetic network. Furthermore, each locus scenario is modeled according to a locus tree through the Multispecies Network Coalescent. More importantly, for inferring the phylogenetic network, we propose a new hybrid approach combining a phylogenetic network method and a machine learning technique. In particular, our hybrid approach, named SO_SCPLOWNARFC_SCPLOW, benefits from advantages of a mathematical phylogenetic method, SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOW, and from the potential of a powerful machine learning classifier, i.e. Approximate Bayesian Computation Random Forest (ABC-RF). These two methods are complementary since SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOW reconstructs network accurately, whereas ABC-RF is able to handle a large amount of data. The originality is twofold. First, prior distributions required for ABC-RF are calibrated thanks to SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOWs estimates. Secondly, ABC-RF relies on summary statistics inspired by phylogenetic network literature. We show, on simulated data, that the SO_SCPLOWNARFC_SCPLOW hybrid approach enjoys very good performances. On rice real data, it infers a scenario with a unique domestication (that of Japonica), followed by three reticulation events involving early Japonica. It highlights two introgression events at the origin of Indica and cAus, and one admixture event responsible for the emergence of cBas. Author summaryToday, in genomics, there is a real need for methods able to infer phylogenetic networks. A phylogenetic network is a directed graph representing events like hybridization, introgression, and horizontal gene transfer. Understanding these complex biological phenomena, essential for crop adaptation, can help breeders when facing challenges like climate change and population growth. Genome-wide diversity analysis thus requires network methods scaling for large data volumes and incorporating fundamental biological phenomena. In this context, we present a new hybrid approach, SO_SCPLOWNARFC_SCPLOW, that benefits from the potential of a powerful machine learning classifier, Approximate Bayesian Computation Random Forest, and from advantages of a mathematical phylogenetic method, SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOW. Consequently, SO_SCPLOWNARFC_SCPLOW is able to handle large data-sets thanks to machine learning and is also based on a deep mathematical theory. On simulated data, our hybrid approach performs very well. When applied to real rice genomic data, it supports a scenario with a single domestication event, that of Japonica. The analysis further highlights the role of early Japonica in the origin of both Indica and circumAus. Finally, it identifies an ancient admixture event, involving circumAus in the emergence of circumBasmati. Together, these findings confirm the importance of early rice history along the Himalayan region.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.