GLADE: Accurate inference of Gains, Losses, Ancestral genomes, and Duplication Events for comparative genomics
Belcher, L. J.; Kelly, S.
Show abstract
Changes in gene content through gain and loss play a key role in the adaptation and diversification of species. Accordingly, our ability to detect and accurately document the history of these changes is important for our understanding the evolutionary trajectories of life on Earth. Here we present GLADE, a tool that accurately reconstructs gene gains, losses, and duplications for a set of species under consideration and uses this information to infer ancestral gene contents for every speciation event in the species tree. GLADE requires as input only a standard OrthoFinder results directory, and outputs the full evolutionary history of every orthogroup, including branch-specific changes and reconstructed ancestral genomes. We benchmark GLADE using both real and simulated data and show that GLADE accurately identifies orthogroup gains, losses, and duplications, and reconstructs ancestral orthogroup sizes with higher precision and overall accuracy than any competitor method. To illustrate the utility of the method, we apply GLADE to a dataset of 78 mammalian genomes and uncover repeated contractions in orthogroups associated with tooth formation on branches leading to ant- and termite-eating mammals - revealing convergent genomic signatures underlying this dietary specialization. GLADE and accompanying documentation and tutorials are freely available at https://github.com/lauriebelch/GLADE/.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ecological predictors of organelle genome evolution: Phylogenetic correlations with taxonomically broad, sparse, unsystematized data 97%
- Improved robustness to gene tree incompleteness, estimation errors, and systematic homology errors with weighted TREE-QMC 97%
- Towards reliable detection of introgression in the presence of among-species rate variation 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.