Predicting locus phylogenetic utility using machine learning
Knyshov, A.; Walling, A. G.; Guccione, C.; Schwartz, R.
Show abstract
Disentangling evolutionary signal from noise in genomic datasets is essential to building phylogenies. The efficiency of current sequencing platforms and workflows has resulted in a plethora of large-scale phylogenomic datasets where, if signal is weak, it can be easily overwhelmed with non-phylogenetic signal and noise. However, the nature of the latter is not well understood. Although certain factors have been investigated and verified as impacting the accuracy of phylogenetic reconstructions, many others (as well as interactions among different factors) remain understudied. Here we use a large simulation-based dataset and machine learning to better understand the factors, and their interactions, that contribute to species tree error. We trained Random Forest regression models on the features extracted from simulated alignments under known phylogenies to predict the phylogenetic utility of the loci. Loci with the worst utility were then filtered out, resulting in an improved signal-to-noise ratio across the dataset. We investigated the relative importance of different features used by the model, as well as how they correspond to the originally simulated properties. We further used the model on several diverse empirical datasets to predict and subset the least reliable loci and re-infer the phylogenies. We measure the impacts of the subsetting on the overall topologies, difficult nodes identified in the original studies, as well as branch length distribution. Our results suggest that subsetting based on the utility predicted by the model can improve the topological accuracy of the trees and their average statistical support, and limits paralogy and its effects. Although the topology generated from the filtered datasets may not always be dramatically different from that generated from unfiltered data, the worst loci consistently yielded different topologies and worst statistical support, indicating that our protocol identified phylogenetic noise in the empirical data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Phylogenetic conflicts, combinability, and deep phylogenomics in plants 97%
- Whole-genomes illuminate the drivers of gene tree discordance and the tempo of tinamou diversification (Aves: Tinamidae) 97%
- Assessing Confidence in Root Placement on Phylogenies: An Empirical Study Using Non-Reversible Models for Mammals 96%
Similar papers in this journal
- Categorical edge-based analyses of phylogenomic data reveal conflicting signals for difficult relationships in the avian tree 97%
- Comprehensive taxon sampling and vetted fossils help clarify the time tree of shorebirds (Aves, Charadriiformes) 95%
- Molecular species delimitation in the primitively segmented spider genus Heptathela endemic to Japanese islands 94%
Similar papers in this journal
- Phylogeographic model selection using convolutional neural networks 93%
- A target capture approach for phylogenomic analyses at multiple evolutionary timescales in rosewoods (Dalbergia spp.) and the legume family (Fabaceae) 93%
- Chromosome-scale inference of hybrid speciation and admixture with convolutional neural networks 92%
Similar papers in this journal
- A consensus phylogenomic approach highlights paleopolyploid and rapid radiation in the history of Ericales 94%
- Species Tree Topology Impacts the Inference of Ancient Whole-Genome Duplications Across the Angiosperm Phylogeny 94%
- Robustness of RADseq for evolutionary network reconstruction from gene trees 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.