Models trained with noisy genomes extend bacterial phenotype prediction into deep time
Koldaeva, A.; Szollosi, G.; Bagrova, O.; Mitchell, J. A. M.; Hugenholtz, P.; Spang, A.; Woodcroft, B. J.; Williams, T. A.
Show abstract
Predicting phenotype from genotype in extant organisms is increasingly tractable through the accumulation of genome sequences and the development of machine-learning algorithms. Here we show that machine learning can be applied to reconstructed ancestral gene content, extending these predictions into the past. We trained models on a diverse set of bacterial phenotypes and found that introducing noise into gene content profiles allows predictions to generalize over larger evolutionary distances. For phenotypes with signal spread across many genes - such as metabolic oxygen use, cell envelope architecture and optimal growth temperature - noise augmentation extends resolution back to the root of the bacterial domain, while for other phenotypes - including GC content and sporulation - the range remains more limited. We therefore conclude that the last bacterial common ancestor (LBCA) was likely an anaerobic, double-membraned, and moderately thermophilic bacterium (46-75{degrees}C). Moreover, this work provides a general approach for learning about the genomic basis of phenotypes and drawing inferences about their early evolution.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Functional Decomposition of Metabolism allows a system-level quantification of fluxes and protein allocation towards specific metabolic functions 95%
- Improved maximum growth rate prediction from microbial genomes by integrating phylogenetic information 94%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.