Building a tRNA thermometer to access the world's biochemical diversity
Cimen, E.; Jensen, S. E.; Buckler, E. S.
Show abstract
ABSTRACTBecause ambient temperature affects biochemical reactions, organisms living in extreme temperature conditions adapt protein composition and structure to maintain biochemical functions. While it is not feasible to experimentally determine optimal growth temperature (OGT) for every known microbial species, organisms adapted to different temperatures have measurable differences in DNA, RNA, and protein composition that allow OGT prediction from genome sequence alone. In this study, we built a model using tRNA sequence to predict OGT. We used tRNA sequences from 100 archaea and 683 bacteria species as input to train two Convolutional Neural Network models. The first pairs individual tRNA sequences from different species to predict which comes from a more thermophilic organism, with accuracy ranging from 0.538 to 0.992. The second uses the complete set of tRNAs in a species to predict optimal growth temperature, achieving a maximum r2 of 0.86; comparable with other prediction accuracies in the literature despite a significant reduction in the quantity of input data. This model improves on previous OGT prediction models by providing a model with minimum input data requirements, removing laborious feature extraction and data preprocessing steps, and widening the scope of valid downstream analyses.Competing Interest StatementThe authors have declared no competing interest.View Full Text
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Analysis of computational codon usage models and their association with translationally slow codons 93%
- DeLUCS: Deep Learning for Unsupervised Clustering of DNA Sequences 92%
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.