Codon usage bias levels predict taxonomic identity and genetic composition
Khomtchouk, B. B.
Show abstract
In this study, we investigate how an organisms codon usage bias levels can serve as a predictor and classifier of various genomic and evolutionary features across the three kingdoms of life (archaea, bacteria, eukarya). We perform secondary analysis of existing genetic datasets to build several artificial intelligence (AI) and machine learning models trained on over 13,000 organisms that show it is possible to accurately predict an organisms DNA type (nuclear, mitochondrial, chloroplast) and taxonomic identity simply using its genetic code (64 codon usage frequencies). By leveraging advanced AI and machine learning methods to accurately identify evolutionary origins and genetic composition from codon usage patterns, our study suggests that the genetic code can be utilized to train accurate machine learning classifiers of taxonomic and phylogenetic features. Our dataset and analyses are made publicly available on Github and the UCI Machine Learning Repository (https://archive.ics.uci.edu/ml/datasets/Codon+usage) to facilitate open-source reproducibility and community engagement.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automated evaluation of multiple sequence alignment methods to handle third generation sequencing errors 93%
- StrainFLAIR: Strain-level profiling of metagenomic samples using variation graphs 93%
- Exploring Environmental Coverages of Species: A New Variable Selection Methodology for Rulesets from the Genetic Algorithm for Ruleset Prediction 92%
Similar papers in this journal
Similar papers in this journal
- AvP: a software package for automatic phylogenetic detection of candidate horizontal gene transfers. 95%
- An assembly-free method of phylogeny reconstruction using short-read sequences from pooled samples without barcodes 95%
- Learning, Visualizing and Exploring 16S rRNA Structure Using an Attention-based Deep Neural Network 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.