Designing minimal E. coli genomes using variational autoencoders
Barnes, C. P.; Buchan, D.; Shcherbakova, A.
Show abstract
Designing minimal bacterial genomes remains a key challenge in synthetic biology. There is currently a lack of efficient tools for the rapid generation of streamlined bacterial genomes, limiting research in this area. Here, using a pangenome dataset for Escherichia coli, we show that variational autoencoders with modified loss functions can successfully create minimised genomes retaining the essential genes identified in the literature. We then sampled new genomes from our fitted model and performed computational validation using an E. coli whole-cell model. We found 6 out of 100 of the sampled genomes were viable in the computer model. These underwent a minimization routine starting from the MG1655 genome giving rise to six new minimal genomes with around a 40 % reduction in size. This study proposes a rapid, machine learning-based approach for bacterial sequence generation, that could accelerate the genomic design process.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Deciphering the Biosynthetic Potential of Microbial Genomes Using a BGC Language Processing Neural Network Model 94%
- TedSim: temporal dynamics simulation of single cell RNA-sequencing data and cell division history 94%
- Accurate assembly of minority viral haplotypes from next-generation sequencing through efficient noise reduction 94%
Similar papers in this journal
- Statistical Analysis of Variability in TnSeq Data Across Conditions Using Zero-Inflated Negative Binomial Regression 95%
- Inference of Genomic Landscapes using Ordered Hidden Markov Models with Emission Densities (oHMMed) 94%
- AmpliDiff: An Optimized Amplicon Sequencing Approach to Estimating Lineage Abundances in Viral Metagenomes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.