Resolving Genome-to-Phenotype Links in Bacteria: Machine-Learned Inference from Downsampled k-mer Representations
Regueira, T. G. B.; Barra, C.; Lund, O.
Show abstract
Standard approaches to bacterial phenotyping often treat the entire genome as the fundamental unit of information, resulting in high-dimensional inputs that may contain significant redundancy. Consequently, current bacterial phenotyping techniques typically rely on the assumption that entire sequences are required for accurate predictions. While downsampling based on min-hashing or prefix filtering has been used for clustering, its utility as a direct input for predictive machine learning remains underexplored. Here, we show that a novel prefix-based downsampling algorithm can reduce the size of genomes while maintaining relatively high predictive accuracy on phenotype prediction tasks. By combining a prefix reduction strategy with the specificity of short k-mers, we developed a method to downsample entire genomes into k-mer frequency matrices and k-mer-on-a-string representations. We found that ensemble models, such as Random Forest and Gradient Boosting, trained on k-mer frequency matrices from downsampled genome representations outperformed more complex deep learning architectures with the same downsampled representation, particularly on datasets with limited data or highly similar genomes. We were able demonstrate explainability by tracing back the k-mers with the most impact on the models to genes coding for the specific phenotype. Our results demonstrate that downsampling genomic data can yield models with good predictive power thus establishing an alternative when using full genomes is infeasible. We present an approach that offers relatively high performance on bacterial phenotyping tasks and demonstrates a path forward towards lightweight Genome Language Models that will enable analysis of entire genomes.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Balrog: A universal protein model for prokaryotic gene prediction 95%
- Learning, Visualizing and Exploring 16S rRNA Structure Using an Attention-based Deep Neural Network 95%
- ConNIS and labeling instability: new statistical methods for improving the detection of essential genes in TraDIS libraries 93%
Similar papers in this journal
- Keeping up with the genomes: efficient learning of our increasing knowledge of the tree of life 96%
- Read-SpaM: assembly-free and alignment-free comparison of bacterial genomes with low sequencing coverage 94%
- AmpliDiff: An Optimized Amplicon Sequencing Approach to Estimating Lineage Abundances in Viral Metagenomes 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.