resLens: genomic language models to enhance antibiotic resistance gene detection
Mollerus, M.; Dittmar, K.; Crandall, K. A.; Rahnavard, A.
Show abstract
The rise of antibiotic resistance necessitates advanced tools to detect and analyze antibiotic resistance genes (ARGs). We present resLens, a family of genomic language models that leverage latent genomic representations to enhance ARG detection and analysis. Unlike alignment-based methods constrained by reference databases, resLens fine-tunes a pre-trained DNA language model on curated ARG datasets, achieving competitive or superior performance in classifying resistance genes across multiple evaluation scenarios, including when ARGs exhibit sequences and mechanisms of resistance dissimilar to those in reference datasets.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Maast: genotyping thousands of microbial strains efficiently 95%
- DIVE: a reference-free statistical approach to diversity-generating and mobile genetic element discovery 94%
- A k-mer-based maximum likelihood method for estimating distances of reads to genomes enables genome-wide phylogenetic placement. 94%
Similar papers in this journal
Similar papers in this journal
- Hi-C-LSTM: Learning representations of chromatin contacts using a recurrent neural network identifies genomic drivers of conformation 94%
- A convolutional neural network highlights mutations relevant to antimicrobial resistance in Mycobacterium tuberculosis 94%
- Convolutional neural networks quantify antibiotic resistance in Mycobacterium tuberculosis with diagnostic grade accuracy and predict treatment response 94%
Similar papers in this journal
- Deciphering the Biosynthetic Potential of Microbial Genomes Using a BGC Language Processing Neural Network Model 94%
- Deciphering enzymatic potential in metagenomic reads through DNA language models 92%
- Accurate assembly of minority viral haplotypes from next-generation sequencing through efficient noise reduction 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.