Back

ALPAR: Automated Learning Pipeline for Antimicrobial Resistance

Yurtseven, A.; Joeres, R.; Kalinina, O. V.

2025-07-11 bioinformatics
10.1101/2025.07.08.663126 bioRxiv
Show abstract

The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially to newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates generation of machine learning-ready data tables and both training of machine learning and execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of distribution of mutations, enhancing its utility for researchers. Our tool is accessible through the ALPAR GitHub page (https://github.com/kalininalab/ALPAR) and installable via conda (https://anaconda.org/kalininalab/ALPAR).

Published in Bioinformatics (predicted rank #14) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.