Back

Protein language modeling and supervised machine learning reveal the functional landscape of antimicrobial resistance genes in African metagenomes

Muigano, M. N.; Kamau, G. W.

2026-01-25 microbiology
10.64898/2026.01.24.701459 bioRxiv
Show abstract

In this work, we analyzed 184 metagenomes from sub-Saharan Africa to characterize the functional landscape of ARGs using a combination of homology-based annotation and protein language model embeddings. We obtained 5,066 high-confidence ARG protein sequences from the African metagenomes, which we compared with 6,052 reference ARGs from CARD using embeddings generated with the ESM-2 protein language model. Additionally, we used a random forest classifier to determine the role of amino acid sequence and physiochemical features in discriminating ARGs and non-ARG sequences. The curated dataset revealed a predominance of ESKAPE pathogens in the resistome. {beta}-lactam resistance was the most prevalent functional class, accounting for 4,046 ARG assignments (79.87%). At the country level, Burkina Faso, Malawi, and Benin exhibited the highest ARG hits per sample, thus demonstrating geographic heterogeneity in ARG burden. Protein language modeling demonstrated that African ARG sequences largely occupied the same functional subspace as globally curated CARD proteins. Supervised machine learning based on protein compositional and physicochemical features achieved high discriminatory performance between ARGs and non-ARGs (accuracy = 0.94, ROC-AUC = 0.99). Feature importance analysis identified amino acid composition, protein length, molecular weight, and isoelectric point as key discriminators, with statistically significant differences between ARG and non-ARG proteins (Mann-Whitney U test, p < 0.001). These results suggest that ARGs in African metagenomes are shaped primarily by ecological filtering and antibiotic selection pressure rather than by the emergence of novel resistance functions. This work provides a functional baseline for AMR surveillance in Africa and highlights the value of protein language models for resistome-scale analyses.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.