Back

The structural context of mutations in proteins predicts their effect on antibiotic resistance

Green, A. G.; Tasmin, M.; Vargas, R.; Farhat, M. R.

2026-06-29 bioinformatics
10.1101/2025.09.23.676583 bioRxiv
Show abstract

In Mycobacterium tuberculosis, a prevalent and deadly pathogen, resistance to antibiotics evolves primarily through non-synonymous mutations in proteins. Sequence-based analyses are currently used to understand the genetic basis of antibiotic resistance, either via genotype-phenotype association, or via signals of convergent evolution. These methods focus on primary sequence and often neglect other biological signals such as protein structural information. We hypothesize that integrating the structural context of mutations improves the prediction of effects on function and phenotype. We curate high confidence structural annotations for the M. tuberculosis proteome from 1,371 crystallography and 2,316 AlphaFold predictions, and combine the structures with mutations from over 31,000 clinical M. tuberculosis isolates. We demonstrate that mutations in proteins known to cause resistance are clustered in 3D space, even in proteins where inactivating mutations at any position are thought to cause resistance. We develop a statistic to search the M. tuberculosis proteome for signal of clustered mutations, finding over 450 proteins that display this signal, many of which have a known relationship with antibiotic resistance. We show that a supervised classifier trained on 3D distance to known resistance sites alone has an F1 score of 94.6% at classifying mutations as resistance-conferring across proteins. This work demonstrates that protein structure provides useful information for categorizing which variants may cause antibiotic resistance, even when the majority of structures are AI-predicted.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Cell Systems
201 papers in training set
Top 0.1%
18.9%
2
Nature Communications
5641 papers in training set
Top 11%
15.4%
3
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 6%
6.9%
4
PLOS Computational Biology
1863 papers in training set
Top 5%
6.9%
5
eLife
5828 papers in training set
Top 24%
4.9%
50% of probability mass above
6
Molecular Biology and Evolution
542 papers in training set
Top 2%
4.4%
7
Scientific Reports
3612 papers in training set
Top 29%
3.5%
8
Molecular Systems Biology
162 papers in training set
Top 0.6%
3.3%
9
Bioinformatics
1204 papers in training set
Top 5%
3.3%
10
Communications Biology
993 papers in training set
Top 8%
2.5%
11
Protein Science
246 papers in training set
Top 2%
2.4%
12
Science Advances
1243 papers in training set
Top 21%
1.5%
13
Nature Methods
385 papers in training set
Top 4%
1.5%
14
Biophysical Journal
631 papers in training set
Top 3%
1.5%
15
Journal of Molecular Biology
232 papers in training set
Top 2%
1.4%
16
PLOS ONE
5266 papers in training set
Top 54%
1.2%
17
Nucleic Acids Research
1281 papers in training set
Top 11%
1.2%
18
Nature Machine Intelligence
70 papers in training set
Top 2%
1.1%
19
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
0.9%
20
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 1%
0.9%
21
Science
477 papers in training set
Top 8%
0.9%
22
Structure
193 papers in training set
Top 2%
0.9%
23
Nature Ecology & Evolution
18 papers in training set
Top 0.4%
0.9%
24
Genome Research
468 papers in training set
Top 7%
0.6%
25
Genome Biology
637 papers in training set
Top 9%
0.6%
26
Genome Biology and Evolution
338 papers in training set
Top 3%
0.6%
27
GENETICS
483 papers in training set
Top 5%
0.6%