Back

SAPPTree: Identification of an S-Acylation Motif Drives a Novel S-Acylation Prediction Program

Guo, A. S.; Luong, V. S.; Petropavlovskiy, A. A.; Dang, A.; Doxey, A. C.; Sanders, S. S.; Martin, D. D. O.

2026-07-03 bioinformatics
10.64898/2026.06.30.735287 bioRxiv
Show abstract

S-acylation, the reversible addition of fatty acids to proteins, has emerged as an abundant post-translational modification that drives protein localization and function. With no known consensus sequence, current prediction programs rely on machine learning algorithms that use short peptide sequences and large proteomic datasets. However, current prediction programs often suggest incorrect sites of S-acylation, leading to wasted experimental time and effort following site-directed mutagenesis and low-throughput validation experiments. Using only experimentally confirmed sites of S-acylation, we sought to identify primary sequence, secondary structure, and tertiary structure features common amongst S-acylation sites to aid in developing more robust prediction tools. In doing so, we identified an S-acylation motif including a cysteine cluster flanked by a hydrophobic stretch, and a positively charged polybasic region found within a helical stretch. These features were combined with known or AlphaFold-predicted structures and additional features including residue depth and solvent accessibility into a random forest model to generate a new and more accurate S-acylation prediction program (SAPP), named SAPPTree. All the processed datasets and complete model training pipeline are available at https://github.com/neurdyphagy-lab/palm-prediction-model, while the webserver is available at http://martintools.sci.uwaterloo.ca/.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Journal of Proteome Research
234 papers in training set
Top 0.4%
11.8%
2
PLOS Computational Biology
1863 papers in training set
Top 3%
10.9%
3
Bioinformatics
1204 papers in training set
Top 3%
9.6%
4
Nature Communications
5641 papers in training set
Top 21%
7.8%
5
Cell Systems
201 papers in training set
Top 0.5%
6.7%
6
Protein Science
246 papers in training set
Top 0.5%
6.7%
50% of probability mass above
7
Nature Methods
385 papers in training set
Top 1%
6.7%
8
Molecular Systems Biology
162 papers in training set
Top 0.3%
5.4%
9
Bioinformatics Advances
203 papers in training set
Top 2%
2.4%
10
Journal of Molecular Biology
232 papers in training set
Top 1%
2.1%
11
eLife
5828 papers in training set
Top 44%
2.1%
12
Molecular & Cellular Proteomics
25 papers in training set
Top 0.2%
2.1%
13
PLOS ONE
5266 papers in training set
Top 50%
1.7%
14
Journal of Biological Chemistry
690 papers in training set
Top 6%
1.3%
15
Communications Biology
993 papers in training set
Top 22%
1.1%
16
Nucleic Acids Research
1281 papers in training set
Top 13%
1.0%
17
Analytical Chemistry
218 papers in training set
Top 2%
1.0%
18
Nature Biotechnology
172 papers in training set
Top 4%
0.8%
19
Briefings in Bioinformatics
354 papers in training set
Top 7%
0.8%
20
Molecular & Cellular Proteomics
158 papers in training set
Top 1%
0.8%
21
Life Science Alliance
285 papers in training set
Top 8%
0.8%
22
Biochemical Journal
91 papers in training set
Top 2%
0.8%
23
Nature Machine Intelligence
70 papers in training set
Top 2%
0.8%
24
Nature Chemical Biology
119 papers in training set
Top 3%
0.6%
25
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 2%
0.6%