Back

PEPstrMOD2: Next-generation tertiary structure prediction of chemically modified and non-natural peptides

Jain, S.; Mehta, N. K.; Raina, S.; Kumar, P.; Varun, ; Raghava, G. P. S.

2026-07-06 bioinformatics
10.64898/2026.06.22.733733 bioRxiv
Show abstract

While most existing methods are limited to predicting the tertiary structures of proteins containing only canonical residues, the PEPstrMOD server (developed in 2015) pioneered structure prediction for chemically modified and non-natural peptides. Despite its widespread use, the original framework was restricted to peptides of 7 to 25 residues and relied on older backbone-prediction algorithms. To address these limitations, we present PEPstrMOD2, which introduces three major advancements over its predecessor. First, it replaces the original in-house coordinate generation with state-of-the-art deep learning (DL) algorithms, leveraging AlphaFold2 and ESMFold for highly accurate initial structure prediction. Secondly, it greatly expands the accessible chemical space through incorporation of new, AMBER force-field compatible library of 257 post-translational modifications (PTMs), 428 non-canonical amino acids (NCAAs), and 243 terminal modifications. Lastly, through the application of native scalability of AlphaFold2 (AF2) and ESMFold (EF), PEPstrMOD2 eliminates the original restrictions of the length, enabling the structural modeling of longer, complex therapeutic peptides and small proteins. We evaluated the performance of PEPstrMOD2 against state-of-the-art methods across three distinct peptide datasets. For the AfCyc dataset consisting of 80 cyclic peptides, PEPstrMOD2 obtained a competitive average atom-level Root Mean Square Deviation (RMSD) of 2.05 angstroms, compared to 1.13 angstroms by AlphaFold3 (AF3) and 1.82 angstroms by AfCycDesign. Remarkably, for the modified peptide ModPep433 dataset, PEPstrMOD2 outperformed AF3, achieving the lower average RMSD score of 4.49 angstroms against 4.67 angstroms of AF3. Furthermore, in the case of the ModPep16 benchmark, PEPstrMOD2 achieved 2.50 angstroms average RMSD value, which is two times more accurate than that of the original PEPstrMOD (5.84 angstroms). In summary, PEPstrMOD2 provides a powerful, high-throughput, and highly accurate platform to facilitate peptide-based drug development and structural biology research. While the original PEPstrMOD was restricted to a web server interface, PEPstrMOD2 is available as both an intuitive webserver and a standalone command-line tool via GitHub, featuring Docker support for easy deployment and reproducible, large-scale modeling pipelines (https://webs.iiitd.edu.in/raghava/pepstrmod/).

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.1%
38.4%
2
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.3%
6.5%
3
Briefings in Bioinformatics
354 papers in training set
Top 1%
6.1%
50% of probability mass above
4
Communications Chemistry
48 papers in training set
Top 0.1%
5.3%
5
Nature Communications
5641 papers in training set
Top 30%
4.7%
6
Bioinformatics
1204 papers in training set
Top 5%
3.1%
7
Journal of Cheminformatics
29 papers in training set
Top 0.2%
3.1%
8
Communications Biology
993 papers in training set
Top 8%
2.6%
9
ACS Omega
105 papers in training set
Top 1%
2.3%
10
Protein Science
246 papers in training set
Top 2%
1.9%
11
PLOS Computational Biology
1863 papers in training set
Top 14%
1.9%
12
PLOS ONE
5266 papers in training set
Top 50%
1.7%
13
Computational and Structural Biotechnology Journal
242 papers in training set
Top 4%
1.7%
14
Journal of Molecular Biology
232 papers in training set
Top 2%
1.7%
15
Journal of Medicinal Chemistry
77 papers in training set
Top 0.6%
1.3%
16
Advanced Science
286 papers in training set
Top 6%
1.3%
17
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 1%
1.1%
18
Scientific Reports
3612 papers in training set
Top 67%
1.1%
19
Nature Machine Intelligence
70 papers in training set
Top 2%
1.0%
20
Journal of Proteome Research
234 papers in training set
Top 2%
0.8%
21
International Journal of Molecular Sciences
494 papers in training set
Top 16%
0.8%
22
Nucleic Acids Research
1281 papers in training set
Top 16%
0.6%
23
Chemical Science
73 papers in training set
Top 2%
0.6%
24
Bioinformatics Advances
203 papers in training set
Top 5%
0.6%