Back

ActSeekN: A Structural-Motif-Based Pipeline for Interpretable Enzyme Function Annotation

Castillo, S.; Gu, C.; Jouhten, P.; Peddinti, G.; Ollila, S. O. H.

2026-04-28 bioinformatics
10.64898/2026.04.24.720574 bioRxiv
Show abstract

Accurate enzyme function annotation remains a major bottleneck in genome analysis despite the rapid expansion of available protein sequence and structure data. Most existing methods rely on sequence similarity or machine-learning representations, which often perform poorly for proteins with low sequence identity or convergent evolutionary histories. Because enzymatic activity is determined by the three-dimensional arrangement of catalytic and binding-site residues, structure-based approaches offer a mechanistically grounded alternative. However, their broader application has been constrained by the limited size and coverage of curated active-site reference databases. To address this challenge, we developed ActSeekN, a structural-motif-based functional annotation pipeline that combines the ActSeek active-site search algorithm with a newly constructed large-scale reference database derived from AlphaFold-predicted structures, UniProt annotations, and curated catalytic residue information. This framework enables rapid and scalable identification of conserved catalytic motifs across structurally related proteins, allowing function to be transferred on the basis of local three-dimensional catalytic geometry rather than global sequence similarity. In this way, ActSeekN overcomes a central limitation of previous structure-based methods by expanding the searchable space of catalytic motifs while retaining mechanistic interpretability. Benchmarking against state-of-the-art machine-learning approaches demonstrates competitive or superior performance. Applications to yeast, human, and Trichoderma reesei proteomes refine existing annotations, complete partial EC assignments, and identify previously unrecognized enzymatic functions, highlighting ActSeekN as a powerful tool for genome annotation and biotechnology. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=99 SRC="FIGDIR/small/720574v1_ufig1.gif" ALT="Figure 1"> View larger version (37K): org.highwire.dtl.DTLVardef@f5da16org.highwire.dtl.DTLVardef@c0faa4org.highwire.dtl.DTLVardef@18765fforg.highwire.dtl.DTLVardef@3960b9_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Nucleic Acids Research
1281 papers in training set
Top 1%
13.0%
2
Bioinformatics
1204 papers in training set
Top 2%
10.8%
3
Nature Communications
5641 papers in training set
Top 18%
9.8%
4
Protein Science
246 papers in training set
Top 0.4%
8.0%
5
Journal of Molecular Biology
232 papers in training set
Top 0.2%
6.8%
6
PLOS Computational Biology
1863 papers in training set
Top 6%
5.6%
50% of probability mass above
7
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
4.4%
8
Bioinformatics Advances
203 papers in training set
Top 1%
4.4%
9
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.3%
10
eLife
5828 papers in training set
Top 37%
2.8%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.4%
12
Advanced Science
286 papers in training set
Top 4%
2.1%
13
Communications Biology
993 papers in training set
Top 16%
1.5%
14
Plant Communications
36 papers in training set
Top 0.6%
1.4%
15
Cell Reports Methods
165 papers in training set
Top 3%
1.1%
16
Nature Methods
385 papers in training set
Top 5%
1.1%
17
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 1.0%
1.1%
18
Scientific Reports
3612 papers in training set
Top 64%
1.1%
19
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.1%
20
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 41%
0.9%
21
Cell Systems
201 papers in training set
Top 5%
0.6%
22
BMC Bioinformatics
457 papers in training set
Top 6%
0.6%
23
Journal of Cheminformatics
29 papers in training set
Top 0.8%
0.6%
24
Molecular & Cellular Proteomics
25 papers in training set
Top 0.4%
0.6%
25
Molecular Biology and Evolution
542 papers in training set
Top 5%
0.6%