Back

DeepKOALA: A Fast and Accurate Deep Learning Framework for KEGG Orthology Assignment

Yu, Z.; Meng, L.; Nguyen, C. H.; Mamitsuka, H.; Kanehisa, M.; Ogata, H.

2026-01-08 bioinformatics
10.64898/2026.01.07.698072 bioRxiv
Show abstract

The KEGG Orthology (KO) system links DNA and protein sequences to biological functions and pathways, providing a curated, fundamental and consistent annotation framework across all domains of life. While accurate, traditional sequence alignment-based annotation methods are computationally expensive, which severely limits their application in large-scale datasets. To address this challenge, we introduce DeepKOALA, a deep learning approach based on Gated Recurrent Units (GRU), which frames KO annotation as an open-set recognition task. This design reduces false positives arising from out-of-scope sequences and, together with a lightweight GRU backbone, enables high-throughput annotation. The GRU-based model was benchmarked against four other deep learning architectures and showed the best balance between speed and accuracy. We further performed a cross-species evaluation of DeepKOALA in comparison with existing KO annotation tools, where it achieved an F1-score of 0.8653 and 37.5-fold acceleration compared with BlastKOALA. We also provide a specialized fragment model for handling incomplete sequences and an optional Multi-domain mode. Together, these features make DeepKOALA a scalable, efficient, and accurate solution for high-throughput function annotation. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=61 SRC="FIGDIR/small/698072v1_ufig1.gif" ALT="Figure 1"> View larger version (15K): org.highwire.dtl.DTLVardef@8c3b53org.highwire.dtl.DTLVardef@8acfa8org.highwire.dtl.DTLVardef@147398borg.highwire.dtl.DTLVardef@112d762_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.