EnzyKAN: Protein Language Model Embeddings and Kolmogorov-Arnold Network Variants for Enzyme Commission Classification with a Proposed Electron-Transfer Physics Feature Framework
R, S.; Reddy, B. R. R.
Show abstract
MotivationComputational enzyme classification has previously utilised sequence homology features and protein language model embeddings. The Kolmogorov-Arnold Network (KAN) paradigm, which uses learnable edge functions rather than fixed ones, has shown promising results in biological sequence tasks. ResultsA fully reproducible investigation of KAN variants for seven-class EC classification on up to 9,516 labelled sequences from the CLEAN benchmark [1] (9,386 for language model experiments). In the sequence only settings, fixed basis KAN variants outperformed an MLP baseline moderately (macro F1 = 0.17-0.29). Utilisation of ESM-2 650M embeddings [2] greatly improved results via 5-fold cross-validation: MLP macro F1 = 0.750 {+/-} 0.009, accuracy = 0.823 {+/-} 0.009; learnable SineKAN macro F1 = 0.716 {+/-} 0.023, accuracy = 0.788 {+/-} 0.019. MLP performed comparably but did not exceed conventional baselines. As an aside, we introduce but do not investigate an approach to EC oxidoreductase sub-classification through the use of a Marcus theory-based electron transfer feature framework. AvailabilityCode and result files are available at https://github.com/sanjuz-cas/ENZYKAN.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.