Sparse autoencoder features from InterPLM predict neuropeptide precursors among secreted proteins
Kulikova, A. V.; Bookout, A. L.; Koch, T. L.; Safavi-Hemami, H.
Show abstract
Neuropeptides are a diverse class of short, secreted signaling molecules that regulate key physiological processes in animals. Despite their important biological roles and increasingly recognized therapeutic value, the discovery of new neuropeptides remains challenging, largely because their short length and high sequence heterogeneity limit the effectiveness of motif- and homology-based approaches. Here, we present a pipeline for neuropeptide precursor prediction that leverages sparse autoencoders (SAEs) from the protein language model InterPLM to decode dense protein language model embeddings into sparse, disentangled features. We identify a small subset of features strongly associated with neuropeptide precursors that achieve high discriminative performance. A logistic regression classifier trained on this reduced feature set, accurately separates human neuropeptide and non-neuropeptide sequences. We then applied this classifier to important model organisms: mouse (Mus musculus), zebrafish (Danio rerio), nematode (Caenorhabditis elegans), and fruit fly (Drosophila melanogaster ) and show that the approach generalizes across diverse species. Overall, InterPLM SAE features provide an interpretable and effective strategy for neuropeptide prediction and enable a trained classifier to predict neuropeptides from large datasets. A web tool for this classifier is freely available at https://biolib.com/ATGCACTGTTCAGGCCTC/SAE-Neuropeptide-Predictor
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ProteinBERT: A universal deep-learning model of protein sequence and function 95%
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 94%
- PIPENN: Protein Interface Prediction with an Ensemble of Neural Nets 94%
Similar papers in this journal
- Mining hidden knowledge: Embedding models of cause-effect relationships curated from the biomedical literature 93%
- Adversarial training improves model interpretability in single-cell RNA-seq analysis 93%
- Refining the cis-regulatory grammar learned by sequence-to-activity models by increasing model resolution 92%
Similar papers in this journal
- Deconvolving multiplexed protease signatures with substrate reduction and activity clustering 94%
- MoCETSE: A mixture-of-convolutional experts and transformer-based model for predicting Gram-negative bacterial secreted effectors 94%
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.