Back

AmyloDeep: pLM-based ensemble model for predicting amyloid propensity from the amino acid sequence

Davtyan, A.; Khachatryan, A.; Petrosyan, R.

2025-09-18 bioinformatics
10.1101/2025.09.16.676495 bioRxiv
Show abstract

Amyloids are predominantly {beta}-sheet-rich, stable protein structures that can maintain their presence in the human body for multiple years. Amyloid protein aggregates contribute to the development of multiple neurodegenerative diseases, such as Alzheimers, Parkinsons, and Huntingtons, and are involved in different vital functions, such as memory formation and immune system function. Here, we used advanced machine learning and deep learning techniques to predict amyloid propensity from the amino acid sequence. First, we aggregated labeled amino acid sequence data from multiple sources, obtaining a roughly balanced dataset of 2366 sequences for binary classification. We leveraged that data to both fine-tune the ESM2 model and to train new models based on protein embeddings from ESM2 and UniRep. The predictions from these models were then unified into a single soft voting ensemble model, yielding highly robust and accurate results. We further made a tool where users can provide the amino acid sequence and get the amyloid formation probabilities of different segments of the input sequence. Users can access the light version of AmyloDeep through the web server at https://amylodeep.com/, and the full model is available as a Python package at https://pypi.org/project/amylodeep/. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=47 SRC="FIGDIR/small/676495v1_ufig1.gif" ALT="Figure 1"> View larger version (13K): org.highwire.dtl.DTLVardef@64615aorg.highwire.dtl.DTLVardef@339a7corg.highwire.dtl.DTLVardef@1e36e14org.highwire.dtl.DTLVardef@4fe717_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.