Back

Developing machine-learning-based amyloid predictors with Cross-Beta DB

Gonay, V.; Dunne, M. P.; Caceres-Delpiano, J.; Kajava, A. V.

2024-02-14 bioinformatics
10.1101/2024.02.12.579644 bioRxiv
Show abstract

Due to shifts in environmental conditions, mutations, or interactions with other biomolecules, some proteins that would normally be soluble can undergo aggregation, resulting in the formation of clumps of amyloid fibrils. Understanding of this phenomenon is of paramount importance due not only to its association with various diseases (including Alzheimers disease), but also due to increasingly abundant evidence for its functional roles. Numerous studies have demonstrated that the propensity to form amyloids is coded by the amino acid sequence and this finding has paved the way for the development of several computational predictors of amyloidogenicity. The ultimate objective of computational methods is to accurately predict the formation of disease-related and functionally relevant amyloids that occur in vivo. These amyloid fibrils are known to form very specific "cross-{beta}" structures of protein regions longer than about 15 residues. Remarkably, despite the significance of the naturally occurring amyloids, there has been a lack of datasets specifically dedicated to them. Hence, we built Cross-Beta DB, a database composed of cross-{beta} amyloids formed in natural conditions. This database is expected to be indispensable for benchmarking amyloid predictors. We used the Cross-Beta DB to train and benchmark several such algorithms, using machine learning. The best-performing of these, the random-forest-based Cross-Beta RF Predictor, demonstrated superior performance over the other existing methods, fostering high expectations for an improved prediction of naturally occurring amyloids.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.