AllerStat: Finding Statistically Significant Allergen-Specific Patterns in Protein Sequences by Machine Learning
Goto, K.; Tamehiro, N.; Yoshida, T.; Hanada, H.; Sakuma, T.; Adachi, R.; Kondo, K.; Takeuchi, I.
Show abstract
Cutting-edge technologies such as genome editing and synthetic biology allow us to produce novel foods and functional proteins. However, their toxicity and allergenicity must be accurately evaluated. Allergic reactions are caused by specific amino-acid sequences in proteins (Allergen Specific Patterns, ASPs), of which, many remain undiscovered. In this study, we introduce a data-driven approach and a machine-learning (ML) method to find undiscovered ASPs. The proposed method enables an exhaustive search for amino-acid sub-sequences whose frequencies are statistically significantly higher in allergenic proteins. As a proof-of-concept (PoC), we created a database containing 21,154 proteins of which the presence or absence allergic reactions are already known, and the proposed method was applied to the database. The detected ASPs in the PoC study were consistent with known biological findings, and the allergenicity prediction accuracy using the detected ASPs was higher than extant approaches. TeaserWe propose a computational method for finding statistically significant allergen-specific amino-acid sequences in proteins.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Attention-aware contrastive learning for predicting T cell receptor-antigen binding specificity 94%
- DeepSS2GO: protein function prediction from secondary structure 94%
- Hierarchical cell-type identifier accurately distinguishes immune-cell subtypes enabling precise profiling of tissue microenvironment with single-cell RNA-sequencing 93%
Similar papers in this journal
Similar papers in this journal
- Integration of multi-source gene interaction networks and omics data with graph attention networks to identify novel disease genes 94%
- Identifying B-cell epitopes using AlphaFold2 predicted structures and pretrained language model 94%
- Prediction of bacterial protein-compound interactions with only positive samples 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.