Sequence-based Optimized Chaos Game Representation and Deep Learning for Peptide/Protein Classification
Huang, B.; Zhang, E.; Chaudhari, R.; Gimperlein, H.
Show abstract
As an effective graphical representation method for 1D sequence (e.g., text), Chaos Game Representation (CGR) has been frequently combined with deep learning (DL) for biological analysis. In this study, we developed a unique approach to encode peptide/protein sequences into CGR images for classification. To this end, we designed a novel energy function and enhanced the encoder quality by constructing a Supervised Autoencoders (SAE) neural network. CGR was used to represent the amino acid sequences and such representation was optimized based on the latent variables with SAE. To assess the effectiveness of our new representation scheme, we further employed convolutional neural network (CNN) to build models to study hemolytic/non-hemolytic peptides and the susceptibility/resistance of HIV protease mutants to approved drugs. Comparisons were also conducted with other published methods, and our approach demonstrated superior performance. Supplementary informationavailable online
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Improving protein function prediction with synthetic feature samples created by generative adversarial networks 95%
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 94%
- A deep learning framework for high-throughput mechanism-driven phenotype compound screening 94%
Similar papers in this journal
- Neural Network Models for Sequence-Based TCR and HLA Association Prediction 95%
- THLANet: A Deep Learning Framework for Predicting TCR-pHLA Binding in Immunotherapy Applications 94%
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.