Exploring an experiment-split method to estimate the generalization ability in new data: DeepKme as an example
Zou, G.; Li, L.
Show abstract
A Large number of predictors have been built based on different data sets for predicting different post-translational modification sites. However, limited to our knowledge, most of them gave an overfitting estimation of their generalization ability in new data because of the intrinsic trait--not considering the experimental sources of the new data--of the cross-validation method. Thus, we proposed and explored a new method--the experiment-split method--imitating the blinded assessment to deal with the overfitting problem in the new data. The experiment-split method logically split the training and test data based on the datas different experimental sources, and the new data can be regarded as the data from different experimental sources. To specifically illustrate the experiment-split method, we combined an actual application, DeepKme--a predictor built by us for the lysine methylation sites, to demonstrate how it be used in the true scenarios. We compared the cross-validation method with the experiment-split method. The result suggested the experiment-split method could effectively relieve the overfitting compared with the cross-validation method and may be widely used in the field of identification participated by multiple experiments. We believe DeepKme would facilitate the related researchers deep thought of the experiment-split method and the overfitting phenomenon, and of course, advance the study of the lysine methylation and similar fields.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 95%
- ISMI-VAE: A Deep Learning Model for Classifying Disease Cells Using Gene Expression and SNV Data 93%
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 93%
Similar papers in this journal
- ICAN: interpretable cross-attention network for identifying drug and target protein interactions 94%
- Using deep maxout neural networks to improve the accuracy of function prediction from protein interaction networks 94%
- Predicting compound-protein interaction using hierarchical graph convolutional networks 94%
Similar papers in this journal
- Deep6mA: a deep learning framework for exploring similar patterns in DNA N6-methyladenine sites across different species 97%
- DNFE: Directed-network flow entropy for detecting the tipping points during biological processes 95%
- Base-resolution prediction of transcription factor binding signals by a deep learning framework 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.