Missing Value Imputation using XGboost for Label-Free Mass Spectrometry-Based Proteomics Data
Song, J.; Yu, C.
Show abstract
The label-free mass spectrometry-based proteomics data inevitably suffer from the problem of missing values. The existence of missing values prevents the downstream analyses which need a complete data matrix. Our motivation is to introduce the state-of-art machine learning algorithm XGboost to realize a method of imputation which can improve the accuracy of imputation. But in practical, XGboost has many parameters need to be tuned to deliver on its potential high performance. Although cross validation may find the best parameters, it is much time-consuming. Alternatively, we empirically determined the parameters to two kinds of base learners of XGboost. To explore the robustness and performance of XGboost based imputation with predetermined parameters, we conducted tests on three benchmark datasets. As a comparative, six common imputation methods were also experimented in terms of normalized root mean squared error and Pearson correlation coefficient. The comparative experimental results indicated that the XGboost based imputation method using the linear base learner is competitive to or out-performs its competitors, including the random forest based imputation, by achieving smaller imputation errors and better structure preservation under the empirical parameters for the three benchmark datasets.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Normalization of RNA-Seq Data using Adaptive Trimmed Mean with Multi-reference 95%
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 95%
- SPCS: A Spatial and Pattern Combined Smoothing Method of Spatial Transcriptomic Expression 95%
Similar papers in this journal
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 94%
- A Machine Learning Model of Microscopic Agglutination Test for Diagnosis of Leptospirosis 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.