Imputation for Lipidomics and Metabolomics (ImpLiMet): Online optimization and method selection for missing data imputation
Ou, H.; Surendra, A.; McDowell, G. S. V.; Hashimoto-Roth, E.; Xia, J.; Bennett, S. A. L.; Cuperlovic-Culf, M.
Show abstract
MotivationMissing values are often unavoidable in modern high-throughput measurements due to various experimental or analytical reasons. Imputation, the process of replacing missing values in a dataset with estimated values, plays an important role in multivariate and machine learning analyses. Three missingness patterns have been conceptualized: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). Each describes unique dependencies between the missing and observed data. As a result, the optimal imputation method for each dataset depends on the type of data, the cause of the missing data, and the nature of relationships between the missing and observed data. The challenge is to identify the optimal imputation solution for a given dataset. ResultsImpLiMet: is a user-friendly UI-platform that enables users to impute missing data using eight different methods. For the users dataset, ImpLiMet can suggest the optimal imputation solution through a grid search-based investigation of the error rate for imputation across three missingness data simulations. The effect of imputation can be visually assessed by histogram, kurtosis and skewness analyses, as well as principal component analysis (PCA) comparing the impact of the chosen imputation method on the distribution and overall behaviour of the data. Availability and implementationImpLiMet is freely available at https://complimet.ca/shiny/implimet/ with software accessible at https://github.com/complimet/ImpLiMet Contactsteffanyann.bennett@uottawa.ca and miroslava.cuperlovic-culf@nrc-cnrc.gca. Supplementary informationSupplementary data are available at Bioinformatics Advances online.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions 93%
- Generating Correlated Data for Omics Simulation 93%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 93%
Similar papers in this journal
- Information-Content-Informed Kendall-tau Correlation: Utilizing Missing Values 95%
- MESSES: Software for Transforming Messy Research Datasets into Clean Submissions to Metabolomics Workbench for Public Sharing 94%
- Major Update and Improved Validation Functionality in the mwtab Python Library and the Metabolomics Workbench File Status Website 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.