slideimp: Efficient Imputation for DNA Methylation Data
Pham, H.; Lombroso, A. P.; Cevik, E. C.; Taylor, H. S.; O'Donnell, K. J.
Show abstract
MotivationThere is a growing need for efficient imputation methods for high-dimensional DNA methylation (DNAm) datasets. Existing microarray imputation approaches, such as k-nearest neighbors (K-NN) or principal component analysis (PCA)-based methods, provide high accuracy but can be computationally intensive, while methods for whole-genome data are not designed for large cohorts. We developed slideimp, an R package that implements sliding window, groupable, parallelized K-NN and optimized PCA imputation to address these limitations. ResultsBenchmarks on microarray DNAm datasets demonstrate that slideimp achieves up to 150x faster runtime and 10x-100x lower memory usage while maintaining comparable or superior accuracy over existing methods and implementations. For K-NN and PCA imputation, slideimp supports grouped imputation which enhances imputation efficiency and accuracy. In a whole-genome dataset, sliding window K-NN imputation substantially increased the correlation of the Horvath 2013 clock with chronological age from 0.131 to 0.477. Additional features include targeted imputation of CpG subsets for K-NN and estimation of imputation accuracy via repeated cross-validation. The efficient and flexible DNAm imputation methods implemented by slideimp can easily be applied to other high-dimensional data types. Availability and ImplementationThe code to fully reproduce all analyses presented in this paper is available on GitHub at https://github.com/hhp94/slideimp_paper. The R package slideimp is also available on GitHub at https://github.com/hhp94/slideimp. Supplemental figures are available online.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Guidelines for cell-type heterogeneity quantification based on a comparative analysis of reference-free DNA methylation deconvolution software 94%
- MethylNet: An Automated and Modular Deep Learning Approach for DNA Methylation Analysis 94%
- DMRscaler: A Scale-Aware Method to Identify Regions of Differential DNA Methylation Spanning Basepair to Multi-Megabase Features 93%
Similar papers in this journal
- scMET: Bayesian modelling of DNA methylation heterogeneity at single-cell resolution 94%
- A systematic evaluation of 41 DNA methylation predictors across 101 data preprocessing and normalization strategies highlights considerable variation in algorithm performance 94%
- ContamLD: Estimation of Ancient Nuclear DNA Contamination Using Breakdown of Linkage Disequilibrium 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.