Imputation strategy for population DNA methylation sequencing data
Duplan, A.; Brandt, S.; Garnier, A.; Tost, J.; Sanchez, L.; Duvaux, L.; Maury, S.; Durufle, H.
Show abstract
BackgroundDNA methylation is a central epigenetic mechanism involved in regulating gene expression and responses to environmental factors. Although it can sometimes be passed down through generations, its heritability remains variable depending on the species and biological context. These characteristics make it a key marker for studying genotype-environment interactions. However, whole-genome sequencing for DNA methylation analysis remains costly when applied to large numbers of individuals, prompting researchers to focus on specific regions of interest. This targeted approach often results in data matrices with missing values for some individuals, which can hinder downstream analyses. ResultsOur study used 200 and 189 poplar and oak individuals from natural populations, respectively. We tested and compared seven methods for missing data imputation in the specific context of targeted DNA methylation sequencing data obtained in the three different DNA methylation contexts in plants (CpG, CHG, and CHH). The comparison of the different imputation result allows to evaluate their performance to determine the most suitable approach for this type of data. Among them, NIPALS, MissForest, and LOESS provided the highest accuracy. NIPALS delivered the best overall performance but with moderate computational cost, MissForest achieved similar accuracy with faster computation, and LOESS offered competitive results suitable for large datasets. ConclusionsOur results provide a reference for the selection of imputation strategies in targeted sequencing studies, improving the reliability of DNA methylation analyses and broadening the applicability of this type of data in epigenomic research. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=127 SRC="FIGDIR/small/644107v2_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@1c98f47org.highwire.dtl.DTLVardef@1dd9cfborg.highwire.dtl.DTLVardef@6d3a43org.highwire.dtl.DTLVardef@10c1d8b_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ARPEGGIO: Automated Reproducible Polyploid EpiGenetic GuIdance workflOw 93%
- A Nextflow pipeline for molecular quantitative trait loci mapping in small sample size datasets with an application in Atlantic salmon 93%
- Comparing methylation levels assayed in GC-rich regions with current and emerging methods 93%
Similar papers in this journal
- LuxHMM: DNA methylation analysis with genome segmentation via Hidden Markov Model 94%
- Probabilistic modeling methods for cell-free DNA methylation based cancer classification 94%
- Improved Quality Metrics for Association and Reproducibility in Chromatin Accessibility Data Using Mutual Information 93%
Similar papers in this journal
- Extraction and high-throughput sequencing of oak heartwood DNA: assessing the feasibility of genome-wide DNA methylation profiling 94%
- Genetic control of the leaf ionome in pearl millet and correlation with root and agromorphological traits 93%
- Identification of differential hypothalamic DNA methylation and gene expression associated with sexual partner preferences in rams 93%
Similar papers in this journal
- Comprehensive benchmarking of software for mapping whole genome bisulfite data: from read alignment to DNA methylation analysis 95%
- Systematic evaluation of cell-type deconvolution pipelines for sequencing-based bulk DNA methylomes 93%
- Molecular Group and Correlation Guided Structural Learning for Multi-Phenotype Prediction 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.