Back

Imputation strategy for population DNA methylation sequencing data

Duplan, A.; Brandt, S.; Garnier, A.; Tost, J.; Sanchez, L.; Duvaux, L.; Maury, S.; Durufle, H.

2025-03-25 bioinformatics
10.1101/2025.03.21.644107 bioRxiv
Show abstract

BackgroundDNA methylation is a central epigenetic mechanism involved in regulating gene expression and responses to environmental factors. Although it can sometimes be passed down through generations, its heritability remains variable depending on the species and biological context. These characteristics make it a key marker for studying genotype-environment interactions. However, whole-genome sequencing for DNA methylation analysis remains costly when applied to large numbers of individuals, prompting researchers to focus on specific regions of interest. This targeted approach often results in data matrices with missing values for some individuals, which can hinder downstream analyses. ResultsOur study used 200 and 189 poplar and oak individuals from natural populations, respectively. We tested and compared seven methods for missing data imputation in the specific context of targeted DNA methylation sequencing data obtained in the three different DNA methylation contexts in plants (CpG, CHG, and CHH). The comparison of the different imputation result allows to evaluate their performance to determine the most suitable approach for this type of data. Among them, NIPALS, MissForest, and LOESS provided the highest accuracy. NIPALS delivered the best overall performance but with moderate computational cost, MissForest achieved similar accuracy with faster computation, and LOESS offered competitive results suitable for large datasets. ConclusionsOur results provide a reference for the selection of imputation strategies in targeted sequencing studies, improving the reliability of DNA methylation analyses and broadening the applicability of this type of data in epigenomic research. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=127 SRC="FIGDIR/small/644107v2_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@1c98f47org.highwire.dtl.DTLVardef@1dd9cfborg.highwire.dtl.DTLVardef@6d3a43org.highwire.dtl.DTLVardef@10c1d8b_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.