Optimizing effect sizes and specificity trumps machine learning when building DNA methylation reference panels for cell-type deconvolution
Guo, X.; Teschendorff, A.
Show abstract
Accurate cell-type deconvolution is critical for correct interpretation of Epigenome-Wide Association Studies. Such cell-type deconvolution involves estimating underlying cell-type fractions in a sample, which is accomplished using a DNA methylation reference panel built from sorted or single-cell DNAm data. Two competing approaches have emerged to build such reference panels, one which uses machine-learning, and another based on optimizing effect size and cell-type specificity. Here we demonstrate that the latter approach is preferable, because, owing to the relatively small number of sorted samples used in building panels, standard machine learning does not optimize effect size and cell-type specificity, causing the model to overfit and underperform when tested in independent data. Furthermore, adult blood panels built from cell-type specific hypomethylated markers improves inference when compared to panels built from hypermethylated ones. These insights provide important guidelines for optimizing the construction of future DNAm reference panels. To aid this task, we have added a function for building an optimized DNAm reference panel to our EpiDISH R-package.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Explainable deep neural networks for predicting sample phenotypes from single-cell transcriptomics 95%
- Systematic evaluation of cell-type deconvolution pipelines for sequencing-based bulk DNA methylomes 95%
- Systematic evaluation of transcriptomics-based deconvolution methods and references using thousands of clinical samples 93%
Similar papers in this journal
- A systematic evaluation of 41 DNA methylation predictors across 101 data preprocessing and normalization strategies highlights considerable variation in algorithm performance 94%
- Prime-seq, efficient and powerful bulk RNA-sequencing 94%
- Benchmark of cellular deconvolution methods using a multi-assay reference dataset from postmortem human prefrontal cortex 94%
Similar papers in this journal
- Epigenetic and transcriptomic reprogramming in monocytes of severe COVID-19 patients reflects alterations in myeloid differentiation and the influence of inflammatory cytokines 94%
- Discovery of CD80 and CD86 as recent activation markers on regulatory T cells by protein-RNA single-cell analysis 94%
- Diagnostic Evidence GAuge of Single cells (DEGAS): A flexible deep-transfer learning framework for prioritizing cells in relation to disease 94%
Similar papers in this journal
- Characterizing the properties of bisulfite sequencing data: maximizing power and sensitivity to identify between-group differences in DNA methylation 94%
- Copy number normalization distinguishes differential signals driven by copy number differences in ATAC-seq and ChIP-seq 93%
- Comparing methylation levels assayed in GC-rich regions with current and emerging methods 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.