An analytical pipeline for DNA Methylation Array Biomarker Studies
Lu, J.; Korbie, D.; Trau, M.
Show abstract
DNA methylation is one of the most commonly studied epigenetic biomarkers, due to its role in disease and development. The Illumina Infinium methylation arrays still remains the most common method to interrogate methylation across the human genome, due to its capabilities of screening over 480, 000 loci simultaneously. As such, initiatives such as The Cancer Genome Atlas (TCGA) have utilized this technology to examine the methylation profile of over 20,000 cancer samples. There is a growing body of methods for pre-processing, normalisation and analysis of array-based DNA methylation data. However, the shape and sampling distribution of probe-wise methylation that could influence the way data should be examined was rarely discussed. Therefore, this article introduces a pipeline that predicts the shape and distribution of normalised methylation patterns prior to selection of the most optimal inferential statistics screen for differential methylation. Additionally, we put forward an alternative pipeline, which employed feature selection, and demonstrate its ability to select for biomarkers with outstanding differences in methylation, which does not require the predetermination of the shape or distribution of the data of interest. AvailabilityThe Distribution test and the feature selection pipelines are available for download at: https://github.com/uqjlu8/DistributionTest
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting cancer origins with a DNA methylation-based deep neural network model 94%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 93%
- Comparative analysis of novel MGISEQ-2000 sequencing platform vs Illumina HiSeq 2500 for whole-genome sequencing 92%
Similar papers in this journal
- Minimum Error Calibration and Normalization for Genomic Copy Number Analysis 94%
- Principal component analysis- and tensor decomposition-based unsupervised feature extraction to select more reasonable differentially methylated cytosines: Optimization of standard deviation versus state-of-the-art methods 93%
- DeepPlnc: Bi-modal Deep Learning for Highly Accurate Plant lncRNA Discovery 91%
Similar papers in this journal
- LuxHMM: DNA methylation analysis with genome segmentation via Hidden Markov Model 94%
- Probabilistic modeling methods for cell-free DNA methylation based cancer classification 93%
- pyCancerSig: subclassifying human cancer with comprehensive single nucleotide, structural and microsatellite mutational signature deconstruction from whole genome sequencing 92%
Similar papers in this journal
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 92%
- Enrichment analysis on regulatory subspaces: a novel direction for the superior description of cellular responses to SARS-CoV-2 91%
- A Novel Tool for Multi-Omics Network Integration and Visualization: A Study of Glioma Heterogeneity 91%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 93%
- Discovering Key Transcriptomic Regulators in Pancreatic Ductal Adenocarcinoma using Dirichlet Process Gaussian Mixture Model 92%
- High tissue-specificity of lncRNAs maximises the prediction of tissue of origin of circulating DNA 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.