Leveraging the largest harmonized epigenomic data collection for metadata prediction validated and augmented over 350,000 public epigenomic datasets
Raby, J.; Frosi, G.; White, F.; Laperle, J.; Jacques, P.-E.
Show abstract
Epigenomic data found in public databases often suffer from issues of non-standardization and incompleteness in their associated metadata. There are currently no automated approaches to validate or correct missing or inaccurate information listed in databases. To tackle this challenge, we harnessed the extensive harmonized data and metadata provided by the EpiATLAS project of the International Human Epigenome Consortium (IHEC) to train EpiClass, a suite of machine learning classifiers that can predict key metadata ([~]98% accuracy), including experimental assay, donor sex, biospecimen and sample cancer status. The development of these classifiers enabled the identification of a few mislabeled and low-quality datasets in the EpiATLAS project, while also completing with high-confidence most of the missing metadata. These classifiers were also validated on ENCODE datasets absent from the initial training, then applied to assess more than 350,000 human ChIP-Seq and RNA-Seq datasets from public repositories. Overall, this effort not only validated the accuracy of the vast majority of assays reported by the original authors, but also unveiled [~]500 datasets with discrepancies, in particular through data swap within series of experiments. More importantly, EpiClass also supplied high-confidence predictions for over 320,000 metadata attributes of the biological sample such as the sex, cancer status and biomaterial type, which had been originally omitted in the majority of cases. Our work introduces the first systematic approach for metadata correction and augmentation, enhancing the quality and reliability of publicly available epigenomic data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PACS allows comprehensive dissection of multiple factors governing chromatin accessibility from snATAC-seq data 97%
- Cross-dataset pan-cancer detection: Correlating cell-free DNA fragment coverage with open chromatin sites across cell types 97%
- Landscape of allele-specific transcription factor binding in the human genome 96%
Similar papers in this journal
Similar papers in this journal
- An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced TF binding 96%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 96%
- Deep-learning prediction of gene expression from personal genomes 96%
Similar papers in this journal
Similar papers in this journal
- Coralysis enables sensitive identification of imbalanced cell types and states in single-cell data via multi-level integration 97%
- SPOTlight:Seeded NMF regression to Deconvolute Spatial Transcriptomics Spots with Single-Cell Transcriptomes 96%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.