Transcriptomics data availability and reusability in the transition from microarray to next-generation sequencing
Rustici, G.; Williams, E.; Barzine, M.; Brazma, A.; Bumgarner, R.; Chierici, M.; Furlanello, C.; Greger, L.; Jurman, G.; Miller, M.; Ouellette, B. F. F.; Quackenbush, J.; Reich, M.; Stoeckert, C. J.; Taylor, R. C.; Trutane, S. C.; Weller, J.; Wilhelm, B.; Winegarden, N.
Show abstract
Over the last two decades, molecular biology has been changed by the introduction of high-throughput technologies. Data sharing requirements have prompted the establishment of persistent data archives. A standardized approach for recording and managing these data was first proposed in the Minimal Information About a Microarray Experiment (MIAME) guidelines. The Minimal Information about a high throughput nucleotide Sequencing Experiment (MINSEQE) proposal was introduced in 2008 as a logical extension of the guidelines to next-generation sequencing (NGS) technologies used for transcriptome analysis. We present a historical snapshot of the data-sharing situation focusing on transcriptomics data from both microarray and RNA-sequencing experiments published between 2009 and 2013, a period during which RNA-seq studies became increasingly popular for transcriptome analysis. We assess how much data from RNA-seq based experiments is actually available in persistent data archives, compared to data derived from microarray based experiments, and evaluate how these types of data differ. Based on this analysis, we provide recommendations to improve RNA-seq data availability, reusability, and reproducibility.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- All of gene expression (AOE): an integrated index for public gene expression databases 95%
- On taming the effect of transcript level intra-condition count variation during differential expression analysis: a story of dogs, foxes and wolves 93%
- An evaluation of RNA-seq differential analysis methods 93%
Similar papers in this journal
- The case for using Mapped Exonic Non-Duplicate (MEND) read counts in RNA-Seq experiments: examples from pediatric cancer datasets 94%
- FAIR Data Station for Lightweight Metadata Management & Validation of Omics Studies 93%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 93%
Similar papers in this journal
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 94%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 94%
- GEGA (Gallus Enriched Gene Annotation): an online tool providing genomics and functional information across 47 tissues for a chicken gene-enriched atlas gathering Ensembl & Refseq genome annotations 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.