bMINTY: Enabling Reproducible Management of High-Throughput Sequencing Analysis Results and their Metadata
Kapelios, K.; Xiropotamos, P.; Manousaki, H.; Sinnis, C.; Kotsira, V.; Dalamagas, T.; GEORGAKILAS, G. K.
Show abstract
Due to the large scale of high-throughput sequencing data generation, the community and publishers have established standards for the dissemination of studies that produce and analyze these data. Despite efforts towards Findable, Accessible, Interoperable and Reproducible (FAIR) science, critical obstacles remain. Best practices are not consistently enforced by scientific publishers, and when they are, essential information is fragmented across the methods section, supplementary materials, and public repositories. When attempting to reproduce scientific findings or reuse published data or analyses, researchers often avoid analyzing sequencing data from the ground up. Instead, they prefer to start directly from the post-sequence-alignment information (e.g., gene expression matrices in transcriptomics). However, existing repositories and workflow-oriented solutions rarely provide a single, portable, queryable resource that integrates this information with the metadata required for downstream reuse. We introduce bMINTY, a locally deployed web application with an intuitive user interface, for structured management of post-alignment workflow data outputs. bMINTY supports metadata for studies, assays, and analysis assets, including workflows, genome assemblies, genomic intervals, and cell-level entities for single-cell assays. Users may export query results in RO-Crate format, providing machine readable data packages and metadata. To the best of current knowledge, bMINTY is the first framework to bundle all this information in publication-ready, portable packaging designed for reuse. These packages can be included as supplementary material with each publication, accompanied by analysis code deposited in public repositories for downstream ad hoc analyses. Together, these practices can promote transparency, efficient reuse of published data, and support FAIR-aligned scientific reproducibility.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepSpaceDB: a spatial transcriptomics atlas for interactive in-depth analysis of tissues and tissue microenvironments 93%
- Datanator: an integrated database of molecular data for quantitatively modeling cellular behavior 93%
- SVCROWS: A User-Defined Tool for Interpreting Significant Structural Variants in Heterogeneous Datasets 92%
Similar papers in this journal
Similar papers in this journal
- Accelerated nanopore basecalling with SLOW5 data format 96%
- FinaleDB: a browser and database of cell-free DNA fragmentation patterns 95%
- Differential Expression Gene Explorer (DrEdGE): A tool for generating interactive online data visualizations for exploration of quantitative transcript abundance datasets 94%
Similar papers in this journal
- The 4D Nucleome Data Portal: a resource for searching and visualizing curated nucleomics data 97%
- Comprehensive generation, visualization, and reporting of quality control metrics for single-cell RNA sequencing data 95%
- Orchestrating and sharing large multimodal data for transparent and reproducible research 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.