A collaborative submission model for building high-quality data resources at scale through partnership
Hilton, J. A.; Chaffer, J.; Chien, J.; Gabdank, I.; Mott, B.; Rutherford, E.; Small, C.; Zamanian, J.; Aevermann, B.; Cherry, J. M.; Klein, T. E.
Show abstract
Community data resources that aggregate datasets across studies are critical infrastructure for modern biomedical research, enabling large-scale analysis and the development of Artificial Intelligence (AI) models. However, building these resources involves a fundamental tension: the desire for a large corpus is often at odds with the need for richness and quality in both data and metadata. We detail how the collaborative submission model - where data contributors partner with dedicated resource curators - has enabled CZ CELLxGENE Discover to become a rapidly growing, widely used community resource for training and testing AI models, performing integrative analysis, validating findings, and generating hypotheses. This partnership leverages contributors intimate study knowledge and curators focus on data reuse and expertise in standardization to improve data quality, metadata accuracy, and contextual richness. This is achieved by motivating researcher participation through tangible benefits while minimizing submission burden. We contrast this collaborative model with contributor-driven and resource-driven approaches, highlighting tradeoffs in scalability, quality assurance, and sustainability. The principles and practices we describe provide a framework for building sustainable, high-quality community data resources across diverse biological data types.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Linking big biomedical datasets to modular analysis with Portable Encapsulated Projects 95%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 94%
- PEPhub: a database, web interface, and API for editing, sharing, and validating biological sample metadata 94%
Similar papers in this journal
- GenoTools: An Open-Source Python Package for Efficient Genotype Data Quality Control and Analysis 91%
- A Comprehensive Benchmarking Study on Computational Tools for Cross-omics Label Transfer from Single-cell RNA to ATAC Data 90%
- MOKA: A pipeline for multi-omics bridged SNP-set kernel association test 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.