Back

GigaScience

Oxford University Press (OUP)

All preprints, ranked by how well they match GigaScience's content profile, based on 212 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
METRIN-KG: A knowledge graph integrating plant metabolites, traits and biotic interactions

Tandon, D.; Mendes de Farias, T.; Allard, P.-M.; Defossez, E.

2025-11-12 bioinformatics 10.1101/2025.08.20.671289 medRxiv
Top 0.1%
60.8%
Show abstract

BackgroundIn recent years, biodiversity data management has emerged as a critical pillar in global conservation efforts. Today, the ability to efficiently collect, structure, and analyze biodiversity data is central to breakthroughs in conservation, drug development, disease monitoring, ecological forecasting, and agri-tech innovation. However, due to the vastness and heterogeneity of biodiversity data, it is often confined to databases for specific research areas in isolated formats and disconnected from other relevant resources. Crucial components of such data in kingdom Plantae comprise of metabolomes - the vast array of compounds produced by plants; traits - measurable characteristics of plants that influence their growth, survival, and reproduction, and that affect ecosystem processes; and biotic interactions - relationships of plants with other living organisms, affecting the ecosystem functions. ResultsIn this work, we present METRIN-KG (MEtabolomes, TRaits, and INteractions-Knowledge Graph) a powerful data structure simplifying the integration of diverse and heterogeneous data resources such as plant metabolomes, traits and biotic interactions. ConclusionsThe proposed knowledge graph provides an interface to interactively search for data relating plant metabolomes, traits and interactions. This, in turn, will facilitate development of research questions in life-sciences. In this context, we provide representative case studies on how to frame queries that can be used to search for relevant data in the knowledge graph.

2
A Metadata-Driven Framework for Strengthening Pathogen Genomics Lessons from SARS-CoV-2

Pavia, M.; OCOnnor, K.; Gonzalez-Hernandez, G.; Scotch, M.

2025-11-06 genetic and genomic medicine 10.1101/2025.11.04.25339514 medRxiv
Top 0.1%
45.1%
Show abstract

During the COVID-19 pandemic, large-scale pathogen sequencing generated millions of SARS-CoV-2 genomes deposited in repositories like GenBank and GISAID. However, most of these records lack detailed patient metadata, such as demographics and clinical outcomes, which limits their utility for large-scale pathogen genomics analyses. While records that are linked to a journal publication might contain such metadata, systematic extraction and linkage to sequence records requires substantial manual effort. In this work, we assess the completeness of metadata in GenBank and demonstrate the value of enriched clinical and demographic annotations for genomic epidemiology. We found that on average GenBank records contained only 21.6% of host metadata, and during our study period [~]0.02% of published articles provided accessible sequence-specific patient metadata. Additionally, using published SARS-CoV-2 genomes and their corresponding journal articles, we constructed an analytical use case in pathogen genomics in which host stratification by clinical and demographic factors enables examination of evolutionary dynamics and clinical outcomes. Our results demonstrate how metadata-enrichment enhances pathogen genomic studies and provide a framework applicable to other pathogens.

3
Killiverse: an interactive multi-omics web resource for killifish

Mittal, A.; Singh, P. P.

2026-06-21 genomics 10.64898/2026.06.16.731504 medRxiv
Top 0.1%
44.8%
Show abstract

BackgroundKillifish have emerged as valuable vertebrate model systems for investigating several disciplines including aging, regeneration, and developmental biology. Multi-omics datasets are increasingly being generated for killifish. However, their reuse remains limited due to computational challenges, largely due to the lack of accessible resources creating a bottleneck in widespread adoption of the killifish model. To address this, we developed Killiverse, a web resource for quick and intuitive exploration of multi-modal omics data dedicated to the model organism. ResultsKilliverse is an interactive, no-code, web-based platform designed for exploration of killifish multi-omics data. The platform aggregates a growing list of datasets including bulk transcriptomes, single-cell and single-nucleus transcriptomes, proteomes, and lipidomes processed through standardized pipelines and genome assemblies. Killiverse supports customized visualization and enables cross-study and cross-species analysis. It provides ortholog mapping to several established model organisms. By combining low-code software development with modern cloud technologies, the platform delivers a scalable browser-accessible application for the community. ConclusionsKilliverse enables rapid hypothesis development through the identification of patterns across studies and species. The ortholog maps allow the findings to be placed in a broader biological context. The platform represents an innovation in genomics data visualization that will serve as a template for future tool development. Killiverse is freely accessible at https://killiverse.org/.

4
Maggot: An ecosystem for sharing metadata within the web of FAIR Data

Jacob, D.; Ehrenmann, F.; David, R.; Tran, J.; Mirande-Ney, C.; Chaumeil, P.

2024-05-29 bioinformatics 10.1101/2024.05.24.595703 medRxiv
Top 0.1%
39.5%
Show abstract

BackgroundDescriptive metadata are crucial for the discovery, reporting and mobilisation of research datasets. Addressing all metadata issues within the Data Management Plan often poses challenges for data producers. Organising and documenting data within data storage entails creating various descriptive metadata. Subsequently, data sharing involves ensuring metadata interoperability in alignment with FAIR principles. Given the tangible nature of these challenges, a real need for management tools has to be addressed to assist data managers to the fullest extent. Moreover, these tools have to meet data producers requirements and be user-friendly as well with minimal training as prerequisites. ResultsWe developed Maggot which stands for Metadata Aggregation on Data Storage, specifically designed to annotate datasets by generating metadata files to be linked into storage spaces. Maggot enables users to seamlessly generate and attach comprehensible metadata to datasets within a collaborative environment. This approach seamlessly integrates into a data management plan, effectively tackling challenges related to data organisation, documentation, storage, and frictionless FAIR metadata sharing within the collaborative group and beyond. Furthermore, for enabling metadata crosswalk, metadata generated with Maggot can be converted for a specific data repository or configured to be exported into a suitable format for data harvesting by third-party applications. ConclusionThe primary feature of Maggot is to ease metadata capture based on a carefully selected schema and standards. Then, it greatly eases access to data through metadata as requested nowadays in projects funded by public institutions and entities such as Europe Commission. Thus, Maggot can be used on one hand to promote good local versus global data management with open data sharing in mind while respecting FAIR principles, and on the other hand to prepare the future EOSC FAIR Web of Data within the framework of the European Open Science Cloud.

5
De novo transcriptome assembly and genome annotation of the fat-tailed dunnart (Sminthopsis crassicaudata)

Ibeh, N.; Feigin, C. Y.; Frankenberg, S. R.; McCarthy, D. J.; Pask, A. J.; Gallego Romero, I.

2023-11-17 genomics 10.1101/2023.11.16.567318 medRxiv
Top 0.1%
39.1%
Show abstract

Marsupials exhibit highly specialized patterns of reproduction and development, making them uniquely valuable for comparative genomics studies with their sister lineage, eutherian (also known as placental) mammals. However, marsupial genomic resources still lag far behind those of eutherian mammals, limiting our insight into mammalian diversity. Here, we present a series of novel genomic resources for the fat-tailed dunnart (Sminthopsis crassicaudata), a mouse-like marsupial that, due to its ease of husbandry and ex-utero development, is emerging as a laboratory model. To enable wider use, we have generated a multi-tissue de novo transcriptome assembly of dunnart RNA-seq reads spanning 12 tissues. This highly representative transcriptome is comprised of 2,093,982 assembled transcripts, with a mean transcript length of 830 bp. The transcriptome mammalian BUSCO completeness score of 93% is the highest amongst all other published marsupial transcriptomes. Additionally, we report an improved fat-tailed dunnart genome assembly which is 3.23 Gb long, organized into 1,848 scaffolds, with a scaffold N50 of 72.64 Mb. The genome annotation, supported by assembled transcripts and ab initio predictions, revealed 21,622 protein-coding genes. Altogether, these resources will contribute greatly towards characterizing marsupial biology and mammalian genome evolution.

6
cheCkOVER: An open framework and AI-ready global crayfish database for next-generation biodiversity knowledge

Parvulescu, L.; Livadariu, D.; Bacu, V. I.; Nandra, C. I.; Stefanut, T. T.; World of Crayfish Contributors,

2025-12-30 bioinformatics 10.64898/2025.12.29.696807 medRxiv
Top 0.1%
38.4%
Show abstract

BackgroundSpecies occurrence records represent the backbone of biodiversity science, yet their utility is often limited to spatial analyses, distribution maps, or presence-absence models. Current biodiversity infrastructures rarely provide computational formats directly usable by modern artificial intelligence (AI) systems, such as large language models (LLMs), which increasingly mediate scientific communication and knowledge synthesis. Open frameworks that convert biodiversity occurrences into structured, machine-accessible, provenance-rich knowledge are therefore essential--particularly those enabling rapid integration of new records, near real-time generation of spatial metrics, and production of both human-interpretable reports and AI-consumable outputs. Such capabilities substantially reduce latency between data acquisition and decision support, while ensuring biodiversity knowledge remains traceable and verifiable in AI-mediated workflows. ResultsWe introduce cheCkOVER, an open framework that converts raw species occurrence datasets into standardized, API-ready, multi-layered outputs: biogeographic descriptors, dynamic distribution maps, summary metrics, and structured JSON geo-narratives following a canonical template. The framework stratifies processing by population origin (indigenous vs. non-indigenous), enabling IUCN-aligned conservation metrics while simultaneously tracking invasion dynamics. Each output embeds standardized citation metadata ensuring full provenance traceability. We applied the pipeline to 111,729 validated crayfish (Astacidea) occurrence records from 465 species, generating comprehensive species packages including indigenous-range classifications (171 endemic, 287 regional, 5 cosmopolitan taxa) and non-indigenous range tracking for 30 invasive species. This proof-of-concept demonstrates how the framework transforms minimal datapoints--validated species occurrences--into interoperable knowledge consumable by both humans and computational systems. The JSON outputs are optimized for retrieval-augmented generation, enabling AI systems to dynamically access and cite biodiversity knowledge with explicit source attribution. ConclusionscheCkOVER is taxon-agnostic and establishes a reproducible pathway from biodiversity occurrences to narrative-ready, AI-interoperable knowledge with immediate public utility via the World of Crayfish(R) platform (https://world.crayfish.ro/), where each species page integrates structured outputs. The open-source framework (GPL-3) combines a generalizable processing pipeline with taxon-specific knowledge products, enabling flexible reuse across conservation research, policy reporting, and AI-driven applications. This minimalist-to-complex design extends the reach of biodiversity data beyond traditional analyses, positioning occurrence repositories as active knowledge engines for next-generation biodiversity informatics. Significance statementBiodiversity infrastructures remain underused by modern AI systems despite their central role in science and society. cheCkOVER embodies a minimalist-to-complex paradigm: from the validated geographic occurrence of a species--a datapoint often perceived as trivial--it derives structured, multi-layered outputs linking distribution, conservation status, and standardized geographic indicators. These outputs are natively optimized for retrieval-augmented generation and other machine-consumable workflows, enabling AI systems to dynamically access and cite biodiversity knowledge with maintained provenance beyond their pre-training corpora. Using a global crayfish dataset as proof-of-concept, we demonstrate how raw occurrence records can scale into rich, interoperable biogeographic knowledge products with immediate value for both human experts and computational systems. This positions biodiversity databases as critical knowledge engines for next-generation science, policy, and societal decision-making, providing standardized outputs directly incorporable into conservation evaluation workflows where transparent, reproducible, and provenance-rich occurrence-based metrics are essential.

7
RAPTOR: A Five-Safes approach to a secure, cloud native and serverless genomics data repository

Shih, C. C.; Chen, J.; Lee, A. S.; Bertin, N.; Hebrard, M.; Khor, C. C.; Li, Z.; Tan, J. H. J.; Meah, W. Y.; Peh, S. Q.; Mok, S. Q.; Sim, K. S.; Liu, J.; Wang, L.; Wong, E.; Li, J.; Tin, A.; Cheng, C.-Y.; Heng, C.-K.; Yuan, J.-M.; Koh, W.-P.; Saw, S. M.; Friedlander, Y.; Sim, X.; Chai, J. F.; Chong, Y. S.; Davila, S.; Goh, L. L.; Lee, E. S.; Wong, T. Y.; Karnani, N.; Leong, K. P.; Yeo, K. K.; Chambers, J. C.; Goh, R. S. M.; Tan, P.; Dorajoo, R.

2022-10-28 genomics 10.1101/2022.10.27.514127 medRxiv
Top 0.1%
34.8%
Show abstract

Genomic researchers are increasingly utilizing commercial cloud platforms (CCPs) to manage their data and analytics needs. Commercial clouds allow researchers to grow their storage and analytics capacity on demand, keeping pace with expanding project data footprints and enabling researchers to avoid large capital expenditures while paying only for IT capacity consumed by their project. Cloud computing also allows researchers to overcome common network and storage bottlenecks encountered when combining or re-analysing large datasets. However, cloud computing presents a new set of challenges. Without adequate security controls, the risk of unauthorised access may be higher for data stored on the cloud. In addition, regulators are increasingly mandating data access patterns and specific security protocols on the storage and use of genomic data to safeguard rights of the study participants. While CCPs provide tools for security and regulatory compliance, utilising these tools to build the necessary controls required for cloud solutions is not trivial as such skill sets are not commonly found in a genomics lab. The Research Assets Provisioning and Tracking Online Repository (RAPTOR) by the Genome Institute of Singapore is a cloud native genomics data repository and analytics platform focusing on security and regulatory compliance. Using a "five-safes" framework (Safe Purpose, Safe People, Safe Settings, Safe Data and Safe Output), RAPTOR provides security and governance controls to data contributors and users leveraging cloud computing for sharing and analysis of large genomic datasets without the risk of security breaches or running afoul of regulations. RAPTOR can also enable data federation with other genomic data repositories using GA4GH community-defined standards, allowing researchers to boost the statistical power of their work and overcome geographic and ancestry limitations of data sets

8
Building a community-driven bioinformatics platform to facilitate Cannabis sativa multi-omics research

Mansueto, L.; Kretzschmar, T.; Mauleon, R.; King, G. J.

2024-10-03 bioinformatics 10.1101/2024.10.02.616368 medRxiv
Top 0.1%
34.4%
Show abstract

Global changes in Cannabis legislation after decades of stringent regulation, and heightened demand for its industrial and medicinal applications have spurred recent genetic and genomics research. An international research community emerged and identified the need for a web portal to host Cannabis-specific datasets that seamlessly integrates multiple data sources and serves omics-type analyses, fostering information sharing. The Tripal platform was used to host public genome assemblies, gene annotations, QTL and genetic maps, gene and protein expression, metabolic profile and their sample attributes. SNPs were called using public resequencing datasets on three genomes. Additional applications, such as SNP-Seek and MapManJS, were embedded into Tripal. A multi-omics data integration web-service API, developed on top of existing Tripal modules, returns generic tables of sample, property, and values. Use-cases demonstrate the APIs utility for various -omics analyses, enabling researchers to perform multi- omics analyses efficiently.

9
Analyzing the Naming Conventions of Life Science Data Resources to Inform Human and Computational Findability

Imker, H. J.; Ou, H.

2025-10-04 bioinformatics 10.1101/2025.10.02.680112 medRxiv
Top 0.1%
32.0%
Show abstract

This study aimed to evaluate the names of life science data resources and consider the impacts on findability, a core feature of the FAIR (Findability, Accessibility, Interoperability, and Reusability) Principles. Utilizing a previously published list of unique data resources, we identified and validated data resources with both common and full names available (n = 1153). From this set, we analyzed characteristics of resource names to identify if any naming conventions have emerged organically. Additionally, since common names are often used in the absence of a resources full name, we performed a test to evaluate our ability to infer any meaning from common names. Our results highlight suboptimal naming practices and a wide-spread opaqueness in common names, which poses challenges to resource identification and retrieval by both human-and computationally-centric methods. These results are informative for those who establish and promote data resources as well as for those who search for data to use in individual research projects, develop data discovery systems, analyze the scientific literature, or assess research infrastructure. The findings underscore the value of findability in the FAIR Principles and the current efforts to develop infrastructure that supports more efficient communication and global connectedness.

10
ScriptManager: a platform for scalable and reproducible high-resolution analysis of genomics datasets

Lang, O. W.; Beer, B.; Zhang, D.; LeSon, C.; Deen, A.; Pugh, F.; Lai, W. K.

2026-06-18 bioinformatics 10.64898/2026.06.14.732163 medRxiv
Top 0.1%
31.1%
Show abstract

BackgroundThe growing diversity of genomic and epigenomic assays has driven a parallel expansion in data formats, analysis workflows, and figure-generation tools. However, tools for analyzing data and assembling publication-quality figures are often specialized to a specific assay, dramatically limiting their interoperability and reproducibility. ResultsWe present the v1.0 release of ScriptManager, a Java-based framework for modular and reproducible analysis and visualization workflows of genomics and epigenomics data. Unlike existing tools specialized for individual assay types, ScriptManager provides a unified and extensible framework for cross-assay visualization and workflow reproducibility. The v1.0 release adds novel analytical modules, GUI session logging, automated unit and integration testing, tutorials, and expanded documentation. It also integrates with the broader reproducibility ecosystem through Singularity containers, Anaconda packaging, and Galaxy XML wrappers. We demonstrate ScriptManagers TagPileup scaling from local single-core execution to a 10,305-job analysis distributed across the Open Science Grid (OSG), with the full workload completing in <2 hours of wall-clock time. ConclusionsScriptManager v1.0 enhances workflow portability, transparency, and reproducibility across a diverse range of high-resolution genomic assays. By coupling a flexible module design with modern reproducibility standards, ScriptManager provides a bridge between exploratory data analysis and formal, publication-ready figure generation. These improvements enable researchers to build, share, and reproduce genomic analyses across diverse computational infrastructures with minimal configuration.

11
Community needs for FAIR pathogen data

van Geest, G.; Thomas-Lopez, D.; Feitzinger, A. A.; Weissgold, L. A.; Halabi, S.; Cuesta, I.; Hjerde, E.; Gurwitz, K. T.; Arora, N.; Neves, A.; Palagi, P. M.; Williams, J. J.

2026-04-15 scientific communication and education 10.64898/2026.04.14.718420 medRxiv
Top 0.1%
30.9%
Show abstract

BackgroundDatasets related to infectious diseases are essential for public health decision-making, yet their reuse remains limited by persistent barriers to data sharing and integration. Achieving data that are Findable, Accessible, Interoperable, and Reusable (FAIR) is widely recognized as essential for accelerating scientific discovery and enabling coordinated responses to emerging threats, but the needs of the global pathogen data community have not been systematically characterized. AimThis study, conducted by the Pathogen Data Network (PDN), aims to identify infrastructural and educational priorities among stakeholders working with infectious disease-related data in order to guide community-responsive support for data sharing and interoperability. MethodsA cross-sectional stakeholder survey was disseminated to a well-defined expert population within PDN networks and via open professional channels. A total of 136 responses from researchers, healthcare professionals, bioinformaticians, and educators were analyzed descriptively to identify prioritized barriers, training needs, and preferred support mechanisms. ResultsRespondents consistently identified structural constraints as the primary impediments to effective data use, including limited funding (74%), data-aggregation challenges (68%), and a shortage of skilled personnel (52%). Respondents identified bioinformatics for infectious disease research (68%) as the highest priority for training, followed by guidance on using the integrated pathogen data and tools portal provided by the PDN, the Pathogens Portal (51%). The Pathogens Portal was also ranked as the most essential PDN resource (72%). Preferred training formats included virtual short courses (68%) and webinars (66%). Notably, while researchers emphasized technical subjects like machine learning, educators prioritized foundational case studies. ConclusionThese findings provide an evidence-based diagnostic of community needs and suggest that barriers to FAIR pathogen data are predominantly systemic rather than purely technological. The survey framework and openly available dataset offer a reusable template for assessing needs in other communities and regions. By aligning training, infrastructure development, and outreach with empirically identified priorities, organizations supporting infectious disease research can strengthen the interoperability and reuse of data and establish a benchmark for future community-driven improvements.

12
AI-readiness for Biomedical Data: Bridge2AI Recommendations

Clark, T.; Caufield, H.; Mohan, J. A.; Al Manir, S.; Amorim, E.; Eddy, J.; Gim, N.; Gow, B.; Goar, W.; Haendel, M.; Hansen, J. N.; Harris, N.; Hermjakob, H.; McWeeney, S. K.; Nebeker, C.; Nikolov, M.; Shaffer, J.; Sheffield, N.; Sheynkman, G.; Stevenson, J.; Mungall, C.; Chen, J. Y.; Wagner, A.; Kong, S. W.; Ghosh, S. S.; Patel, B.; Williams, A.; Munoz-Torres, M. C.

2024-10-25 bioinformatics 10.1101/2024.10.23.619844 medRxiv
Top 0.1%
30.9%
Show abstract

Biomedical research and clinical practice are in the midst of a transition toward significantly increased use of artificial intelligence (AI) and machine learning (ML) methods. These advances promise to enable qualitatively deeper insight into complex challenges formerly beyond the reach of analytic methods and human intuition while placing increased demands on ethical and explainable artificial intelligence (XAI), given the opaque nature of many deep learning methods. The U.S. National Institutes of Health (NIH) has initiated a significant research and development program, Bridge2AI, aimed at producing new "flagship" datasets designed to support AI/ML analysis of complex biomedical challenges, elucidate best practices, develop tools and standards in AI/ML data science, and disseminate these datasets, tools, and methods broadly to the biomedical community. An essential set of concepts to be developed and disseminated in this program along with the data and tools produced are criteria for AI-readiness of data, including critical considerations for XAI and ethical, legal, and social implications (ELSI) of AI technologies. NIH Bridge to Artificial Intelligence (Bridge2AI) Standards Working Group members prepared this article to present methods for assessing the AI-readiness of biomedical data and the data standards perspectives and criteria we have developed throughout this program. While the field is rapidly evolving, these criteria are foundational for scientific rigor and the ethical design and application of biomedical AI methods.

13
AutoXAI4Omics: an Automated Explainable AI tool for Omics and tabular data

Strudwick, J.; Gardiner, L.-J.; Denning-James, K.; Haiminen, N.; Evans, A.; Kelly, J.; Madgwick, M.; Utro, F.; Seabolt, E.; Gibson, C.; Bedi, B.; Clayton, D.; Howell, C.; PARIDA, L.; Carrieri, A. P.

2024-03-28 bioinformatics 10.1101/2024.03.25.586460 medRxiv
Top 0.1%
30.7%
Show abstract

Machine learning (ML) methods have the potential of detailed insights of complex biological systems and today are increasingly used to analyse omics data for tasks such as the discovery of novel biomarkers and phenotype prediction. It can be extremely beneficial and powerful for scientists, domain experts, to easily run sophisticated, robust, and interpretable ML pipelines without the need for an in depth understanding of the code needed to train, tune, optimise ML algorithms. They can then focus on the biological interpretation and validation of the results and insights generated by ML models. Here, we present an entirely automated open-source explainable AI tool, AutoXAI4Omics, that performs classification and regression tasks from omics and tabular numerical data. AutoXAI4Omics accelerates scientific discovery by automating processes and decisions made by AI experts, e.g., selection of the best feature set, hyper-tuning of different ML algorithms and selection of the best ML model for a specific task and dataset. Prior to ML analysis AutoXAI4Omics incorporates feature filtering options that are tailored to specific omic data types. Moreover, the insights into the predictions that are provided by the tool through explainability analysis highlight associations between omic feature values and the targets under investigation e.g., predicted phenotypes, facilitating the discovery of actionable insights. AutoXAI4Omics is at: https://github.com/IBM/AutoXAI4Omics. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=188 HEIGHT=200 SRC="FIGDIR/small/586460v1_ufig1.gif" ALT="Figure 1"> View larger version (36K): org.highwire.dtl.DTLVardef@366327org.highwire.dtl.DTLVardef@a7b559org.highwire.dtl.DTLVardef@7319aforg.highwire.dtl.DTLVardef@9b5030_HPS_FORMAT_FIGEXP M_FIG C_FIG

14
SODAR: enabling, modeling, and managing multi-omics integration studies

Nieminen, M.; Stolpe, O.; Kuhring, M.; Weiner, J.; Pett, P.; Beule, D.; Holtgrewe, M.

2022-08-22 bioinformatics 10.1101/2022.08.19.504516 medRxiv
Top 0.1%
30.5%
Show abstract

Scientists employing omics in life science studies face challenges such as the modeling of multi assay studies, recording of all relevant parameters, and managing many samples with their metadata. They must manage many large files that are the results of the assays or subsequent computation. Users with diverse backgrounds, ranging from computational scientists to wet-lab scientists, have dissimilar needs when it comes to data access, with programmatic interfaces being favored by the former and graphical ones by the latter. We introduce SODAR, the system for omics data access and retrieval. SODAR is a software package that addresses these challenges by providing a web-based graphical user interface for managing multi assay studies and describing them using the ISA (Investigation, Study, Assay) data model and the ISA-Tab file format. Data storage is handled using the iRODS data management system, which handles large quantities of files and substantial amounts of data. SODAR also offers programmable APIs and command line access for metadata and file storage. SODAR supports complex omics integration studies and can be easily installed. The software is written in Python 3 and freely available at https://github.com/bihealth/sodar-server under the MIT license.

15
Metadata Collector: An Open-Source Platform for Standardized Metadata Management in Multi Centre Sequencing Projects

Liguori, R.; Ferrazzi, F.

2026-06-07 bioinformatics 10.64898/2026.06.05.730314 medRxiv
Top 0.1%
30.4%
Show abstract

BackgroundNext-generation sequencing (NGS) projects generate increasingly complex metadata that are critical for reproducibility, interoperability, and compliance with FAIR principles. Nevertheless, metadata curation in multi-institutional settings often still relies on spreadsheets, manual data entry and curation, as well as non-standardized terminology. These practices frequently result in incomplete or inconsistent annotations, hinder metadata sharing, and delay submission to public repositories. ResultsWe developed Metadata Collector as a React/API/PostgreSQL web platform and deployed it on a Kubernetes cluster within a large German research consortium. The platform implements a flexible, machine-readable metadata model for experimental data and integrates customizable templates, controlled vocabularies designed to support future ontology integration, and a complete event-based versioning model. Since deployment, Metadata Collector has been used across 32 projects involving RNA-seq, scRNA-seq, ATAC-seq and multiomics datasets, representing over 700 annotated samples contributed by multiple consortium partners. The platform is designed for use by non-computational researchers as well as centralized facilities and can be integrated into existing research data management infrastructures. ConclusionsMetadata Collector embeds standardization early in the metadata lifecycle, ensuring consistent, FAIR-aligned, and reproducible metadata across distributed research groups. Its modular, open-source architecture supports both local and consortium-scale deployments and provides a foundation for future extensions, including multi-omics support and integration with laboratory information management systems and automated submission pipelines. AvailabilityOpen-source under MIT license: https://github.com/spang-lab/metadata-collector Graphical abstract of Metadata Collector O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=150 SRC="FIGDIR/small/730314v1_ufig1.gif" ALT="Figure 1"> View larger version (63K): org.highwire.dtl.DTLVardef@c7e913org.highwire.dtl.DTLVardef@9701acorg.highwire.dtl.DTLVardef@1eed3eborg.highwire.dtl.DTLVardef@9af69a_HPS_FORMAT_FIGEXP M_FIG C_FIG

16
geneXplore: An Interactive Browser for X Chromosome-Wide Association Study Results

Cook, N.; Boulais-Richard, J.; Zeng, Y.; Yang, C.; Budde, J.; Taliun, D.; Gagliano Taliun, S. A.; Cruchaga, C.; Belloy, M. E.

2026-07-14 neurology 10.64898/2026.07.14.26357489 medRxiv
Top 0.1%
27.0%
Show abstract

Summary: The X chromosome comprises approximately 5% of the human genome and encodes over 800 protein-coding genes, many of which exhibit sex-differentiated expression patterns due to escape from X chromosome inactivation (XCI) mechanisms. Despite its relevance to sex differences in complex traits, the X chromosome is routinely excluded from genome-wide association studies due to analytical challenges, and when analyzed, the impact of escape from XCI or sex is limitedly explored. No dedicated, publicly accessible browser for X chromosome-wide association study (XWAS) summary statistics currently exists, creating a barrier to systematic investigation of X-linked contributions to human traits. Here, we present geneXplore, an interactive web browser based on the PheWeb2 implementation, tailored for XWAS summary statistics across 1,944 phenotypes while distinguishing random XCI (rXCI), escape from XCI (eXCI), and sex-stratified analyses. Users can explore results via interactive plots (Manhattan and Miami, PheWAS and LocusZoom), searchable tables and access to cross-database lookup, with full summary statistics available for download. Availability and Implementation: geneXplore is freely available at https://genexplore.wustl.edu/ with no registration required and will be maintained for a minimum of two years following publication. Source code is available at https://github.com/Belloy-Lab/geneXplore_XWAS_Browser under an MIT license.

17
Machine learning-based prediction of memory requirements for metagenomic assembly in high-performance computing environments

Zierep, P. F.; Faack, S.; Beracochea, M.; Sanchez, S.; Batut, B.; Finn, R. D.; Gruening, B. A.

2026-05-13 microbiology 10.64898/2026.05.12.724571 medRxiv
Top 0.1%
27.0%
Show abstract

Metagenomic assembly can be a computationally intensive step in microbiome analysis, with memory requirements that vary widely depending on input data characteristics. In workflow systems like Galaxy and large-scale platforms like MGnify, which run thousands of heterogeneous jobs, inaccurate memory allocation drives job failures and costly retries when underestimated, and reduces throughput when overestimated. Current approaches rely primarily on heuristic rules based on input file size or sample metadata, which often fail to generalize across diverse datasets. In this study, we present a machine learning-based framework for predicting memory requirements of metagenomic assembly using metaSPAdes. We analyzed 300 assembly jobs from diverse biomes and evaluated 18 predictive models using combinations of input file size, biome classification, and sequence-derived k-mer features. K-mer profiles were computed from raw sequencing data and summarized into statistical descriptors capturing sequence complexity and diversity. Model performance was assessed using both conventional regression metrics and a production-oriented cost function that accounts for retry policies and resource waste in high-performance computing environments. Our results show that machine learning models can outperform commonly used heuristics. In particular, models incorporating biome information achieved the best performance and can be tuned to favor conservative predictions that reduce job failure rates. Simpler models based solely on input file size also performed competitively, offering a practical alternative for systems with limited feature availability. When evaluated under realistic workload distributions, predictive approaches reduced total memory waste by several million gigabyte-hours per 1,000 jobs compared to static allocation strategies. These findings demonstrate that data-driven resource prediction can substantially improve efficiency in metagenomic workflows. The proposed framework is adaptable to different computational environments and provides a foundation for integrating predictive resource allocation into large-scale bioinformatics platforms beyond Galaxy.

18
MOAflow: how re-design a pipeline with Nextflow streamlines data analysis

Tartaglia, J.; Giorgioni, M.; Cattivelli, L.; Faccioli, P.

2026-03-30 bioinformatics 10.64898/2026.03.26.713914 medRxiv
Top 0.1%
27.0%
Show abstract

BackgroundAdvances in high-throughput DNA sequencing technologies have dramatically reduced the time and cost required to generate genomic data. As sequencing is no longer a limiting factor, increasing attention must be paid to optimizing the analyses of the large-scale datasets produced. Efficient processing of such data is essential to reduce computational time and operational costs. In this context, workflow management systems (WMSs) have become key instruments for orchestrating complex bioinformatic pipelines. Among these systems, Nextflow has emerged as one of the most widely adopted solutions in bioinformatics. MethodsTo improve scalability and computational efficiency, we employed Nextflow to re-design an already existing pipeline dedicated to the analysis of MNase-defined cistrome-Occupancy (MOA-seq) data. The re-engineering process focused on modularizing the workflow and integrating containerization technologies to ensure reproducibility and easier deployment across heterogeneous computing environments. ResultsThe resulting workflow, named MOAflow, represents a modernized and fully containerized pipeline for MOA-seq data analysis. With only Docker and Nextflow required, the pipeline guarantees high portability and reproducibility. The data of the original article was used to benchmark the new pipeline. Its outputs closely match those of the original study with minor variations. ConclusionsMOAflow demonstrates how the adoption of robust WMS can substantially enhance the performance and usability of pre-existing bioinformatic pipelines. By leveraging containerization and Nextflow, it ensures consistent results across platforms while minimizing setup complexity. This work highlights the value of modern WMS-driven approaches in meeting the computational demands.

19
On the many advantages of using the VariantExperiment class to store, exchange and analyze SARS-CoV-2 genomic data and associated metadata

Ambroise, J.; Gatto, L.; Hurel, J.; Bearzatto, B.; Gala, J.-L.

2021-04-06 bioinformatics 10.1101/2021.04.05.438328 medRxiv
Top 0.1%
26.6%
Show abstract

On Friday, 19 March 2021, WHO organized a virtual global workshop highlighting the need for a globally coordinated plan to increase SARS-CoV-2 genetic sequencing capacities to detect SARS-CoV-2 mutations and variants, and to monitor virus genomic evolution worldwide. One week later, in another virtual meeting, it focused on sero epidemiology for SARS-CoV-2 variants of concern and variants of interest. Efficient monitoring of the virus relies on the storage, handling and sharing of the genomic data and the associated metadata. In this manuscript, we demonstrate how the Bioconductor VariantExperiment class addresses these needs, offering a robust and efficient solution to the requirements laid out by the WHO.

20
From expansion to consolidation: two decades ofGene Ontology evolution

Pitarch, B.; Pazos, F.; Chagoyen, M.

2026-03-06 bioinformatics 10.64898/2026.03.04.709507 medRxiv
Top 0.1%
26.4%
Show abstract

The Gene Ontology (GO) is a long-standing, community-maintained knowledge resource that underpins the functional annotation of gene products across numerous biological databases. Released regularly, GO and its associated annotations form a large, continuously evolving dataset whose temporal dynamics have direct consequences for data reuse, versioning, and reproducibility. Because analytical results derived from GO are inherently tied to specific ontology and annotation releases, a systematic understanding of how GO changes over time is essential for transparent interpretation and long-term reuse of GO-based analyses. Here, we present a comprehensive temporal characterization of the Gene Ontology and its annotations spanning 21 years of publicly available releases. Treating successive ontology and annotation versions as longitudinal research data, we quantify changes in ontology structure, term composition, relationships, and annotation content across time and across three representative annotation resources. Our analysis reveals sustained growth of GO over its lifetime, accompanied by marked structural reorganization, particularly affecting high-level, general ontology terms. Notably, across multiple structural and annotation metrics, we identify a transition toward increased stability beginning around 2017, consistent with a maturation phase of the resource. This work provides a reference framework for researchers who rely on GO releases for data integration, benchmarking, and reproducible functional analysis.