GigaScience
◐ Oxford University Press (OUP)
All preprints, ranked by how well they match GigaScience's content profile, based on 212 papers previously published here. The average preprint has a 0.16% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Tandon, D.; Mendes de Farias, T.; Allard, P.-M.; Defossez, E.
Show abstract
BackgroundIn recent years, biodiversity data management has emerged as a critical pillar in global conservation efforts. Today, the ability to efficiently collect, structure, and analyze biodiversity data is central to breakthroughs in conservation, drug development, disease monitoring, ecological forecasting, and agri-tech innovation. However, due to the vastness and heterogeneity of biodiversity data, it is often confined to databases for specific research areas in isolated formats and disconnected from other relevant resources. Crucial components of such data in kingdom Plantae comprise of metabolomes - the vast array of compounds produced by plants; traits - measurable characteristics of plants that influence their growth, survival, and reproduction, and that affect ecosystem processes; and biotic interactions - relationships of plants with other living organisms, affecting the ecosystem functions. ResultsIn this work, we present METRIN-KG (MEtabolomes, TRaits, and INteractions-Knowledge Graph) a powerful data structure simplifying the integration of diverse and heterogeneous data resources such as plant metabolomes, traits and biotic interactions. ConclusionsThe proposed knowledge graph provides an interface to interactively search for data relating plant metabolomes, traits and interactions. This, in turn, will facilitate development of research questions in life-sciences. In this context, we provide representative case studies on how to frame queries that can be used to search for relevant data in the knowledge graph.
Jacob, D.; Ehrenmann, F.; David, R.; Tran, J.; Mirande-Ney, C.; Chaumeil, P.
Show abstract
BackgroundDescriptive metadata are crucial for the discovery, reporting and mobilisation of research datasets. Addressing all metadata issues within the Data Management Plan often poses challenges for data producers. Organising and documenting data within data storage entails creating various descriptive metadata. Subsequently, data sharing involves ensuring metadata interoperability in alignment with FAIR principles. Given the tangible nature of these challenges, a real need for management tools has to be addressed to assist data managers to the fullest extent. Moreover, these tools have to meet data producers requirements and be user-friendly as well with minimal training as prerequisites. ResultsWe developed Maggot which stands for Metadata Aggregation on Data Storage, specifically designed to annotate datasets by generating metadata files to be linked into storage spaces. Maggot enables users to seamlessly generate and attach comprehensible metadata to datasets within a collaborative environment. This approach seamlessly integrates into a data management plan, effectively tackling challenges related to data organisation, documentation, storage, and frictionless FAIR metadata sharing within the collaborative group and beyond. Furthermore, for enabling metadata crosswalk, metadata generated with Maggot can be converted for a specific data repository or configured to be exported into a suitable format for data harvesting by third-party applications. ConclusionThe primary feature of Maggot is to ease metadata capture based on a carefully selected schema and standards. Then, it greatly eases access to data through metadata as requested nowadays in projects funded by public institutions and entities such as Europe Commission. Thus, Maggot can be used on one hand to promote good local versus global data management with open data sharing in mind while respecting FAIR principles, and on the other hand to prepare the future EOSC FAIR Web of Data within the framework of the European Open Science Cloud.
Ibeh, N.; Feigin, C. Y.; Frankenberg, S. R.; McCarthy, D. J.; Pask, A. J.; Gallego Romero, I.
Show abstract
Marsupials exhibit highly specialized patterns of reproduction and development, making them uniquely valuable for comparative genomics studies with their sister lineage, eutherian (also known as placental) mammals. However, marsupial genomic resources still lag far behind those of eutherian mammals, limiting our insight into mammalian diversity. Here, we present a series of novel genomic resources for the fat-tailed dunnart (Sminthopsis crassicaudata), a mouse-like marsupial that, due to its ease of husbandry and ex-utero development, is emerging as a laboratory model. To enable wider use, we have generated a multi-tissue de novo transcriptome assembly of dunnart RNA-seq reads spanning 12 tissues. This highly representative transcriptome is comprised of 2,093,982 assembled transcripts, with a mean transcript length of 830 bp. The transcriptome mammalian BUSCO completeness score of 93% is the highest amongst all other published marsupial transcriptomes. Additionally, we report an improved fat-tailed dunnart genome assembly which is 3.23 Gb long, organized into 1,848 scaffolds, with a scaffold N50 of 72.64 Mb. The genome annotation, supported by assembled transcripts and ab initio predictions, revealed 21,622 protein-coding genes. Altogether, these resources will contribute greatly towards characterizing marsupial biology and mammalian genome evolution.
Parvulescu, L.; Livadariu, D.; Bacu, V. I.; Nandra, C. I.; Stefanut, T. T.; World of Crayfish Contributors,
Show abstract
BackgroundSpecies occurrence records represent the backbone of biodiversity science, yet their utility is often limited to spatial analyses, distribution maps, or presence-absence models. Current biodiversity infrastructures rarely provide computational formats directly usable by modern artificial intelligence (AI) systems, such as large language models (LLMs), which increasingly mediate scientific communication and knowledge synthesis. Open frameworks that convert biodiversity occurrences into structured, machine-accessible, provenance-rich knowledge are therefore essential--particularly those enabling rapid integration of new records, near real-time generation of spatial metrics, and production of both human-interpretable reports and AI-consumable outputs. Such capabilities substantially reduce latency between data acquisition and decision support, while ensuring biodiversity knowledge remains traceable and verifiable in AI-mediated workflows. ResultsWe introduce cheCkOVER, an open framework that converts raw species occurrence datasets into standardized, API-ready, multi-layered outputs: biogeographic descriptors, dynamic distribution maps, summary metrics, and structured JSON geo-narratives following a canonical template. The framework stratifies processing by population origin (indigenous vs. non-indigenous), enabling IUCN-aligned conservation metrics while simultaneously tracking invasion dynamics. Each output embeds standardized citation metadata ensuring full provenance traceability. We applied the pipeline to 111,729 validated crayfish (Astacidea) occurrence records from 465 species, generating comprehensive species packages including indigenous-range classifications (171 endemic, 287 regional, 5 cosmopolitan taxa) and non-indigenous range tracking for 30 invasive species. This proof-of-concept demonstrates how the framework transforms minimal datapoints--validated species occurrences--into interoperable knowledge consumable by both humans and computational systems. The JSON outputs are optimized for retrieval-augmented generation, enabling AI systems to dynamically access and cite biodiversity knowledge with explicit source attribution. ConclusionscheCkOVER is taxon-agnostic and establishes a reproducible pathway from biodiversity occurrences to narrative-ready, AI-interoperable knowledge with immediate public utility via the World of Crayfish(R) platform (https://world.crayfish.ro/), where each species page integrates structured outputs. The open-source framework (GPL-3) combines a generalizable processing pipeline with taxon-specific knowledge products, enabling flexible reuse across conservation research, policy reporting, and AI-driven applications. This minimalist-to-complex design extends the reach of biodiversity data beyond traditional analyses, positioning occurrence repositories as active knowledge engines for next-generation biodiversity informatics. Significance statementBiodiversity infrastructures remain underused by modern AI systems despite their central role in science and society. cheCkOVER embodies a minimalist-to-complex paradigm: from the validated geographic occurrence of a species--a datapoint often perceived as trivial--it derives structured, multi-layered outputs linking distribution, conservation status, and standardized geographic indicators. These outputs are natively optimized for retrieval-augmented generation and other machine-consumable workflows, enabling AI systems to dynamically access and cite biodiversity knowledge with maintained provenance beyond their pre-training corpora. Using a global crayfish dataset as proof-of-concept, we demonstrate how raw occurrence records can scale into rich, interoperable biogeographic knowledge products with immediate value for both human experts and computational systems. This positions biodiversity databases as critical knowledge engines for next-generation science, policy, and societal decision-making, providing standardized outputs directly incorporable into conservation evaluation workflows where transparent, reproducible, and provenance-rich occurrence-based metrics are essential.
Shih, C. C.; Chen, J.; Lee, A. S.; Bertin, N.; Hebrard, M.; Khor, C. C.; Li, Z.; Tan, J. H. J.; Meah, W. Y.; Peh, S. Q.; Mok, S. Q.; Sim, K. S.; Liu, J.; Wang, L.; Wong, E.; Li, J.; Tin, A.; Cheng, C.-Y.; Heng, C.-K.; Yuan, J.-M.; Koh, W.-P.; Saw, S. M.; Friedlander, Y.; Sim, X.; Chai, J. F.; Chong, Y. S.; Davila, S.; Goh, L. L.; Lee, E. S.; Wong, T. Y.; Karnani, N.; Leong, K. P.; Yeo, K. K.; Chambers, J. C.; Goh, R. S. M.; Tan, P.; Dorajoo, R.
Show abstract
Genomic researchers are increasingly utilizing commercial cloud platforms (CCPs) to manage their data and analytics needs. Commercial clouds allow researchers to grow their storage and analytics capacity on demand, keeping pace with expanding project data footprints and enabling researchers to avoid large capital expenditures while paying only for IT capacity consumed by their project. Cloud computing also allows researchers to overcome common network and storage bottlenecks encountered when combining or re-analysing large datasets. However, cloud computing presents a new set of challenges. Without adequate security controls, the risk of unauthorised access may be higher for data stored on the cloud. In addition, regulators are increasingly mandating data access patterns and specific security protocols on the storage and use of genomic data to safeguard rights of the study participants. While CCPs provide tools for security and regulatory compliance, utilising these tools to build the necessary controls required for cloud solutions is not trivial as such skill sets are not commonly found in a genomics lab. The Research Assets Provisioning and Tracking Online Repository (RAPTOR) by the Genome Institute of Singapore is a cloud native genomics data repository and analytics platform focusing on security and regulatory compliance. Using a "five-safes" framework (Safe Purpose, Safe People, Safe Settings, Safe Data and Safe Output), RAPTOR provides security and governance controls to data contributors and users leveraging cloud computing for sharing and analysis of large genomic datasets without the risk of security breaches or running afoul of regulations. RAPTOR can also enable data federation with other genomic data repositories using GA4GH community-defined standards, allowing researchers to boost the statistical power of their work and overcome geographic and ancestry limitations of data sets
Mansueto, L.; Kretzschmar, T.; Mauleon, R.; King, G. J.
Show abstract
Global changes in Cannabis legislation after decades of stringent regulation, and heightened demand for its industrial and medicinal applications have spurred recent genetic and genomics research. An international research community emerged and identified the need for a web portal to host Cannabis-specific datasets that seamlessly integrates multiple data sources and serves omics-type analyses, fostering information sharing. The Tripal platform was used to host public genome assemblies, gene annotations, QTL and genetic maps, gene and protein expression, metabolic profile and their sample attributes. SNPs were called using public resequencing datasets on three genomes. Additional applications, such as SNP-Seek and MapManJS, were embedded into Tripal. A multi-omics data integration web-service API, developed on top of existing Tripal modules, returns generic tables of sample, property, and values. Use-cases demonstrate the APIs utility for various -omics analyses, enabling researchers to perform multi- omics analyses efficiently.
Smith, B. S.; Smith, L. A.; Lee, J.-H.; Cahill, J. A.; Graim, K.
Show abstract
A plethora of studies have identified shared molecular mechanisms involved in tumor development across humans and other mammalian species. While these two-species analyses advance understanding of human disease, extending them across many species would provide evolutionary insight into molecular mechanisms driving human cancers. However, this expansion requires knowledge transfer and harmonization across species. Genomic differences between species, including variation in genome annotation quality, have historically hindered multi-species large-scale atlas creation. To overcome these challenges, we present Paipu, a comprehensive pipeline designed to streamline querying, preprocessing, harmonization, and retrieval of large-scale RNA-seq data and associated metadata from the NCBI Sequence Read Archive (SRA). Paipu facilitates multi-species analysis by creating a harmonized atlas from user-defined search terms and species. It consists of three components: reference genome preparation, SRA metadata retrieval, and RNA-seq data processing. We apply Paipu to 188 cancer-related terms in 239 non-human mammalian species, creating a harmonized atlas of 3,484 RNA-seq samples spanning 17 species and 35 cancers. This pan-mammalian pan-cancer atlas enables myriad comparative genomics analyses that leverage genetic variation to better understand rare human cancers. As such, Paipu serves as a resource for cross-species cancer genomics and supports atlas creation for any set of species and search terms. Graphical Abstract
Imker, H. J.; Ou, H.
Show abstract
This study aimed to evaluate the names of life science data resources and consider the impacts on findability, a core feature of the FAIR (Findability, Accessibility, Interoperability, and Reusability) Principles. Utilizing a previously published list of unique data resources, we identified and validated data resources with both common and full names available (n = 1153). From this set, we analyzed characteristics of resource names to identify if any naming conventions have emerged organically. Additionally, since common names are often used in the absence of a resources full name, we performed a test to evaluate our ability to infer any meaning from common names. Our results highlight suboptimal naming practices and a wide-spread opaqueness in common names, which poses challenges to resource identification and retrieval by both human-and computationally-centric methods. These results are informative for those who establish and promote data resources as well as for those who search for data to use in individual research projects, develop data discovery systems, analyze the scientific literature, or assess research infrastructure. The findings underscore the value of findability in the FAIR Principles and the current efforts to develop infrastructure that supports more efficient communication and global connectedness.
Clark, T.; Caufield, H.; Mohan, J. A.; Al Manir, S.; Amorim, E.; Eddy, J.; Gim, N.; Gow, B.; Goar, W.; Haendel, M.; Hansen, J. N.; Harris, N.; Hermjakob, H.; McWeeney, S. K.; Nebeker, C.; Nikolov, M.; Shaffer, J.; Sheffield, N.; Sheynkman, G.; Stevenson, J.; Mungall, C.; Chen, J. Y.; Wagner, A.; Kong, S. W.; Ghosh, S. S.; Patel, B.; Williams, A.; Munoz-Torres, M. C.
Show abstract
Biomedical research and clinical practice are in the midst of a transition toward significantly increased use of artificial intelligence (AI) and machine learning (ML) methods. These advances promise to enable qualitatively deeper insight into complex challenges formerly beyond the reach of analytic methods and human intuition while placing increased demands on ethical and explainable artificial intelligence (XAI), given the opaque nature of many deep learning methods. The U.S. National Institutes of Health (NIH) has initiated a significant research and development program, Bridge2AI, aimed at producing new "flagship" datasets designed to support AI/ML analysis of complex biomedical challenges, elucidate best practices, develop tools and standards in AI/ML data science, and disseminate these datasets, tools, and methods broadly to the biomedical community. An essential set of concepts to be developed and disseminated in this program along with the data and tools produced are criteria for AI-readiness of data, including critical considerations for XAI and ethical, legal, and social implications (ELSI) of AI technologies. NIH Bridge to Artificial Intelligence (Bridge2AI) Standards Working Group members prepared this article to present methods for assessing the AI-readiness of biomedical data and the data standards perspectives and criteria we have developed throughout this program. While the field is rapidly evolving, these criteria are foundational for scientific rigor and the ethical design and application of biomedical AI methods.
Strudwick, J.; Gardiner, L.-J.; Denning-James, K.; Haiminen, N.; Evans, A.; Kelly, J.; Madgwick, M.; Utro, F.; Seabolt, E.; Gibson, C.; Bedi, B.; Clayton, D.; Howell, C.; PARIDA, L.; Carrieri, A. P.
Show abstract
Machine learning (ML) methods have the potential of detailed insights of complex biological systems and today are increasingly used to analyse omics data for tasks such as the discovery of novel biomarkers and phenotype prediction. It can be extremely beneficial and powerful for scientists, domain experts, to easily run sophisticated, robust, and interpretable ML pipelines without the need for an in depth understanding of the code needed to train, tune, optimise ML algorithms. They can then focus on the biological interpretation and validation of the results and insights generated by ML models. Here, we present an entirely automated open-source explainable AI tool, AutoXAI4Omics, that performs classification and regression tasks from omics and tabular numerical data. AutoXAI4Omics accelerates scientific discovery by automating processes and decisions made by AI experts, e.g., selection of the best feature set, hyper-tuning of different ML algorithms and selection of the best ML model for a specific task and dataset. Prior to ML analysis AutoXAI4Omics incorporates feature filtering options that are tailored to specific omic data types. Moreover, the insights into the predictions that are provided by the tool through explainability analysis highlight associations between omic feature values and the targets under investigation e.g., predicted phenotypes, facilitating the discovery of actionable insights. AutoXAI4Omics is at: https://github.com/IBM/AutoXAI4Omics. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=188 HEIGHT=200 SRC="FIGDIR/small/586460v1_ufig1.gif" ALT="Figure 1"> View larger version (36K): org.highwire.dtl.DTLVardef@366327org.highwire.dtl.DTLVardef@a7b559org.highwire.dtl.DTLVardef@7319aforg.highwire.dtl.DTLVardef@9b5030_HPS_FORMAT_FIGEXP M_FIG C_FIG
Nieminen, M.; Stolpe, O.; Kuhring, M.; Weiner, J.; Pett, P.; Beule, D.; Holtgrewe, M.
Show abstract
Scientists employing omics in life science studies face challenges such as the modeling of multi assay studies, recording of all relevant parameters, and managing many samples with their metadata. They must manage many large files that are the results of the assays or subsequent computation. Users with diverse backgrounds, ranging from computational scientists to wet-lab scientists, have dissimilar needs when it comes to data access, with programmatic interfaces being favored by the former and graphical ones by the latter. We introduce SODAR, the system for omics data access and retrieval. SODAR is a software package that addresses these challenges by providing a web-based graphical user interface for managing multi assay studies and describing them using the ISA (Investigation, Study, Assay) data model and the ISA-Tab file format. Data storage is handled using the iRODS data management system, which handles large quantities of files and substantial amounts of data. SODAR also offers programmable APIs and command line access for metadata and file storage. SODAR supports complex omics integration studies and can be easily installed. The software is written in Python 3 and freely available at https://github.com/bihealth/sodar-server under the MIT license.
Pavlidis, P.; Mancarci, B. O.; Maximo, A.; Yan, C.; Schwartz, R. A.
Show abstract
We describe an automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource. Gemma is a hand-curated database of reprocessed transcriptomic studies, currently covering over 23,000 human, mouse and rat data sets largely drawn from the Gene Expression Omnibus (GEO). We developed a pipeline that uses both traditional (mechanical) and large-language models to produce detailed ontology-anchored, sample- and experiment-level annotations in accordance with our established curation guidelines. In this report, we describe benchmarking the pipeline and investigations aimed at evaluating readiness of the v1.1 Gemma curation agent for production use. Overall, performance is near that of human curators, at approximately 1/20th the cost and at least 100 times the speed. We also present preliminary exploration of triage methods for identifying agent curations that are more likely to contain errors, and thus can be forwarded for human review. We discuss the potential place of such curation approaches in bioinformatics ecosystems. Besides the software, our deliverables include the benchmark set of 500 studies and an evaluation framework that can be used to further develop the pipeline or compare to other approaches.
Ambroise, J.; Gatto, L.; Hurel, J.; Bearzatto, B.; Gala, J.-L.
Show abstract
On Friday, 19 March 2021, WHO organized a virtual global workshop highlighting the need for a globally coordinated plan to increase SARS-CoV-2 genetic sequencing capacities to detect SARS-CoV-2 mutations and variants, and to monitor virus genomic evolution worldwide. One week later, in another virtual meeting, it focused on sero epidemiology for SARS-CoV-2 variants of concern and variants of interest. Efficient monitoring of the virus relies on the storage, handling and sharing of the genomic data and the associated metadata. In this manuscript, we demonstrate how the Bioconductor VariantExperiment class addresses these needs, offering a robust and efficient solution to the requirements laid out by the WHO.
Jariwala, S.; Long-Fox, B. L.; Berardini, T. Z.
Show abstract
Curation of biological and paleontological datasets is a labor-intensive process that requires standardization and validation to ensure data integrity. In particular, manual curation of datasets is prone to human errors such as typographical errors, inconsistent formatting, and incomplete metadata, which hinder reproducibility and compliance with Findability, Accessibility, Interoperability, and Reusability (FAIR) principles. Artificial Intelligence (AI) offers a transformative solution for enhancing research efficiency by automating data validation, improving accuracy, and streamlining curation workflows. This study explores the integration of an AI-assisted curation tool developed for MorphoBank, an open access repository established to enhance standardization and usability of morphological character datasets. Specifically, this work presents an AI tool designed to extract, structure, and standardize morphological character data from published literature into the NEXUS file format, a widely used format for phylogenetic analyses. This tool leverages machine learning techniques, including Large Language Models (LLMs), to automate the extraction of character names and states from text in various formats, reducing manual data entry errors and improving data completeness. The system enables efficient conversion of matrix-only files into complete, machine- and human-readable datasets that include key character metadata. By automating these tasks, the tool significantly accelerates dataset curation while improving accuracy and standardization. This approach increases the FAIRness of the data and offers a scalable framework for extending AI-assisted curation to standardize other biological datasets. Our findings demonstrate the value of AI in scientific data curation and advancing data reuse in paleontology, systematics, and evolutionary biology.
Ruiz, J. L.; Terron-Camero, L. C.; Castillo-Gonzalez, J.; Fernandez-Rengel, I.; Delgado, M.; Gonzalez-Rey, E.; Andres-Leon, E.
Show abstract
SummaryIn the current context of transcriptomics democratization, there is an unprecedented surge in the number of studies and datasets. However, advances are hampered by aspects such as the reproducibility crisis, and lack of standardization, in particular with scarce reanalyses of secondary data. reanalyzerGSE, is a user-friendly pipeline that aims to be an all-in-one automatic solution for locally available transcriptomic data and those found in public repositories, thereby encouraging data reuse. With its modular and expandable design, reanalyzerGSE combines cutting-edge software to effectively address simple and complex transcriptomic studies ensuring standardization, up to date reference genome, reproducibility, and flexibility for researchers. Availability and implementationThe reanalyzerGSE open-source code and test data are freely available at both https://github.com/BioinfoIPBLN/reanalyzerGSE and 10.5281/zenodo.XXXX under the GPL3 license. Supplementary data are available.
Oziolor, E. M.; Sullivan, S.; Mangelson, H.; Eacker, S. M.; Agostino, M.; Whiteley, L. O.; Cook, J.; Koza-Taylor, P.
Show abstract
The cynomolgus macaque is a non-human primate model, heavily used in biomedical research, but with outdated genomic resources. Here we have used the latest long-read sequencing technologies in order to assemble a fully phased, chromosome-level assembly for the cynomolgus macaque. We have built a hybrid assembly with PacBio, 10x Genomics, and HiC technologies, resulting in a diploid assembly that spans a length of 5.1 Gb with a total of 16,741 contigs (N50 of 0.86Mb) contained in 370 scaffolds (N50 of 138 Mb) positioned on 42 chromosomes (21 homologous pairs). This assembly is highly homologous to former assemblies and identifies novel inversions and provides higher confidence in the genetic architecture of the cynomolgus macaque genome. A demographic estimation is also able to capture the recent genetic bottleneck in the Mauritius population, from which the sequenced individual originates. We offer this resource as an enablement for genetic tools to be built around this important model for biomedical research.
Chrysostomakis, I.; Mozer, A.; Bruno Di-Nizo, C.; Fischer, D.; Sargheini, N.; von der Mark, L.; Huettel, B.; Astrin, J. J.; Toepfer, T.; Boehne, A.
Show abstract
BackgroundReference genomes have a wide range of applications. Yet, we are from a complete genomic picture for the tree of life. We here contribute another piece to the puzzle by providing a high-quality reference genome for the Ural Owl (Strix uralensis), a species of conservation concern and efforts affected by habitat destruction and climate change. ResultsWe generated a reference genome assembly for the Ural Owl based on high-fidelity (HiFi) long reads and chromosome conformation capture (Hi-C) data. It figures amongst the best avian genome assemblies currently available (BUSCO completeness of 99.94 %). The primary assembly had a size of 1.38 Gb with a scaffold N50 of 90.1 Mb, while the alternative assembly had a size of 1.3 Gb and a scaffold N50 of 17.0 Mb. We show an exceptionally high repeat content (21.07 %) that is different from those of other bird taxa with repeat extensions. We confirm a Strix characteristic chromosomal fusion and support the observation that bird microchromosomes have a higher density of genes, associated with a reduction in gene length due to shorter introns. An analysis of gene content provides evidence of changes in the keratin gene repertoire as well as modifications of metabolism genes of owls. This opens an avenue of research if this is related to flight adaptations. The population size history of the Ural Owl decreased over long periods of time with increases during the Eemian interglacial and stable size during the last glacial period. Ever since it is declining to its currently lowest effective population size. We also investigated cell culture of progressive passages as a tool for genetic resources. Karyotyping of passages confirmed no large variants, while a SNP analysis revealed a low presence of short variants across cell passages. ConclusionsThe established reference genome is a valuable resource for ongoing conservation efforts, but also for (avian) comparative genomics research. Further research is needed to determine whether cell culture passages can be safely used in genomic research.
Webb, A. E.; Wolf, S. W.; Traniello, I. M.; Kocher, S. D.
Show abstract
The exponential growth in biological data generation has created an urgent need for efficient, reproducible computational analysis workflows. Here, we present pipemake, a computational platform designed to streamline the development and implementation of efficient and reproducible Snakemake workflows. pipemake creates modular pipelines that can be seamlessly integrated or removed from the platform without requiring reconfiguration of the core system, enabling flexible adaptation of workflows to different analytical needs across diverse fields. To demonstrate the platforms capabilities, we created and implemented pipelines to reanalyze two distinct biological datasets. First, we recreated a population genomics analysis of the socially flexible halictid bee, Lasioglossum albipes, using pipemake-generated workflows for de novo genome annotation, processing of variant data, dimensionality reduction, and a genome-wide association study (GWAS). We then used pipemake to analyze behavioral tracking data from the common eastern bumble bee, Bombus impatiens. In both cases, pipemake workflows produced results consistent with published findings while substantially reducing hands-on analysis time. Overall, pipemakes modular design allows researchers to easily modify existing pipelines or develop new ones without software development expertise. Beyond streamlining workflow creation, pipemake leverages the full Snakemake ecosystem to enable parallel processing, automated error recovery, and comprehensive analysis documentation. These features make pipemake an efficient and accessible solution for analyzing complex biological datasets. pipemake is freely available as a conda package or direct download at https://github.com/kocherlab/pipemake
Shintani, M.; Andrade, D.; Bono, H.
Show abstract
Although the Gene Expression Omnibus and other public repositories are expanding rapidly, curation across these databases has not kept pace. Data reuse is often hindered by unstandardized metadata comprising unstructured text. To address this, we developed a workflow that combines retrieval via an application programming interface with semantic filtering using large language models (LLMs) for automated curation. We benchmarked multiple LLMs using metadata from 150 candidate Arabidopsis RNA sequencing projects to classify samples treated with exogenous abscisic acid and their controls. Simple keyword searches yielded many false positives (F1=0.59); classification using LLMs significantly improved performance. Several open-weight models achieved a nearly perfect performance (F1>0.98), comparable to that of closed models. We also found that utilizing LLM confidence scores enables high-confidence cases to be processed automatically. These results suggest that open-weight LLMs can support scalable and reproducible metadata curation in local environments, providing a foundation for accelerating public dataset reuse.
Boiten, J.-W.; Azevedo, R.; van Bochove, K.; Cavelaars, M.; Dekker, A.; Fijneman, R. J. A.; Lansberg, P.; van der Linden, W.; Mons, B.; Stathonikos, N.; Verheul, H. M. W.; Beliën, J. A. M.; Meijer, G. A.
Show abstract
Translating new technology and biological findings into clinical applications is hampered by insufficient translational research IT. The Dutch Translational research IT (TraIT) initiative organizes, deploys, and manages data and workflows in an on-line "office suite", supplemented with efficient training and user support. TraIT has been adopted by a wide user community providing an excellent large-scale demonstrator for the nation-wide Health-RI initiative.