Back

GigaScience

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match GigaScience's content profile, based on 212 papers previously published here. The average preprint has a 0.16% match score for this journal, so anything above that is already an above-average fit.

1
Detailed curation of biological samples and experimental designs for genomics using LLM-supported agentic workflows

Pavlidis, P.; Mancarci, B. O.; Maximo, A.; Yan, C.; Schwartz, R. A.

2026-08-02 bioinformatics 10.64898/2026.07.30.741874 medRxiv
Top 0.1%
30.2%
Show abstract

We describe an automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource. Gemma is a hand-curated database of reprocessed transcriptomic studies, currently covering over 23,000 human, mouse and rat data sets largely drawn from the Gene Expression Omnibus (GEO). We developed a pipeline that uses both traditional (mechanical) and large-language models to produce detailed ontology-anchored, sample- and experiment-level annotations in accordance with our established curation guidelines. In this report, we describe benchmarking the pipeline and investigations aimed at evaluating readiness of the v1.1 Gemma curation agent for production use. Overall, performance is near that of human curators, at approximately 1/20th the cost and at least 100 times the speed. We also present preliminary exploration of triage methods for identifying agent curations that are more likely to contain errors, and thus can be forwarded for human review. We discuss the potential place of such curation approaches in bioinformatics ecosystems. Besides the software, our deliverables include the benchmark set of 500 studies and an evaluation framework that can be used to further develop the pipeline or compare to other approaches.

2
A Standardized Methodology for FAIRness Assessment and Multi-Dimensional Scoring in Agrosystem Research Data Infrastructures

Haleem, A. U.; Arend, D.; Etukala, J. R.; Mazon, E. R.; Schmidt, M.; Jung, J.; Martini, D.; Usadel, B.; Neidiger, C.; Ulrich, R.; Lange, M.

2026-07-22 bioinformatics 10.64898/2026.07.17.738915 medRxiv
Top 0.1%
22.4%
Show abstract

The NFDI-consortium FAIRagro has established a systematic framework for evaluating the FAIRness of Research Data Infrastructures (RDI) within the German agrosystem research landscape. While FAIR principles are widely accepted, their practical implementation by RDIs remains challenging. By operationalizing the FAIR principles into a reproducible multi-dimensional scoring methodology, this initiative addresses the critical need for a transparent and citable benchmark of RDIs that moves beyond simple compliance. This paper details the underlying assessment criteria, comprising 20 aggregated core metrics, the iterative community-driven validation process, and the integration of these metrics into the FAIRagro Search Hub. This framework evaluates RDIs, like repositories or databases, instead of sampling hosted data sets, across the four distinct categories of FAIR independently, yielding granular, pillar-specific ratings. Unlike aggregate scoring models, which can inadvertently mask technical deficiencies by averaging performance across categories, this multi-dimensional approach ensures that a repositorys distinct strengths and bottlenecks remain fully visible. Our findings demonstrate that standardized scoring not only clarifies data accessibility for users but also highlights specific operational gaps, allowing repository providers to identify precisely where the service implementation can be enhanced. By establishing this data-driven service in the agronomy domain, we provide a scalable template for the broader NFDI and EOSC ecosystems to foster a culture of excellence in research data stewardship.

3
LightAlign: a lightweight pairwise aligner for memory-constrained HiFi read assembly

Liu, J.; Zhang, J.

2026-08-05 bioinformatics 10.64898/2026.07.30.741934 medRxiv
Top 0.1%
22.0%
Show abstract

IntroductionCurrent de novo genome assembly tools often demand substantial memory resources, and their execution typically relies on high-performance computing (HPC) clusters. This dependency limits their use in resource-constrained settings. Furthermore, mainstream third-generation sequencing assembly and alignment tools usually require explicit detection of overlap regions between reads, a process that often entails significant computational and storage overhead. ResultsTo address this issue, we developed LightAlign, a lightweight alignment tool for HiFi data that innovatively uses sequence-derived fuzzy features and reduces the peak memory usage during overlap detection. ConclusionsWhen combined with miniasm, LightAlign generated bacterial draft assemblies while maintaining peak memory usage below 1 GB and completed overlap generation for the tested eukaryotic datasets within 1.88 GB RAM.

4
MPGEM: A harmonized and transcriptome-complete resource for large-scale reuse of legacy human microarray data

Gupta, S.; Verma, A. K.; Jana, S.; Ahmad, S.

2026-08-25 bioinformatics 10.64898/2026.08.20.746052 medRxiv
Top 0.1%
18.9%
Show abstract

Abstract Background: Legacy microarray datasets provide an extensive record of human transcriptomic biology, but their reuse is constrained by differences in platform design, preprocessing, measurement scale, and gene coverage. Platforms measuring only subsets of genes cannot readily be integrated with higher-coverage platforms, limiting large-scale analysis and computational modeling. Results: We developed Multi-Platform Gene Expression Matrix (MPGEM), a computational framework and resource for harmonizing and completing gene-expression profiles across heterogeneous microarray platforms. MPGEM uses a Reference Quantile Distribution (RQD) and generalized Reference Subset Quantile Distribution (RSQD) framework to transform profiles with different gene coverage onto a common quantitative scale. The MPGEM Engine, a multilayer perceptron, predicts expression of unmeasured genes from genes shared across platforms. Applied to Affymetrix GPL570, GPL571, and GPL96, MPGEM uses GPL570 as a 19,320- gene reference space comprising 12,712 predictor and 6,608 target genes. The resulting resource contains 207,135 human gene-expression profiles across 19,320 genes. Evaluation using masked GPL570 profiles yielded mean sample-wise Pearson and Spearman correlations of 0.944 and 0.939, respectively, and mean gene-wise correlations of 0.830 and 0.825. The lowest-performing 5% of target genes achieved a mean Pearson correlation of 0.683. MPGEM showed comparable or higher predictive performance than baseline mean imputation and K-nearest-neighbor approaches. Conclusions: MPGEM transforms heterogeneous, partially measured legacy microarray profiles into a harmonized, transcriptome-complete representation, facilitating their reuse for large-scale transcriptomic analysis, biomarker discovery, systems biology, and machine learning. The framework, trained models, and expression resource are provided as open-source resources.

5
Programmatic access to ICTV virus taxonomy through a public ontology API

Lieutaud, P.; McLaughlin, j.; Hendrickson, R. C.; David, R.; Parkinson, H.; Lefkowitz, E.; Dempsey, D.; Coutard, B.

2026-06-16 bioinformatics 10.64898/2026.06.16.732600 medRxiv
Top 0.1%
17.1%
Show abstract

The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. To address this, we developed a public and sustainable solution leveraging ontology-based APIs. Successive ICTV Master Species List (MSL) releases were transformed into a structured ontology and deployed as a unified representation through the Ontology Lookup Service (OLS). The framework also provides ICTV-NCBI mappings and helper libraries for integration into downstream systems. This enables, for the first time, public programmatic retrieval of current and historical virological taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness. Key pointsO_LIICTV provides the official taxonomy for classifying viruses and naming virus taxa, but lacks standardised programmatic access. C_LIO_LITransforming ICTV data into an ontology enables semantic, machine-actionable access across releases via ontology-based APIs. C_LIO_LIICTV-NCBI mappings support interoperability across bioinformatics resources. C_LIO_LIThe framework enables programmatic resolution of current and historical viral taxa. C_LIO_LIThis approach provides a reusable model for exposing biological datasets through public APIs. C_LI

6
From Abandoned Scripts to FAIR Community Pipelines: Rescuing Orphan Bioinformatics Workflows with nf-core - Lessons from Light-Sheet Fluorescence Microscopy

Schwitalla, C.; Kuhn Cuellar, L.; Hoertenhuber, M.; Grote, N.; Woller, T.; Lamberti, I.; Pavie, B.; Kuestner, T.; Kyere, F. A.; Curtin, I.; Stein, J. L.; Nahnsen, S.

2026-08-03 bioinformatics 10.64898/2026.07.29.741447 medRxiv
Top 0.1%
15.5%
Show abstract

BackgroundResearch software is essential for modern data analysis but is often developed and maintained by a small number of researchers. When developers leave, software may become orphaned, limiting reuse and risking the loss of valuable domain knowledge and computational methods. While the FAIR Principles for Research Software (FAIR4RS) provide an essential foundation for improving the reuse of research software, compliance with these principles alone does not guarantee practical reusability. Here, we investigate whether orphaned scientific software can be systematically rescued and transformed into sustainable, reusable workflows using established software engineering practices and community standards. FindingsWe re-engineered the abandoned MATLAB-based NuMorph toolkit for large-scale light-sheet microscopy image analysis into nf-core/lsmquant, a Nextflow-based workflow developed according to nf-core community guidelines. The re-engineered workflow preserved the original scientific methods at comparable computational cost while improving the softwares FAIRness, portability, and reproducibility. Integration into the nf-core ecosystem provides a community-driven framework that supports software sustainability through distributed maintenance and shared development practices, while the modular workflow architecture simplified adaptation of nf-core/lsmquant to additional light-sheet microscopy datasets beyond the original application ConclusionOur work demonstrates that orphaned scientific software can be successfully rescued through systematic re-engineering guided by FAIR and software sustainability principles. By transforming a legacy codebase into a community-maintained workflow, we preserve valuable domain-specific methods while improving usability, maintainability, and reproducibility. This approach provides a practical strategy for recovering orphan research software and integrating it into modern, reusable research ecosystems.

7
PhaGAMeToo: A semi-automated workflow for merging structural and functional annotation of phage genomes and generation of a GenBank file

Demircioglu, E.; Bole, M.; da Rocha, U. N.; Kallies, R.

2026-08-18 bioinformatics 10.64898/2026.08.09.738482 medRxiv
Top 0.1%
14.9%
Show abstract

MotivationAnalysing and concatenating phage annotation is time-consuming. Further, the output of phage annotation tools cannot be directly submitted to public repositories. To deal with these issues, we developed PhaGAMeToo. This command-line workflow for Linux integrates the functional annotations of two major viral annotation tools (Pharokka and VIBRANT), enabling faster and more accurate functional annotation. Furthermore, the workflow provides merged annotations as submission-ready GenBank files. ResultsPhaGAMeToo uses three steps to generate submission-ready GenBank files. The user uses the reoriented viral genomes as inputs for Pharokka and VIBRANT. Pharokka and VIBRANT-generated files are parsed through the PhaGAMeToo workflow to produce a merged GenBank file. Further, PhaGAMeToo also enables the use of BLASTP to annotate hypothetical proteins not identified by Pharokka and VIBRANT. It then merges the results into a submission-ready GenBank file(s). We tested PhaGAMeToo in three different Use Cases. We analysed reference and uncultivated viral genomes manually curated or directly recovered using MuDoGeR in our Use Cases. In the Use Case 1, we analysed four different NCBI reference genomes. In the Use Cases 2 and 3, we analysed seven recently described huge phage genomes and 56 uncultivated viral genomes recovered from 30 soil metagenomes, respectively. Availability and implementationThe source code, documentation, and installation instructions for PhaGAMeToo are available at https://github.com/NFDI4Microbiota/PhaGAMeToo ContactRene.Kallies@uba.de; ebrardemircioglu25@hacettepe.edu.tr Supplementary informationSupplementary data will be made available upon publication.

8
oxo-flow: compiled, memory-safe bioinformatics workflow orchestration

Wang, S.

2026-06-15 bioinformatics 10.64898/2026.06.11.731578 medRxiv
Top 0.1%
13.3%
Show abstract

Bioinformatics analyses depend on workflow engines to coordinate dozens of computational tools across complex dependency chains. The most widely adopted engines--Snakemake, Nextflow, the Common Workflow Language (CWL), and the Workflow Description Language (WDL)--run on interpreted or just-in-time (JIT) compiled language runtimes, incurring hundreds of milliseconds of startup latency and providing no compile-time safety guarantees from the host language. We developed oxo-flow, a workflow engine written in Rust that compiles to a single native binary. On an Apple M5 processor, oxo-flow parses, validates, and dry-runs a production-scale workflow in roughly 22 milliseconds--before Snakemake or Nextflow have finished loading their runtime environments. Peak memory usage is 16 megabytes, representing six- to seven-fold reductions relative to Snakemake and Nextflow. Dry-run latency is essentially independent of workflow size: a hundred-fold increase in rule count adds approximately 0.4 milliseconds. oxo-flow integrates 31 command-line tools, a REST interface with 60 endpoints, an embedded web application, and native cluster submission into a single 10-megabyte binary. It provides per-rule environment isolation across seven backends, checkpoint-based fault tolerance with cryptographic output verification, and a formal installation and operational qualification protocol for regulated laboratory environments. Ten curated workflows and three demonstration pipeline repositories are available. oxo-flow is freely available under Apache License 2.0 at https://github.com/Traitome/oxo-flow.

9
Survey and Evaluation of Applied Containerization Practices in Bioinformatics

Nagraj, V.; Turner, S. D.; Magee, N.

2026-07-18 bioinformatics 10.64898/2026.07.17.739198 medRxiv
Top 0.1%
13.1%
Show abstract

Containerization enables portable, reproducible, and scalable scientific computing. However, container development, documentation, and deployment practices can vary widely, even within a single domain. Understanding patterns of how containers are implemented in real-world settings can inform community guidelines, quality scoring rubrics, and systems to automate software review. Focusing on bioinformatics as an exemplar domain, we surveyed published software tools to find applied containerization examples. After identifying >250 tools to review, we annotated metadata regarding version control, asset provenance, and general container image health and tested the ability to build and pull the container images in a purpose-built evaluation platform (https://github.com/vpnagraj/socr8s). The majority of images tested did not build successfully in our environment. Of the tool characteristics we tracked, the strongest predictor of build success was version control activity. Tools with commits in the preceding two years were roughly twice as likely to build. Deeper assessment of failures identified a variety of issues, including missing assets and broken dependency chains. While many tools had published images accessible in open registries, 24% could neither be built nor pulled. Images tended to be large, with median size of 1.57GB for those that were able to be built. Where image specification files were available, we found that more advanced container orchestration and build techniques were uncommon. Our results highlight areas of improvement in how containerized tools are built and maintained in bioinformatics and scientific computing software in general.

10
PhenoStream: A Cyberinfrastructure for Automated and AI-Based Crop Trait Extraction from Aerial Imagery

Varela, S.; Ruhter, J.; Sacks, E.; Zheng, X.; Allen, D.; Hale, A.; Landry, C.; Kuang, X.; Long, B.; Zhu, Y.; Proma, S.; Kaur, S.; Jarquin, D.; Morrison, J.; Leakey, A.

2026-08-30 plant biology 10.64898/2026.08.26.747008 medRxiv
Top 0.1%
13.0%
Show abstract

The integration of digital technologies for high-throughput field phenotyping is critical for accelerating crop improvement in agriculture. However, extracting traits from remote sensing data remains constrained by fragmented workflows, manual intervention, and limited interoperability among existing tools, resulting in delays that hinder timely biological insight and decision-making. To address these challenges, we present PhenoStream (Phenotyping Streaming), a scalable, end-to-end cyberinfrastructure designed to automate the full lifecycle of aerial imagery-based phenotyping, from data acquisition to plot- and genotype-level inference. The framework integrates automated data ingestion from distributed field sites, geospatial processing, and AI-enabled trait extraction within a unified, user-accessible graphical interface. Its modular and extensible architecture supports adaptable trait modeling and seamless integration of new data sources, enabling deployment across diverse crops, environments, and experimental designs. We demonstrate the system across a large multi-location field trial network of bioenergy crops, where it enables high-throughput characterization of spatiotemporal growth dynamics, genotype-by-environment (GxE) interactions, and predictive modeling of key agronomic traits. By significantly reducing processing latency and manual effort, the platform facilitates near-real-time analysis and reproducible workflows. This work establishes a generalizable and scalable pathway for operationalizing very-high-spatial resolution aerial phenotyping in agricultural research. By bridging data acquisition and analytics, the end-to-end cyberinfrastructure provides a foundation for integrating heterogeneous and unstructured data streams--including remote sensing, environmental, and management data--toward data-driven decision making in agriculture.

11
Topology-Based Query Framework for Longitudinal Omics Trajectories

Zounemat-Kermani, N.; Richardson, M.; Faiz, A.; Wang, S.; Sun, K.; Vuckovic, D.; van den Berge, M.; Maitland-van der Zee, A. H.; Sayers, I.; Dahlen, S.-E.; Brightling, C. E.; Siddiqui, S.; Chung, K. F.; Nawijn, M. C.; Chadeau-Hyam, M.; Adcock, I. M.

2026-08-21 bioinformatics 10.64898/2026.08.18.745427 medRxiv
Top 0.1%
13.0%
Show abstract

1 Abstract 1.1 Background Many longitudinal omics studies contain only a small number of repeated measurements collected before, during, or after an intervention. Existing approaches, including mixed-effects models and generalized additive models, estimate temporal effects but do not generally provide a discrete representation of trajectory topology that can be queried directly across experimental groups. 1.2 Methods We developed LongOmicsTraj, an open-source R package for topology-based representation and querying of short longitudinal omics trajectories. The framework encodes the direction of change between adjacent visits as up, down, or flat, with the ordered sequence defining an Ordinal Trajectory State (OTS). LongOmicsTraj operates downstream of trajectory estimation and can therefore be applied to empirical summaries or model-derived visit-level estimates, including those from linear mixed-effects models, generalized additive models, and polynomial regression, following a maSigPro-style time-course formulation [1]. OTS labels provide a common representation for topology-based querying, cross-group comparison, and evaluation of higherlevel representations such as trajectory clusters. We evaluated the framework using controlled simulations and bronchial biopsy transcriptomic data from the GLUCOLD corticosteroid intervention study (GEO accession GSE36221), measured at baseline, 6 months, and 30 months. The biological analysis compared continued inhaled corticosteroid (ICS) treatment, ICS withdrawal after 6 months, and placebo. 1.3 Results In simulations, LongOmicsTraj recovered predefined stable, monotonic, transient, rebound, and oscillatory trajectories with high accuracy when longitudinal signal was sufficiently clear, with performance declining under high-noise conditions and depending partly on the upstream estimator. In GLUCOLD, comparator-aware topology queries reduced 20,358 measured transcripts to 168 genes showing a corticosteroid response that was maintained during continued treatment, reversed following withdrawal, and was not reproduced under placebo. The selected genes included established corticosteroid-response genes and were enriched for immune-cell migration, chemotaxis, cell adhesion, and extracellular-matrix organisation. Topology-aware evaluation of FlexMix trajectory clusters additionally revealed substantial within-cluster temporal heterogeneity, with topology purities of approximately 46% to 60%. 1.4 Conclusions LongOmicsTraj provides a compact, directly queryable representation of temporal direction and order in short longitudinal omics studies. It complements existing longitudinal estimation and clustering methods by making trajectory structure explicit, enabling structured cross-group queries and quantification of temporal heterogeneity within trajectory clusters.

12
CREPAS: a reproducible nascent chromatin sequencing analysis pipeline for epigenome replication studies

Ruiz-Perez, S.; Du, Q.; Biran, A.; Groth, A.; Alcaraz, N.

2026-06-25 bioinformatics 10.64898/2026.06.21.732899 medRxiv
Top 0.1%
12.8%
Show abstract

Chromatin-based genomics data are essential for understanding genome regulation and the mechanisms underlying epigenetic memory. Recent methods such as ChOR-seq and SCAR-seq assess histone modifications and chromatin-associated proteins during and after replication, capturing chromatin states that contribute to memory across cell divisions. Current tools for chromatin data analysis lack scalability and reproducibility across computing infrastructures, offer limited parameters, and are applicable only to a few sequencing techniques, ignoring the information from nascent chromatin assays. To address these challenges, we developed CREPAS, a Nextflow pipeline for analyzing nascent and parental chromatin sequencing data, including ChIP-seq, ChOR-seq, SCAR-seq, OK-seq, ATAC-seq, CUT&RUN, and CUT&Tag, and derivative protocols. CREPAS provides an end-to-end solution, from quality control to advanced analyses, including downsampling, peak calling, annotation, and visualization. By harnessing quantitative assays such as qChIP-seq and qChOR-seq, the normalization methods in CREPAS allow to compare the restoration kinetics of individual marks or proteins across replication timepoints. Moreover, the pipeline includes calculations such as fork directionality and partitioning using OK-seq and SCAR-seq data, linking replication dynamics to epigenetic inheritance. CREPAS is a valuable resource that enhances the efficiency and reproducibility of nascent chromatin sequencing data analyses, enabling the study of chromatin replication and propagation of epigenetic states. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/732899v1_ufig1.gif" ALT="Figure 1"> View larger version (24K): org.highwire.dtl.DTLVardef@29ad26org.highwire.dtl.DTLVardef@26b9acorg.highwire.dtl.DTLVardef@67dcb8org.highwire.dtl.DTLVardef@cbc299_HPS_FORMAT_FIGEXP M_FIG C_FIG

13
An evaluation of clustering and assembly strategies from Iso-Seq data in the absence of reference genomes in non-model animals

Eleftheriadi, K.; Vazquez-Valls, M.; Fernandez, R.

2026-07-08 evolutionary biology 10.1101/2025.09.18.677004 medRxiv
Top 0.1%
12.8%
Show abstract

Transcriptome assembly enables the recovery of expressed genes and isoforms, but the optimal strategy for reconstructing transcriptomes from long-read sequencing remains unresolved. In particular, establishing best practices for generating accurate gene models and selecting representative isoforms is essential for comparative genomics, as for orthology inference typically only the longest isoform per gene model is included. Here, we systematically compare clustering and de novo assembly methods using PacBio Iso-Seq data from 17 animal lineages spanning seven phyla, most of them non-model species, with the goal of investigating which methodology is more adequate to select one isoform per gene model, in the absence of specific pipelines to do so. We evaluate four approaches: isoseq cluster, CD-HIT, RNA-Bloom2 and isONform. We benchmark them with short-reads using Trinity, assessing assembly quality with BUSCO completeness, short-read mapping rates, coding sequence recovery, and longest isoform prediction. Our results show that CD-HIT clustering at high similarity thresholds ([≥]99%) yields the most complete and coding-rich long-read transcriptomes, rivaling Trinity while avoiding its high redundancy. Consensus-based methods such as isoseq cluster and isONform recover fewer single-copy orthologs (mirrored in a lower BUSCO score) and achieve lower mapping rates, while RNA-Bloom2 provide intermediate performance with reduced duplication. Together, these findings establish, to date, CD-HIT as a robust and practical strategy for transcriptome reconstruction from long-read data when genomic references are unavailable. By benchmarking de novo methods across a taxonomically broad dataset, this work defines the realistic capabilities of long-read transcriptome reconstruction in the absence of a reference genome and provides practical guidance for deriving high-quality gene models and selecting representative isoforms for orthology inference in non-model species.

14
Global vascular plants reveal persistent gaps across taxa and ecoregions

Maciel, E. A.

2026-08-28 evolutionary biology 10.64898/2026.08.24.746674 medRxiv
Top 0.1%
12.7%
Show abstract

Biodiversity aggregators such as GBIF provide unprecedented access to global biodiversity data, yet their representativeness remains uneven across space and taxa. This study examined the spatial and taxonomic structure of global vascular plant data available on GBIF. Six filters were applied to the GBIF vascular plant dataset, resulting in the removal of 54% of all records. Together, the filters explained more than 90% of the identified spatial issues, with duplicate and missing coordinates accounting for most of the variation. A higher number of occurrence records was associated with a greater number of spatial issues. Record distributions became progressively more even at finer taxonomic levels, from orders to species. The time series of occurrences for species, genera, and families increased sharply after 1800 and continued to rise, with no apparent stabilisation. Of the 824 ecoregions covered, 73 accounted for 72% of all occurrence records. These ecoregions spanned all continents but were strongly concentrated in Europe, followed by North America and Oceania. The analyses reveal four key patterns: (1) data volume is positively associated with spatial issues; (2) a small number of taxa account for a large proportion of records, whereas many are represented by relatively few; (3) occurrence data aggregated by GBIF have increased continuously since 1800; and (4) record coverage remains highly uneven across the world's ecoregions. These results highlight the substantial contribution of biodiversity data aggregators to expanding access to biological information while demonstrating the persistent spatial and taxonomic biases that shape their contents. Such biases should be explicitly considered when assessing data completeness and quality and when using aggregated occurrence records to infer global biodiversity patterns.

15
A chromosome-level assembly of an aquatic passerine bird, the northern white-throated dipper, Cinclus cinclus cinclus (Linnaeus, 1758)

Strand, M. A.; Toerresen, O. K.; Skage, M.; Ferrari, G.; Tooming-Klunderud, A.; Johnsen, A.; Jakobsen, K. S.

2026-08-24 genomics 10.64898/2026.08.20.746034 medRxiv
Top 0.1%
12.6%
Show abstract

We present a chromosome-level genome assembly of a female Norwegian white-throated dipper (Cinclus cinclus cinclus) generated using Oxford Nanopore Technologies (ONT) long reads and Hi-C scaffolding. The assembly comprises two pseudo-haplotypes, hap1 (1186 Mb) and hap2 (1115 Mb), with 96.7% and 94.4% of sequences assigned to chromosome-scale scaffolds, respectively. Both pseudo-haplotypes contain 40 autosomes, with the Z and W sex chromosomes assigned to hap1. Compared with the PacBio HiFi-based C. c. gularis reference assembly bCinCin1.1.pri, which contains 38 autosomes, sequence represented as a single dot-chromosome (chr 36) is resolved into three distinct dot-chromosomes (chr 36, 39, and 40), a configuration supported by Hi-C contact patterns. BUSCO completeness was high for hap1 (99.2%) and hap2 (95.0%), with 19,003 and 17,746 predicted protein-coding genes, respectively. Compared with the HiFi-based C. c. gularis reference and HiFi-based assemblies generated from the same individual, the ONT-derived assemblies were substantially less fragmented and recovered more sequence from the smallest chromosomes. Synteny was otherwise largely conserved between subspecies. HiFi depletion increased strongly from macrochromosomes to micro- and dot-chromosomes, and HiFi-depleted regions were enriched for repeats and predicted non-B-DNA-associated features, particularly G-quadruplexes and direct repeats, whereas ONT coverage remained comparatively stable. These results show that conventional genome-wide assembly metrics can obscure substantial differences in the recovery of repeat-rich avian dot-chromosomes and highlight the value of chromosome-aware evaluation and ONT sequencing for recovering these regions.

16
dbverse scales spatial omics analysis with embedded analytical databases

Ruiz, E. C.; Jarzabek, V.; Chen, J. G.; Rizvanov, T.; Amin, I.; Dries, R.

2026-08-14 bioinformatics 10.64898/2026.08.09.743742 medRxiv
Top 0.1%
12.5%
Show abstract

Spatial omics datasets are increasing in size and complexity, exceeding the memory of standard computers and thereby limiting data analysis. Here we present dbverse, a framework for larger-than-memory matrix, spatial and genomic data analysis in embedded analytical databases. Benchmarks show dbverse provides orders of magnitude runtime improvements relative to established in-memory and file-backed methods for core operations in single-cell and spatial omics analysis. We integrated dbverse with Giotto Suite, scaling end-to-end preprocessing of millions of cells and enabling spatial alternative polyadenylation analysis as demonstrated on a Visium HD 3' ovarian clear cell carcinoma sample. The dbverse framework provides an interoperable database foundation for larger-than-memory spatial omics analysis on ordinary computers.

17
PanGBank: a large-scale resource of precomputed microbial pangenomes built with PPanGGOLiN

Mainguy, J.; Lemane, T.; Bazin, A.; Arnoux, J.; Gautreau, G.; Medigue, C.; Calteau, A.; Vallenet, D.

2026-08-09 bioinformatics 10.64898/2026.08.05.742796 medRxiv
Top 0.1%
12.5%
Show abstract

PanGBank (https://pangbank.genoscope.cns.fr) is a comprehensive open-access database providing precomputed prokaryotic pangenomes at a broad taxonomic scale. Built upon PPanGGOLiN partitioned pangenome graphs, PanGBank addresses the growing need for large-scale comparative genomics through a standardized, regularly updated, and fully accessible resource. The initial release comprises two complementary collections covering more than 4,600 prokaryotic species from the Genome Taxonomy Database (GTDB), encompassing over 393,000 genomes: GTDB all, maximizing taxonomic and environmental diversity through the inclusion of MAGs and SAGs, and GTDB refseq, focusing on high-quality, annotation-rich genomes. Each species-level pangenome integrates graph-based statistical partitions into persistent, shell, and cloud gene families, together with regions of genomic plasticity (panRGP) and co-localized functional modules (panModule). PanGBank offers multiple access modes, including a REST API, a command-line interface (PanGBank-cli), and an interactive web interface. By combining large-scale pangenome resources with advanced graph-based analyses, PanGBank provides a scalable framework for exploring microbial diversity, genome evolution, functional variation, and the dissemination of adaptive traits across prokaryotic populations, as illustrated by a use case on Acinetobacter baumannii pangenome investigating the distribution and evolution of antimicrobial resistance determinants. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=63 SRC="FIGDIR/small/742796v1_ufig1.gif" ALT="Figure 1"> View larger version (23K): org.highwire.dtl.DTLVardef@186b88dorg.highwire.dtl.DTLVardef@1be33d0org.highwire.dtl.DTLVardef@3bc596org.highwire.dtl.DTLVardef@292401_HPS_FORMAT_FIGEXP M_FIG C_FIG

18
The recount3 Python package for programmatic access to uniformly processed RNA-seq data

Alsalihi, A.; Flight, R. M.; Moseley, H. N. B.

2026-06-20 bioinformatics 10.64898/2026.06.17.732943 medRxiv
Top 0.1%
12.0%
Show abstract

The recount3 online resource provides tens of thousands of uniformly processed RNA-seq samples across human and mouse from major sequencing repositories like the Sequence Read Archive. While access to these datasets has traditionally been centered in the R/Bioconductor ecosystem, the growing prominence of Python in bioinformatics and machine learning necessitates native, efficient tooling for Python users. Therefore, we present the recount3 Python package with robust application programming interface (API) and command-line interface (CLI) for discovering, downloading, and materializing recount3 resources. The software orchestrates uniform resource locator (URL) resolution, persistent on-disk caching, and the automatic parsing of data into analysis-ready data structures, including Pandas DataFrames and BiocPy RangedSummarizedExperiment objects. The recount3 Python package drastically lowers the barrier to entry for large-scale utilization of RNA-seq data in Python-based computational pipelines, bridging the gap between massive public transcriptomic data and modern machine learning ecosystems.

19
trAIt: Species-by-Trait Data Retrieval using Large Language Models

Balaji, S.; Martinson, K. A.; Schellenberger, J. S.; Koley, J.; Inman, C. M.; Hofmann, H. A.; Young, R. L.; Harpak, A.

2026-06-24 bioinformatics 10.64898/2026.06.19.732660 medRxiv
Top 0.1%
11.9%
Show abstract

Biological research often requires information about species traits. Manual literature collation can be time-consuming and miss parts of the literature. To address this gap, we developed trAIt, a publicly available software for the retrieval of characteristics of species from scientific literature catalogued in the Europe PubMed Central (PubMed) database. trAIt provides a graphical user interface (GUI) in which users specify species and characteristics of interest. Leveraging a large language model (LLM), trAIt retrieves relevant papers, combines their content through a consensus-based summarization model, and outputs a species-by-characteristic table. For a case study involving frog species, trAIt recovered 47.1% of trait-species combinations in 2.75 hours, while an expert curator independently recovered 62.4% over months. The consensus-based summarization substantially aids accuracy compared to single-source extraction. Across three case studies of vertebrate taxa, an expert confirmed the accuracy of 70.9% of trait-species entries recovered by trAIt. We observed considerable variation across taxa in trAIts accuracy, which is possibly due to heterogeneity in open-access literature availability and inconsistencies in species and trait terminology. In sum, our analysis suggests that LLM-based tools can accelerate biological data synthesis but should be used to support domain experts research, rather than replace their judgment.

20
Kiosc: an integrated platform for managing bioinformatics data analysis containers

Marotta, F.; Stolpe, O.; Obermayer, B.; Weiner, J.; Holtgrewe, M.; Beule, D.; Nieminen, M.

2026-08-24 bioinformatics 10.64898/2026.08.20.745983 medRxiv
Top 0.1%
11.8%
Show abstract

In many bioinformatic data analysis projects, it is convenient to visualize plots and results through an interactive web app or dashboard. These interactive reports can then be shared with customers, collaborators, or the general public. Publishing and sharing these apps is not straightforward, becoming especially cumbersome when the number of projects and customers start growing. Docker containers offer a convenient way to package, distribute, and run interactive web apps, and their use is already widespread in the bioinformatics community. We developed Kiosc to simplify the orchestration of containerized web apps, organize them into projects, and regulate access control. We implemented it as a web server based on the Django framework, with a user- and admin-friendly interface as well as a REST API for programmatic tasks. Users can select Docker containers packaging apps like Plotly Dash, Shiny, or Quarto, and configure them to display the results of their analysis. Kiosc runs the containers with the appropriate network configuration and acts as a proxy to the web services running inside the containers. We have been maintaining a Kiosc instance for more than 5 years, serving 321 containers in 150 projects across multiple institutions. In this article, we introduce the main functionality in Kiosc and describe four use-cases that show how Kiosc can prove helpful to the broader bioinformatics community, such as configuring and running web apps for the interactive visualization of workflow results, and publishing companion apps for scientific articles. Kiosc is a self-hosted platform for publishing web apps, which doesn't require significant expertise in either Docker or network administration to be deployed. It provides a similar service to Kubernetes, but with a convenient web interface and much lower administration overhead.