Back

GigaScience

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match GigaScience's content profile, based on 212 papers previously published here. The average preprint has a 0.16% match score for this journal, so anything above that is already an above-average fit.

1
MPGEM: A harmonized and transcriptome-complete resource for large-scale reuse of legacy human microarray data

Gupta, S.; Verma, A. K.; Jana, S.; Ahmad, S.

2026-08-25 bioinformatics 10.64898/2026.08.20.746052 medRxiv
Top 0.1%
18.9%
Show abstract

Abstract Background: Legacy microarray datasets provide an extensive record of human transcriptomic biology, but their reuse is constrained by differences in platform design, preprocessing, measurement scale, and gene coverage. Platforms measuring only subsets of genes cannot readily be integrated with higher-coverage platforms, limiting large-scale analysis and computational modeling. Results: We developed Multi-Platform Gene Expression Matrix (MPGEM), a computational framework and resource for harmonizing and completing gene-expression profiles across heterogeneous microarray platforms. MPGEM uses a Reference Quantile Distribution (RQD) and generalized Reference Subset Quantile Distribution (RSQD) framework to transform profiles with different gene coverage onto a common quantitative scale. The MPGEM Engine, a multilayer perceptron, predicts expression of unmeasured genes from genes shared across platforms. Applied to Affymetrix GPL570, GPL571, and GPL96, MPGEM uses GPL570 as a 19,320- gene reference space comprising 12,712 predictor and 6,608 target genes. The resulting resource contains 207,135 human gene-expression profiles across 19,320 genes. Evaluation using masked GPL570 profiles yielded mean sample-wise Pearson and Spearman correlations of 0.944 and 0.939, respectively, and mean gene-wise correlations of 0.830 and 0.825. The lowest-performing 5% of target genes achieved a mean Pearson correlation of 0.683. MPGEM showed comparable or higher predictive performance than baseline mean imputation and K-nearest-neighbor approaches. Conclusions: MPGEM transforms heterogeneous, partially measured legacy microarray profiles into a harmonized, transcriptome-complete representation, facilitating their reuse for large-scale transcriptomic analysis, biomarker discovery, systems biology, and machine learning. The framework, trained models, and expression resource are provided as open-source resources.

2
PhaGAMeToo: A semi-automated workflow for merging structural and functional annotation of phage genomes and generation of a GenBank file

Demircioglu, E.; Bole, M.; da Rocha, U. N.; Kallies, R.

2026-08-18 bioinformatics 10.64898/2026.08.09.738482 medRxiv
Top 0.1%
14.9%
Show abstract

MotivationAnalysing and concatenating phage annotation is time-consuming. Further, the output of phage annotation tools cannot be directly submitted to public repositories. To deal with these issues, we developed PhaGAMeToo. This command-line workflow for Linux integrates the functional annotations of two major viral annotation tools (Pharokka and VIBRANT), enabling faster and more accurate functional annotation. Furthermore, the workflow provides merged annotations as submission-ready GenBank files. ResultsPhaGAMeToo uses three steps to generate submission-ready GenBank files. The user uses the reoriented viral genomes as inputs for Pharokka and VIBRANT. Pharokka and VIBRANT-generated files are parsed through the PhaGAMeToo workflow to produce a merged GenBank file. Further, PhaGAMeToo also enables the use of BLASTP to annotate hypothetical proteins not identified by Pharokka and VIBRANT. It then merges the results into a submission-ready GenBank file(s). We tested PhaGAMeToo in three different Use Cases. We analysed reference and uncultivated viral genomes manually curated or directly recovered using MuDoGeR in our Use Cases. In the Use Case 1, we analysed four different NCBI reference genomes. In the Use Cases 2 and 3, we analysed seven recently described huge phage genomes and 56 uncultivated viral genomes recovered from 30 soil metagenomes, respectively. Availability and implementationThe source code, documentation, and installation instructions for PhaGAMeToo are available at https://github.com/NFDI4Microbiota/PhaGAMeToo ContactRene.Kallies@uba.de; ebrardemircioglu25@hacettepe.edu.tr Supplementary informationSupplementary data will be made available upon publication.

3
PhenoStream: A Cyberinfrastructure for Automated and AI-Based Crop Trait Extraction from Aerial Imagery

Varela, S.; Ruhter, J.; Sacks, E.; Zheng, X.; Allen, D.; Hale, A.; Landry, C.; Kuang, X.; Long, B.; Zhu, Y.; Proma, S.; Kaur, S.; Jarquin, D.; Morrison, J.; Leakey, A.

2026-08-30 plant biology 10.64898/2026.08.26.747008 medRxiv
Top 0.1%
13.0%
Show abstract

The integration of digital technologies for high-throughput field phenotyping is critical for accelerating crop improvement in agriculture. However, extracting traits from remote sensing data remains constrained by fragmented workflows, manual intervention, and limited interoperability among existing tools, resulting in delays that hinder timely biological insight and decision-making. To address these challenges, we present PhenoStream (Phenotyping Streaming), a scalable, end-to-end cyberinfrastructure designed to automate the full lifecycle of aerial imagery-based phenotyping, from data acquisition to plot- and genotype-level inference. The framework integrates automated data ingestion from distributed field sites, geospatial processing, and AI-enabled trait extraction within a unified, user-accessible graphical interface. Its modular and extensible architecture supports adaptable trait modeling and seamless integration of new data sources, enabling deployment across diverse crops, environments, and experimental designs. We demonstrate the system across a large multi-location field trial network of bioenergy crops, where it enables high-throughput characterization of spatiotemporal growth dynamics, genotype-by-environment (GxE) interactions, and predictive modeling of key agronomic traits. By significantly reducing processing latency and manual effort, the platform facilitates near-real-time analysis and reproducible workflows. This work establishes a generalizable and scalable pathway for operationalizing very-high-spatial resolution aerial phenotyping in agricultural research. By bridging data acquisition and analytics, the end-to-end cyberinfrastructure provides a foundation for integrating heterogeneous and unstructured data streams--including remote sensing, environmental, and management data--toward data-driven decision making in agriculture.

4
Topology-Based Query Framework for Longitudinal Omics Trajectories

Zounemat-Kermani, N.; Richardson, M.; Faiz, A.; Wang, S.; Sun, K.; Vuckovic, D.; van den Berge, M.; Maitland-van der Zee, A. H.; Sayers, I.; Dahlen, S.-E.; Brightling, C. E.; Siddiqui, S.; Chung, K. F.; Nawijn, M. C.; Chadeau-Hyam, M.; Adcock, I. M.

2026-08-21 bioinformatics 10.64898/2026.08.18.745427 medRxiv
Top 0.1%
13.0%
Show abstract

1 Abstract 1.1 Background Many longitudinal omics studies contain only a small number of repeated measurements collected before, during, or after an intervention. Existing approaches, including mixed-effects models and generalized additive models, estimate temporal effects but do not generally provide a discrete representation of trajectory topology that can be queried directly across experimental groups. 1.2 Methods We developed LongOmicsTraj, an open-source R package for topology-based representation and querying of short longitudinal omics trajectories. The framework encodes the direction of change between adjacent visits as up, down, or flat, with the ordered sequence defining an Ordinal Trajectory State (OTS). LongOmicsTraj operates downstream of trajectory estimation and can therefore be applied to empirical summaries or model-derived visit-level estimates, including those from linear mixed-effects models, generalized additive models, and polynomial regression, following a maSigPro-style time-course formulation [1]. OTS labels provide a common representation for topology-based querying, cross-group comparison, and evaluation of higherlevel representations such as trajectory clusters. We evaluated the framework using controlled simulations and bronchial biopsy transcriptomic data from the GLUCOLD corticosteroid intervention study (GEO accession GSE36221), measured at baseline, 6 months, and 30 months. The biological analysis compared continued inhaled corticosteroid (ICS) treatment, ICS withdrawal after 6 months, and placebo. 1.3 Results In simulations, LongOmicsTraj recovered predefined stable, monotonic, transient, rebound, and oscillatory trajectories with high accuracy when longitudinal signal was sufficiently clear, with performance declining under high-noise conditions and depending partly on the upstream estimator. In GLUCOLD, comparator-aware topology queries reduced 20,358 measured transcripts to 168 genes showing a corticosteroid response that was maintained during continued treatment, reversed following withdrawal, and was not reproduced under placebo. The selected genes included established corticosteroid-response genes and were enriched for immune-cell migration, chemotaxis, cell adhesion, and extracellular-matrix organisation. Topology-aware evaluation of FlexMix trajectory clusters additionally revealed substantial within-cluster temporal heterogeneity, with topology purities of approximately 46% to 60%. 1.4 Conclusions LongOmicsTraj provides a compact, directly queryable representation of temporal direction and order in short longitudinal omics studies. It complements existing longitudinal estimation and clustering methods by making trajectory structure explicit, enabling structured cross-group queries and quantification of temporal heterogeneity within trajectory clusters.

5
Global vascular plants reveal persistent gaps across taxa and ecoregions

Maciel, E. A.

2026-08-28 evolutionary biology 10.64898/2026.08.24.746674 medRxiv
Top 0.1%
12.7%
Show abstract

Biodiversity aggregators such as GBIF provide unprecedented access to global biodiversity data, yet their representativeness remains uneven across space and taxa. This study examined the spatial and taxonomic structure of global vascular plant data available on GBIF. Six filters were applied to the GBIF vascular plant dataset, resulting in the removal of 54% of all records. Together, the filters explained more than 90% of the identified spatial issues, with duplicate and missing coordinates accounting for most of the variation. A higher number of occurrence records was associated with a greater number of spatial issues. Record distributions became progressively more even at finer taxonomic levels, from orders to species. The time series of occurrences for species, genera, and families increased sharply after 1800 and continued to rise, with no apparent stabilisation. Of the 824 ecoregions covered, 73 accounted for 72% of all occurrence records. These ecoregions spanned all continents but were strongly concentrated in Europe, followed by North America and Oceania. The analyses reveal four key patterns: (1) data volume is positively associated with spatial issues; (2) a small number of taxa account for a large proportion of records, whereas many are represented by relatively few; (3) occurrence data aggregated by GBIF have increased continuously since 1800; and (4) record coverage remains highly uneven across the world's ecoregions. These results highlight the substantial contribution of biodiversity data aggregators to expanding access to biological information while demonstrating the persistent spatial and taxonomic biases that shape their contents. Such biases should be explicitly considered when assessing data completeness and quality and when using aggregated occurrence records to infer global biodiversity patterns.

6
A chromosome-level assembly of an aquatic passerine bird, the northern white-throated dipper, Cinclus cinclus cinclus (Linnaeus, 1758)

Strand, M. A.; Toerresen, O. K.; Skage, M.; Ferrari, G.; Tooming-Klunderud, A.; Johnsen, A.; Jakobsen, K. S.

2026-08-24 genomics 10.64898/2026.08.20.746034 medRxiv
Top 0.1%
12.6%
Show abstract

We present a chromosome-level genome assembly of a female Norwegian white-throated dipper (Cinclus cinclus cinclus) generated using Oxford Nanopore Technologies (ONT) long reads and Hi-C scaffolding. The assembly comprises two pseudo-haplotypes, hap1 (1186 Mb) and hap2 (1115 Mb), with 96.7% and 94.4% of sequences assigned to chromosome-scale scaffolds, respectively. Both pseudo-haplotypes contain 40 autosomes, with the Z and W sex chromosomes assigned to hap1. Compared with the PacBio HiFi-based C. c. gularis reference assembly bCinCin1.1.pri, which contains 38 autosomes, sequence represented as a single dot-chromosome (chr 36) is resolved into three distinct dot-chromosomes (chr 36, 39, and 40), a configuration supported by Hi-C contact patterns. BUSCO completeness was high for hap1 (99.2%) and hap2 (95.0%), with 19,003 and 17,746 predicted protein-coding genes, respectively. Compared with the HiFi-based C. c. gularis reference and HiFi-based assemblies generated from the same individual, the ONT-derived assemblies were substantially less fragmented and recovered more sequence from the smallest chromosomes. Synteny was otherwise largely conserved between subspecies. HiFi depletion increased strongly from macrochromosomes to micro- and dot-chromosomes, and HiFi-depleted regions were enriched for repeats and predicted non-B-DNA-associated features, particularly G-quadruplexes and direct repeats, whereas ONT coverage remained comparatively stable. These results show that conventional genome-wide assembly metrics can obscure substantial differences in the recovery of repeat-rich avian dot-chromosomes and highlight the value of chromosome-aware evaluation and ONT sequencing for recovering these regions.

7
dbverse scales spatial omics analysis with embedded analytical databases

Ruiz, E. C.; Jarzabek, V.; Chen, J. G.; Rizvanov, T.; Amin, I.; Dries, R.

2026-08-14 bioinformatics 10.64898/2026.08.09.743742 medRxiv
Top 0.1%
12.5%
Show abstract

Spatial omics datasets are increasing in size and complexity, exceeding the memory of standard computers and thereby limiting data analysis. Here we present dbverse, a framework for larger-than-memory matrix, spatial and genomic data analysis in embedded analytical databases. Benchmarks show dbverse provides orders of magnitude runtime improvements relative to established in-memory and file-backed methods for core operations in single-cell and spatial omics analysis. We integrated dbverse with Giotto Suite, scaling end-to-end preprocessing of millions of cells and enabling spatial alternative polyadenylation analysis as demonstrated on a Visium HD 3' ovarian clear cell carcinoma sample. The dbverse framework provides an interoperable database foundation for larger-than-memory spatial omics analysis on ordinary computers.

8
PanGBank: a large-scale resource of precomputed microbial pangenomes built with PPanGGOLiN

Mainguy, J.; Lemane, T.; Bazin, A.; Arnoux, J.; Gautreau, G.; Medigue, C.; Calteau, A.; Vallenet, D.

2026-08-09 bioinformatics 10.64898/2026.08.05.742796 medRxiv
Top 0.1%
12.5%
Show abstract

PanGBank (https://pangbank.genoscope.cns.fr) is a comprehensive open-access database providing precomputed prokaryotic pangenomes at a broad taxonomic scale. Built upon PPanGGOLiN partitioned pangenome graphs, PanGBank addresses the growing need for large-scale comparative genomics through a standardized, regularly updated, and fully accessible resource. The initial release comprises two complementary collections covering more than 4,600 prokaryotic species from the Genome Taxonomy Database (GTDB), encompassing over 393,000 genomes: GTDB all, maximizing taxonomic and environmental diversity through the inclusion of MAGs and SAGs, and GTDB refseq, focusing on high-quality, annotation-rich genomes. Each species-level pangenome integrates graph-based statistical partitions into persistent, shell, and cloud gene families, together with regions of genomic plasticity (panRGP) and co-localized functional modules (panModule). PanGBank offers multiple access modes, including a REST API, a command-line interface (PanGBank-cli), and an interactive web interface. By combining large-scale pangenome resources with advanced graph-based analyses, PanGBank provides a scalable framework for exploring microbial diversity, genome evolution, functional variation, and the dissemination of adaptive traits across prokaryotic populations, as illustrated by a use case on Acinetobacter baumannii pangenome investigating the distribution and evolution of antimicrobial resistance determinants. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=63 SRC="FIGDIR/small/742796v1_ufig1.gif" ALT="Figure 1"> View larger version (23K): org.highwire.dtl.DTLVardef@186b88dorg.highwire.dtl.DTLVardef@1be33d0org.highwire.dtl.DTLVardef@3bc596org.highwire.dtl.DTLVardef@292401_HPS_FORMAT_FIGEXP M_FIG C_FIG

9
Kiosc: an integrated platform for managing bioinformatics data analysis containers

Marotta, F.; Stolpe, O.; Obermayer, B.; Weiner, J.; Holtgrewe, M.; Beule, D.; Nieminen, M.

2026-08-24 bioinformatics 10.64898/2026.08.20.745983 medRxiv
Top 0.1%
11.8%
Show abstract

In many bioinformatic data analysis projects, it is convenient to visualize plots and results through an interactive web app or dashboard. These interactive reports can then be shared with customers, collaborators, or the general public. Publishing and sharing these apps is not straightforward, becoming especially cumbersome when the number of projects and customers start growing. Docker containers offer a convenient way to package, distribute, and run interactive web apps, and their use is already widespread in the bioinformatics community. We developed Kiosc to simplify the orchestration of containerized web apps, organize them into projects, and regulate access control. We implemented it as a web server based on the Django framework, with a user- and admin-friendly interface as well as a REST API for programmatic tasks. Users can select Docker containers packaging apps like Plotly Dash, Shiny, or Quarto, and configure them to display the results of their analysis. Kiosc runs the containers with the appropriate network configuration and acts as a proxy to the web services running inside the containers. We have been maintaining a Kiosc instance for more than 5 years, serving 321 containers in 150 projects across multiple institutions. In this article, we introduce the main functionality in Kiosc and describe four use-cases that show how Kiosc can prove helpful to the broader bioinformatics community, such as configuring and running web apps for the interactive visualization of workflow results, and publishing companion apps for scientific articles. Kiosc is a self-hosted platform for publishing web apps, which doesn't require significant expertise in either Docker or network administration to be deployed. It provides a similar service to Kubernetes, but with a convenient web interface and much lower administration overhead.

10
CyChat: a conversational Cytoscape app for no-code, reproducible network analysis

Liebold, J.; Stahl, M.; Schulze, J.-O.; Razavi, M. M.; Bader, G. B.; Kurtz, S.; Baumbach, J.

2026-09-01 bioinformatics 10.64898/2026.08.28.747833 medRxiv
Top 0.1%
9.7%
Show abstract

Network-based analyses of molecular interactions are useful for interpreting high-throughput omics data and identifying therapeutic targets. Cytoscape is the standard platform for these tasks, but users face a trade-off between accessible graphical workflows that are difficult to document and reproducible automation in Python or R that requires programming expertise. General-purpose coding assistants can generate Cytoscape Automation scripts, but remain external to Cytoscape. We present CyChat, a Cytoscape Desktop app that integrates a chat interface and a large language model (LLM) agent into the application. CyChat translates natural language into executable Cytoscape Automation workflows, runs generated Python code, and exports chat sessions with executed code as standalone Jupyter notebooks. To reduce setup barriers, CyChat includes an embedded Python runtime and supports both cloud-based and locally hosted LLMs. CyChat was evaluated across ten Cytoscape workflows using seven LLM providers, each represented by one LLM. The strongest configuration achieves a pass rate above 99%. In a qualitative evaluation based on a published network visualization, CyChat completes the task in 1.5-5 minutes, compared with 15-20 minutes for manual GUI workflows by computational biologists. CyChat is available through the Cytoscape App Store at https://apps.cytoscape.org/apps/cychat.

11
PlantAI: A Multi-Agent System for Plant Functional Genomics Analysis and Biological Knowledge Interpretation

Wu, T.; Yang, Z.; Shi, J.; Zou, M.; Wu, Y.; Jiang, S.; Xia, C.; Kong, L.; Yang, L.; Xia, Z.

2026-08-18 bioinformatics 10.64898/2026.08.14.744760 medRxiv
Top 0.2%
9.4%
Show abstract

Plant functional genomics requires the integration of sequence, expression, evolutionary, regulatory and literature evidence. However, the corresponding analyses are often distributed across disparate programs, scripts and databases, creating substantial barriers to task organization and result interpretation. Here, we present PlantAI, a multi-agent system that integrates bioinformatics analysis, project-level process tracking and knowledge-assisted interpretation. A Main Agent coordinates two complementary routes: an analysis route that invokes bioinformatics tools for RNA-seq and gene-family analyses, and a knowledge route that uses PlantAI-RAG for knowledge retrieval and evidence synthesis. PlantAI-RAG currently contains 31,207 plant-science literature records, comprising approximately 3.82 million normalized entities and 8.25 million literature-supported relation assertions. In an evaluation using plant-science questions, it achieved a Gold evidence-assertion recall of 86.7%, while strict accuracy ranged from 77% to 82% across three independent evaluator models. We further demonstrate an end-to-end task using 24 rice RNA-seq libraries collected under salt stress, spanning transcriptome analysis, candidate-family screening, HXK/HKL family analysis and knowledge-assisted interpretation, and prioritize OsHXK8 for experimental validation. By preserving analysis artifacts, run manifests, logs and environment records, PlantAI supports result verification and repeat execution while linking project-derived results to traceable literature evidence. Together, these capabilities provide an integrated and auditable framework to support plant functional genomics research.

12
Interpretable biomarker discovery from small-sample microarray datasets using XGBoost rank aggregation and SVM-RFECV

Prapty, M. M.; Rahman, M. S.

2026-08-21 bioinformatics 10.64898/2026.08.13.744652 medRxiv
Top 0.2%
9.0%
Show abstract

MotivationHigh-dimensional microarray datasets remain valuable for cancer biomarker discovery, but their small sample sizes make robust and interpretable feature selection challenging. Efficient workflows are needed to derive compact gene signatures while preserving biological interpretability. ResultsWe developed a two-stage biomarker-discovery workflow that combines cross-validated XGBoost rank aggregation with support vector machine recursive feature elimination and cross-validation (SVM-RFECV) to identify compact candidate biomarker panels. The workflow was evaluated on 21 public binary and multiclass microarray datasets using repeated stratified cross-validation for internal validation. Across the dataset collection, the selected panels demonstrated strong internal discriminative performance while remaining sufficiently compact for downstream biological interpretation. SHAP analysis identified dataset- and class-specific discriminative genes, and functional enrichment analysis supported the biological coherence of representative consensus signatures. The proposed workflow provides an interpretable and reproducible framework for candidate biomarker discovery from small-sample microarray datasets. AvailabilitySource code and processed outputs are freely available at https://github.com/mashiyat-mahjabin-prapty/microarray-feature-selection.

13
ASAREE: An Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation

Moran, J.; Freda, P. J.; Ghosh, A.; Hernandez, M. E.; Moore, J. H.

2026-08-25 bioinformatics 10.64898/2026.08.20.746074 medRxiv
Top 0.2%
9.0%
Show abstract

Summary: Agentic AI platforms enable the engineering of autonomous workflows but are not designed for experimentation and hypothesis testing. ASAREE (Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation), is an open-source platform to address this gap. ASAREE creates agents, connects to MCP servers and tools, and designs factorial experiments through a visual interface or Python SDK. It records a full provenance trace for every run and routes all model calls through a provider-agnostic bridge that supports local deployments, ensuring data privacy. As a use-case, we use ASAREE to evaluate key design choices in a mutli-agent machine learning pipeline. Across a 2 x 2 x 2 factorial design, more advanced models, greater reasoning effort, and critic agent use significantly increased compute time, token use, cost, and feature count without improving predictive performance. The lowest-cost baseline, Claude Sonnet 5 with medium effort and no critic, achieved the highest mean PR AUC while Claude Opus 5 with extra high effort and a critic agent cost 15.5x more (USD) and ran 13.1x longer while performing worse on average. These findings highlight ASAREE as a robust framework for evaluating agentic system performance and resource efficiency.

14
nf-core/genomeqc: a best-practice pipeline for comparing genome and assembly quality

Wyatt, C. D. R.; Duarte Frutos, F.; Turner, S. D.; Cerqueira De Araujo, A.; Rashid, U.; Begley, V.; Sumner, S.

2026-08-20 bioinformatics 10.64898/2026.08.20.745971 medRxiv
Top 0.2%
8.8%
Show abstract

The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome quality therefore requires integrating multiple complementary quality metrics that are often generated by independent tools. Here, we present nf-core/genomeqc, a workflow for assessing and comparing genome assemblies. The pipeline accepts RefSeq/GenBank accessions for automatic genome and annotation retrieval, or local genome (FASTA) and annotation (GFF3/GTF) files. It integrates complementary analyses of assembly contiguity, gene completeness, annotation quality, repeat content and other quality metrics using tools such as BUSCO, QUAST, Merqury, and AGAT, before combining the results on a phylogenetic tree for visualisation and comparison across species. GenomeQC is implemented in Nextflow within the nf-core framework, providing an accessible, reproducible, scalable and community-driven workflow for genome quality assessment.

15
Dynamic Hierarchical Interleaved Bloom Filter: An Updatable Index for Large-Scale Fast Sequence Search

Seiler, E.; Willemsen, M.; Piro, V. C.; Reinert, K.

2026-08-30 bioinformatics 10.64898/2026.08.26.747224 medRxiv
Top 0.2%
8.7%
Show abstract

Motivation: A continued decrease in sequencing costs has facilitated the exponential increase in available sequencing data, with public databases like the European Nucleotide Archive (ENA) and Sequence Read Archive (SRA) reaching well in the order of petabases. This has been the incentive to develop more scalable tools for common bioinformatics tasks. One such task is the approximate searching of short sequence patterns like genes or reads in reference data sets. In recent years, a variety of indexing data structures have been proposed for searching large sequencing databases. The state-of-the-art index, the Hierarchical Interleaved Bloom Filter (HIBF) was first-in-class to index one million samples. To be useful for expanding repositories, it must be extended to support dynamic updates. Results: In this paper, we introduce a scalable and updatable sequence-search index by extending the HIBF with partial rebuilding to support efficient updates. We demonstrate the Dynamic HIBF's capacity for large-scale data by iteratively creating an index from over 100 TB of compressed reads across more than 39,000 full human RNA-Seq samples, updated in consecutive batches of 100. To benchmark against state-of-the-art tools, we evaluated incremental performance on a subset of 5,000 samples sub-sampled to 1% of their original read depth. In this comparative setting, the dynamic HIBF completed the sequential insertion of all 5,000 samples within 5 hours--24 to 65 times faster than competing methods and twice as fast as the static HIBF.

16
nf_xpatial: A Reproducible Framework for Standardized Preprocessing and Clustering of Xenium Data

Potter, L. A.; Trull, A.; Kumar, N.; Drake, O. R.; Nogueira, M.; Peters, J.; Heinsbroek, J. A.; Day, J. J.; Worthey, E. A.; Ianov, L.

2026-08-29 bioinformatics 10.64898/2026.08.25.747147 medRxiv
Top 0.2%
8.0%
Show abstract

Recent advances in spatial transcriptomics have enabled the profiling of increasingly larger numbers of genes while retaining single-cell and subcellular resolution in situ. However, standardized bioinformatics workflows for analyzing these datasets have lagged behind, with existing pipelines focusing primarily on image processing and cell segmentation. To address this gap, we present nf_xpatial, a best-practices Nextflow pipeline for the downstream analysis of 10x Genomics Xenium data. The pipeline performs quality control, filtering, log and cell area normalization, multi-sample integration, and both expression-driven and spatially informed clustering across systematic parameter sweeps, allowing users to evaluate and compare clustering resolutions and spatial modeling parameters within a single reproducible run. Overall, nf_xpatial streamlines the processing of Xenium data from platform outputs to integrated single-cell and spatial clustering datasets, providing a standardized starting point from which biologists can fine-tune parameters and proceed to hypothesis-driven spatial analyses.

17
Correction of the cytosine deamination artifacts in FFPE-based sequencing experiments

Płonka, W.; Kostka, D.; Lalik, A.; Kurpas, M.; Dinh, K. N.; Sitkiewicz, M.; Kimmel, M.; Rzyman, W.; Jaksik, R.

2026-08-19 bioinformatics 10.64898/2026.08.11.744151 medRxiv
Top 0.2%
8.0%
Show abstract

Formalin-fixed, paraffin-embedded (FFPE) tissues remain an essential resource for molecular studies, yet formalin-induced cytosine deamination introduces characteristic C>T/G>A artifacts that compromise the accuracy of next-generation sequencing (NGS) analyses. Numerous computational methods and enzymatic DNA repair strategies have been proposed to reduce these artifacts, but no systematic comparison across tools and experimental conditions exists. Here, we evaluate the performance of seven computational approaches (SOBDetector, Ideafix, MicroSEC, FFPolish, DeepOmics FFPE/FFPE-PLUS, FFPErase) together with the NEBNext(R) FFPE DNA Repair Mix v2, a multi-enzyme repair system applied during DNA preparation. Using three independent datasets, one based on whole genome sequencing (CGCI-BL) and two on whole exome sequencing (TCGA-PC and SUT-LUAD, the latter containing enzymatically repaired samples), and matched fresh-frozen samples as the gold standard, we assess precision, sensitivity, and artifact reduction efficiency across all methods. We further examine the potential synergy between enzymatic repair and post-sequencing computational filtering. Our results provide practical guidelines for FFPE artifact correction and demonstrate that enzymatic treatment provides the best results, while among the computational methods, FFPErase offers the most robust reduction of cytosine deamination artifacts while maximizing the retention of true somatic variants. KEY MESSAGESO_LIFormalin fixation in FFPE samples introduces artifacts that can significantly affect the accuracy of NGS analyses. C_LIO_LIAmong the evaluated approaches, enzymatic repair using NEBNext(R) FFPE DNA Repair Mix v2 achieves the most effective reduction of sequencing artifacts. C_LIO_LIComputational methods vary in performance, with FFPErase showing the most robust balance between artifact removal and retention of true somatic variants. C_LIO_LICombining enzymatic repair with computational filtering did not lead to consistent improvements in performance across datasets. C_LI

18
SeqDesk: a sequencing-facility management system for standards-compliant and FAIR (meta)data submission

Muench, P. C.; Robertson, G.; McHardy, A. C.

2026-08-11 bioinformatics 10.64898/2026.08.05.743014 medRxiv
Top 0.2%
7.8%
Show abstract

Achieving FAIR compliance requires both standardized metadata and infrastructure for data deposition, yet in practice a large fraction of sequencing studies is still published without the persistent, standards-compliant metadata that reuse depends on. Collecting MIxS-compliant metadata is complex: environment-specific checklists can contain hundreds of fields, and the effort is magnified when metadata is assembled retrospectively at publication time rather than captured throughout the project. We developed SeqDesk, an open-source data management system for sequencing facilities that is designed so that FAIR-compliant public data is produced as the natural output of routine operations. Its current scope is microbial sequencing data, covering metagenomes as well as isolate genomes, for which it supports the corresponding MIxS checklists. SeqDesk gives a sequencing facility a configurable order-and-tracking system for sequencing projects, captures and validates MIxS-compliant metadata aligned with ENA checklists at project initiation, runs bioinformatics analyses through Nextflow pipelines, and brokers submission to the European Nucleotide Archive, all within the institutions own infrastructure. By embedding standards-compliant metadata capture into the sequencing-facility workflow rather than bolting it on at submission, SeqDesk shortens the path from sample to reusable public data. The underlying checklist model is generic, so support can be extended to further data types and metadata standards beyond the microbial domain. SeqDesk is free and open source under the Apache 2.0 licence and available at https://seqdesk.org, with a live demonstration at https://seqdesk.org/#demo.

19
Prediction of plant organismal complexity based on transcription factor annotation: an AI approach

Varshney, D.; Tajjar, M. H.; de Vries, J.; Hutter, F.; Rensing, S. A.

2026-08-22 evolutionary biology 10.64898/2026.08.18.745462 medRxiv
Top 0.2%
7.7%
Show abstract

How morphological complexity evolves is still enigmatic. While there is evidence in algae and plants as well as animals that diversification of the repertoire of transcription factors (TF) is causative for evolution of organismal complexity, there are many examples from lineages that follow their own way of complexity evolution, for example by expansion of particular families. For land plants, correlation of the size of the TF complement with number of cell types (as a proxy for morphological complexity) has been shown, and several families were identified as candidates to drive complexity evolution. Here, we expand a previously available dataset of cell type numbers from 12 to 82 proteomes and introduce a four class body plan scheme. We find that the total TF complement correlates with the number of cell types of Archaeplastida (primary plastid bearing plants and algae). We used TabPFN (Tabular Prior-data Fitted Network) for binary (uni- vs. multicellularity) as well as for four class Bauplan classification. TabPFN is able to predict the morphological complexity with high accuracy. This approach allows to determine organismal complexity based on the gene space of an organism. Based on our results, we can confirm that plant morphological evolution is driven by gain and expansion of TF families.

20
pastForward: a Snakemake pipeline for ancient and historical DNA with eukaryote-wide taxonomic screening and tracking of copy-number variation

Saadain, S.; Kapun, M.; Kofler, R.

2026-08-13 bioinformatics 10.64898/2026.08.07.743613 medRxiv
Top 0.2%
7.7%
Show abstract

Ancient and historical DNA has the potential to resolve many open questions in biology. While pipelines for processing ancient and historical DNA exist, none combine user-friendly, configurable processing with copy number variation tracking and targeted taxonomic profiling. Therefore, we developed pastForward, a fully automated Snakemake pipeline that integrates all analysis steps from raw reads to damage-rescaled BAM files in a single reproducible workflow. It performs ancient and historical DNA processing, including adapter trimming, read merging, deduplication, damage assessment, quality rescaling, and generates interactive reports summarizing the endogenous read content, library complexity, and breadth and depth coverage statistics. These reports allow users to rapidly assess the quality of sequencing data. It handles single- and paired-end NGS libraries. Mapping to multiple reference sequences is supported, facilitating co-analysis of host and endosymbiont sequences and genotyping of marker genes such as COI. pastForward further integrates two novel tools. ECMSD (Efficient Comprehensive Mitochondrial Sequence Detector) screens each library for eukaryotic DNA by aligning reads against a mitochondrial reference database. The presence of bacteria, archaea and viruses is detected in parallel with Centrifuge. REVEAL (Read-based Estimation and visualization of Element Abundance and Loci) quantifies and visualizes copy number variation of genetic features, such as transposable elements (TEs) or gene duplications. Two case studies demonstrate the usage of the pipeline. Using pastForward on dog genomic time series, including Neolithic samples, we confirm that the copy number of AMY2B, which encodes the starch-digesting enzyme amylase, increased during domestication. From historical D. melanogaster genomes, we recover the recent invasion of the transposable element opus. It is absent in specimens from the 1800s and present from 1933 onward. By efficiently processing large numbers of samples, pastForward facilitates longitudinal tracking of genomic features in diverse species.