Back

Embarrassingly_FASTA: Enabling Recomputable, Population-Scale Pangenomics by Reducing Commercial Genome Processing Costs from $100 to less than $1

Walsh, D. J.; Njie, e. G.

2026-02-04 bioinformatics
10.64898/2026.02.02.703356 bioRxiv
Show abstract

Computational preprocessing has become the dominant bottleneck in genomics, frequently exceeding sequencing costs and constraining population-scale analysis, even as large repositories grow from tens of petabytes toward exabyte-scale storage to support World Genome Models. Legacy CPU-based workflows require many hours to days per 30x human genome, driving many repositories to distribute aligned or derived intermediates such as BAM and VCF files rather than raw FASTQ data. These intermediates embed reference- and model-dependent assumptions that limit reproducibility and impede reanalysis as reference genomes, including pangenomes, continue to evolve. Although recent work has established that GPUs can dramatically accelerate genomic pipelines, enabling large-cohort processing to shrink from years to days given sufficient parallelism, such workflows remain cost-prohibitive. Here, we introduce Embarrassingly_FASTA, a GPU-accelerated preprocessing pipeline built on NVIDIA Parabricks that fundamentally changes the economics of genomic data management. By rendering intermediate files transient rather than archival, Embarrassingly_FASTA enables retention of raw FASTQ data and reliable use of highly discounted ephemeral cloud infrastructure such as spot instances, reducing compute spend from [~]$17/genome (CPU on-demand) to <$1/genome (GPU spot), and commercial secondary-analysis pricing from [~]$120/genome to compute spend under $1/genome. We demonstrate the impact of this efficiency using a simulated large-cohort pangenome build-up (using variant-union accumulation as a proxy for diversity growth) in Caenorhabditis elegans and humans, highlighting the long tail of unsampled human genetic diversity. Beyond GPU kernels, Embarrassingly_FASTA contributes a transient-intermediate lifecycle and spot-friendly orchestration that makes FASTQ retention and routine recomputation economically viable. Embarrassingly_FASTA thus provides enabling infrastructure for recomputable, population-scale pangenomics and next-generation genomic models. Non-Expert DescriptionReading a persons complete DNA sequence has become fast and inexpensive, but turning that raw data into something scientists can analyze is now one of the biggest obstacles in modern genetics. Today, processing a single genome can take many hours or even days, which makes it difficult and expensive to study large populations or reanalyze data when better methods become available. As a result, many databases store only partially processed DNA instead of the original data, limiting future discoveries. In this work, we present a new system that dramatically speeds up this processing step using graphics processing units (GPUs), the same type of hardware used in modern artificial intelligence. With our approach, a human genome can be processed in about 35 minutes instead of more than 15 hours, and at a fraction of the cost. This makes it practical to keep the original DNA data and reprocess it whenever new tools or reference genomes become available, rather than being locked into outdated results. We also show that this speed and affordability allow researchers to explore genetic diversity at an unprecedented scale. By analyzing both human genomes and those of a small worm species commonly used in research, we demonstrate how new genetic variations continue to emerge as more individuals are studied, especially in humans, where much diversity remains unexplored. Overall, our work helps remove a major barrier to studying DNA at population scale and lays the foundation for future genetic models that could better explain disease, evolution, and human health. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=112 SRC="FIGDIR/small/703356v2_ufig1.gif" ALT="Figure 1"> View larger version (69K): org.highwire.dtl.DTLVardef@11b8fc3org.highwire.dtl.DTLVardef@7b8b60org.highwire.dtl.DTLVardef@fb5a3borg.highwire.dtl.DTLVardef@1e11276_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.