Embarrassingly_FASTA: Enabling Recomputable, Population-Scale Pangenomics by Reducing Commercial Genome Processing Costs from $100 to less than $1
Walsh, D. J.; Njie, e. G.
Show abstract
Computational preprocessing has become the dominant bottleneck in genomics, frequently exceeding sequencing costs and constraining population-scale analysis, even as large repositories grow from tens of petabytes toward exabyte-scale storage to support World Genome Models. Legacy CPU-based workflows require many hours to days per 30x human genome, driving many repositories to distribute aligned or derived intermediates such as BAM and VCF files rather than raw FASTQ data. These intermediates embed reference- and model-dependent assumptions that limit reproducibility and impede reanalysis as reference genomes, including pangenomes, continue to evolve. Although recent work has established that GPUs can dramatically accelerate genomic pipelines, enabling large-cohort processing to shrink from years to days given sufficient parallelism, such workflows remain cost-prohibitive. Here, we introduce Embarrassingly_FASTA, a GPU-accelerated preprocessing pipeline built on NVIDIA Parabricks that fundamentally changes the economics of genomic data management. By rendering intermediate files transient rather than archival, Embarrassingly_FASTA enables retention of raw FASTQ data and reliable use of highly discounted ephemeral cloud infrastructure such as spot instances, reducing compute spend from [~]$17/genome (CPU on-demand) to <$1/genome (GPU spot), and commercial secondary-analysis pricing from [~]$120/genome to compute spend under $1/genome. We demonstrate the impact of this efficiency using a simulated large-cohort pangenome build-up (using variant-union accumulation as a proxy for diversity growth) in Caenorhabditis elegans and humans, highlighting the long tail of unsampled human genetic diversity. Beyond GPU kernels, Embarrassingly_FASTA contributes a transient-intermediate lifecycle and spot-friendly orchestration that makes FASTQ retention and routine recomputation economically viable. Embarrassingly_FASTA thus provides enabling infrastructure for recomputable, population-scale pangenomics and next-generation genomic models. Non-Expert DescriptionReading a persons complete DNA sequence has become fast and inexpensive, but turning that raw data into something scientists can analyze is now one of the biggest obstacles in modern genetics. Today, processing a single genome can take many hours or even days, which makes it difficult and expensive to study large populations or reanalyze data when better methods become available. As a result, many databases store only partially processed DNA instead of the original data, limiting future discoveries. In this work, we present a new system that dramatically speeds up this processing step using graphics processing units (GPUs), the same type of hardware used in modern artificial intelligence. With our approach, a human genome can be processed in about 35 minutes instead of more than 15 hours, and at a fraction of the cost. This makes it practical to keep the original DNA data and reprocess it whenever new tools or reference genomes become available, rather than being locked into outdated results. We also show that this speed and affordability allow researchers to explore genetic diversity at an unprecedented scale. By analyzing both human genomes and those of a small worm species commonly used in research, we demonstrate how new genetic variations continue to emerge as more individuals are studied, especially in humans, where much diversity remains unexplored. Overall, our work helps remove a major barrier to studying DNA at population scale and lays the foundation for future genetic models that could better explain disease, evolution, and human health. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=112 SRC="FIGDIR/small/703356v2_ufig1.gif" ALT="Figure 1"> View larger version (69K): org.highwire.dtl.DTLVardef@11b8fc3org.highwire.dtl.DTLVardef@7b8b60org.highwire.dtl.DTLVardef@fb5a3borg.highwire.dtl.DTLVardef@1e11276_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 96%
- Analysis-ready VCF at Biobank scale using Zarr 95%
- Identifying, understanding, and correcting technical biases on the sex chromosomes in next-generation sequencing data 95%
Similar papers in this journal
- Blackbird: structural variant detection using synthetic and low-coverage long-reads 95%
- AnnSQL: A Python SQL-based package for fast large-scale single-cell genomics analysis using minimal computational resources 94%
- Comparative genome analysis using sample-specific string detection in accurate long reads 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.