SLOW5: a new file format enables massive acceleration of nanopore sequencing data analysis
Gamaarachchi, H.; Samarakoon, H.; Jenner, S. P.; Ferguson, J. M.; Amos, T. G.; Hammond, J. M.; Saadat, H.; Smith, M. A.; Parameswaran, S.; Deveson, I. W.
Show abstract
Nanopore sequencing is an emerging genomic technology with great potential. However, the storage and analysis of nanopore sequencing data have become major bottlenecks preventing more widespread adoption in research and clinical genomics. Here, we elucidate an inherent limitation in the file format used to store raw nanopore data - known as FAST5 - that prevents efficient analysis on high-performance computing (HPC) systems. To overcome this, we have developed SLOW5, an alternative file format that permits efficient parallelisation and, thereby, acceleration of nanopore data analysis. For example, we show that using SLOW5 format, instead of FAST5, reduces the time and cost of genome-wide DNA methylation profiling by an order of magnitude on common HPC systems, and delivers consistent improvements on a wide range of different architectures. With a simple, accessible file structure and a ~25% reduction in size compared to FAST5, SLOW5 format will deliver substantial benefits to all areas of the nanopore community.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepCell Kiosk: Scaling deep learning-enabled cellular image analysis with Kubernetes 94%
- Haplotype-aware variant calling enables high accuracy in nanopore long-reads using deep neural networks 94%
- Multiscale Analysis of Pangenome Enables Improved Representation of Genomic Diversity For Repetitive And Clinical Relevant Genes 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.