Back

CRAM compression: practical across-technologies considerations for large-scale sequencing projects

Al Ali, A.; Kandavel, P. K.; Al Mabrazi, H.; Carvalho, G.; Kusuma, V.; Katagi, G.; Elavalli, S.; Yousif, A.; Akhter, M. R.; Mafofo, J.; Magalhaes, T.; Quilez, J.

2022-12-22 bioinformatics
10.1101/2022.12.21.521516 bioRxiv
Show abstract

CRAM is an efficient format to store high-throughput sequencing data and it has been widely adopted. We thus plan to use CRAM for the Emirati Genome Program, which aims to sequence the genomes of ~1 million nationals in the United Arab Emirates using short- and long-read sequencing technologies (Illumina, MGI and Oxford Nanopore Sequencing). We conducted a pilot study on the three technologies before start using CRAM at scale. We found CRAM achieved 40-70% compression depending on the sequencing platform. As expected, CRAM compression was data lossless and did not alter variant calls. In our cloud, we observed compression speeds 0.7-1.4 GB per minute, varying on the sequencing platform too. This translates into ~1-2 hours using a single CPU to compress a ~30X human whole-genome sequencing sample. Despite its wide use, we found little publicly available information about CRAM compression rate, speed, losslessness and parallelization, especially across many sequencing platforms. This work will have direct application for Emirati Genome Program and provide practical considerations for other large-scale sequencing efforts.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.