Back

Engineering the Compression of Sequencing Reads

Kowalski, T.; Grabowski, S. P.

2020-05-02 bioinformatics
10.1101/2020.05.01.071720 bioRxiv
Show abstract

MotivationFASTQ remains among the widely used formats for high-throughput sequencing data. Despite advances in specialized FASTQ compressors, they are still imperfect in terms of practical performance tradeoffs. ResultsWe present a multi-threaded version of Pseudogenome-based Read Compressor (PgRC), an in-memory algorithm for compressing the DNA stream, based on the idea of building an approximation of the shortest common superstring over high-quality reads. The current version, v1.2, practically preserves the compression ratio and decompression speed of the previous one, reducing the compression time by a factor of about 4-5 on a 6-core/12-thread machine. AvailabilityPgRC 1.2 can be downloaded from https://github.com/kowallus/PgRC. Contactsgrabow@kis.p.lodz.pl

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.