NYX: Format-aware, learned compression across omics file types
Patsakis, M.; Chronopoulos, T.; Mouratidis, I.; Georgakopoulos-Soares, I.
Show abstract
Genomic data repositories continue to grow as sequencing technologies improve, with the NCBI SRA alone exceeding 47 PB. General-purpose compressors treat bioinformatics files as unstructured byte streams and fail to exploit the structured nature of omics data. We present NYX, a format-aware compression system for FASTA, FASTQ, VCF, WIG, H5AD, and BED files. NYX combines lightweight, reversible preprocessing and is build upon the OpenZL framework to take advantage of inherent data structure, delivering high compression ratios while preserving fast and lossless compression. Across representative datasets in the target formats, NYX achieves substantially higher speed than format-specific compressors while maintaining or improving compression ratio.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- kmtricks: Efficient and flexible construction of Bloom filters for large sequencing data collections 96%
- FroM Superstring to Indexing: a space-efficient index for unconstrained k-mer sets using the Masked Burrows-Wheeler Transform (MBWT) 95%
- ScaleSC: A superfast and scalable single cell RNA-seq data analysis pipeline powered by GPU. 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.