ViReaDB: A user-friendly database for compactly storing viral sequence data and rapidly computing consensus genome sequences
Moshiri, N.
Show abstract
MotivationIn viral molecular epidemiology, reconstruction of consensus genomes from sequence data is critical for tracking mutations and variants of concern. However, storage of the raw sequence data can become prohibitively large, and computing consensus genome from sequence data can be slow and requires bioinformatics expertise. ResultsViReaDB is a user-friendly database system for compactly storing viral sequence data and rapidly computing consensus genome sequences. From a dataset of 1 million trimmed mapped SARS-CoV-2 reads, it is able to compute the base counts and the consensus genome in 16 minutes, store the reads alongside the base counts and consensus in 50 MB, and optionally store just the base counts and consensus (without the reads) in 300 KB. AvailabilityViReaDB is freely available on PyPI (https://pypi.org/project/vireadb) and on GitHub (https://github.com/niemasd/ViReaDB) as an open-source Python software project. Contactniema@ucsd.edu
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MerCat2: a versatile k-mer counter and diversity estimator for database-independent property analysis obtained from omics data 95%
- AnnSQL: A Python SQL-based package for fast large-scale single-cell genomics analysis using minimal computational resources 94%
- Panalyze: automated virus pangenome variation graph construction and analysis 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.