Logan: Planetary-Scale Genome Assembly Surveys Life's Diversity
Chikhi, R.; Raffestin, B.; Korobeynikov, A.; Edgar, R. C.; Babaian, A.
Show abstract
The breadth of lifes diversity is unfathomable, but public nucleic acid sequencing data offers a window into the dispersion and evolution of genetic diversity across Earth. However the rapid growth and accumulation of sequence data have outpaced efficient analysis capabilities. The largest collection of freely available sequencing data is the Sequence Read Archive (SRA), comprising 27.3 million datasets or 5 x 1016 basepairs. To realize the potential of the SRA, we constructed Logan, a massive sequence assembly transforming short reads into long contigs and compressing the data over 100-fold, enabling highly efficient petabase-scale analysis. We created Logan-Search, a k-mer index of Logan for free planetary-scale sequence search, returning matches in minutes. We used Logan contigs to identify >200 million plastic-degrading enzyme homologs, and validate novel enzymes with catalytic activities exceeding current reference standards. Further, we vastly expand the known diversity of proteins (30-fold over UniRef50), plasmids (22-fold over PLSDB), P4 satellites (4.5-fold), and the recently described Obelisk RNA elements (3.7-fold). Logan also enables ecological and biomedical data mining, such as global tracking of antimicrobial resistance genes and the characterization of viral reactivation across millions of human BioSamples. By transforming the SRA, Logan democratizes access to the worlds public genetic data and opens frontiers in biotechnology, molecular ecology, and global health.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Uncovering 2-D toroidal representations in grid cell ensemble activity during 1-D behavior 97%
- Widespread transfer of mobile antibiotic resistance genes within individual gut microbiomes revealed through bacterial Hi-C 96%
- Ascertaining cells' synaptic connections and RNA expression simultaneously with massively barcoded rabies virus libraries 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.