Lapidary: Identifying and reporting amino acid sequences in metagenomes using sequence reads and Diamond
Bloomfield, S. J.; Zomer, A. L.; Mather, A. E.
Show abstract
Genome and metagenome comparisons rely on identifying genetic elements that differ or are in common between samples. These genetic elements can be identified by assembling sequenced reads and identifying the genetic element in the assembly, or by aligning nucleotide sequences in the reads to the nucleotide sequences of a reference genetic element. The first relies on the complete assembly of the genetic element of interest, and the second relies on a reference sequence represented in nucleotides. This is particularly challenging with metagenome data, where the genetic elements, including genes, are often fragmented because sequences are shared between different species in the metagenomic data, resulting in contig breaks in or around genetic elements. This presents a difficulty when identifying genetic elements through the first approach. A common approach with metagenomes is to map reads against reference nucleotide sequences and extract the depth and coverage from those reference sequences. However, currently no software exists to identity and report genetic elements using DNA-protein alignments in metagenomes. We have developed the software Lapidary to identify the identity, coverage, depth, and most likely sequence of amino acid sequences from both genome and metagenome read files. We tested the effectiveness of the method against simulated, genomic and metagenomic read datasets. Lapidary is more sensitive than assembly methods for metagenomic data that often have fragmented assemblies but is less sensitive when assemblies are more complete, as is the case with genomic data.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 95%
- MinION Sequencing of colorectal cancer tumour microbiomes - a comparison with amplicon-based and RNA-Sequencing 95%
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 95%
Similar papers in this journal
- Optimizing experimental design for genome sequencing and assembly with Oxford Nanopore Technologies 93%
- The Crown Pearl V2: an improved genome assembly of the European freshwater pearl mussel Margaritifera margaritifera (Linnaeus, 1758) 93%
- Whole Genome Sequencing and Assembly of the House Sparrow, Passer domesticus 93%
Similar papers in this journal
- Contamination detection and microbiome exploration with GRIMER 94%
- PathoGFAIR: a collection of FAIR and adaptable (meta)genomics workflows for (foodborne) pathogens detection and tracking 94%
- Global ocean resistome revealed: exploring Antibiotic Resistance Genes (ARGs) abundance and distribution on TARA oceans samples through machine learning tools 94%
Similar papers in this journal
- FANGORN: A quality-checked and publicly available database of full-length 16S-ITS-23S rRNA operon sequences 95%
- Identifying the best PCR enzyme for library amplification in NGS 94%
- Evaluation of the accuracy of bacterial genome reconstruction with Oxford Nanopore R10.4.1 long-read-only sequencing 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.