FrameRate: learning the coding potential of unassembled metagenomic reads
Liu-Wei, W.; Aubrey, W.; Clare, A.; Hoehndorf, R.; Creevey, C. J.; Dimonaco, N. J.
Show abstract
MotivationMetagenomic assembly is a slow and computationally intensive process and despite needing iterative rounds for improvement and completeness the resulting assembly often fails to incorporate many of the input sequencing reads. This is further complicated when there is reduced read-depth and/or artefacts which result in chimeric assemblies both of which are especially prominent in the assembly of metagenomic datasets. Many of these limitations could potentially be overcome by exploiting the information content stored in the reads directly and thus eliminating the need for assembly in a number of situations. ResultsWe explored the prediction of coding potential of DNA reads by training a machine learning model on existing protein sequences. Named FrameRate, this model can predict the coding frame(s) from unassembled DNA sequencing reads directly, thus greatly reducing the computational resources required for genome assembly and similarity-based inference to pre-computed databases. Using the eggNOG-mapper function annotation tool, the predicted coding frames from FrameRate were functionally verified by comparing to the results from full-length protein sequences reconstructed with an established metagenome assembly and gene prediction pipeline from the same metagenomic sample. FrameRate captured equivalent functional profiles from the coding frames while reducing the required storage and time resources significantly. FrameRate was also able to annotate reads that were not represented in the assembly, capturing this missing information. As an ultra-fast read-level assembly-free coding profiler, FrameRate enables rapid characterisation of almost every sequencing read directly, whether it can be assembled or not, and thus circumvent many of the problems caused by contemporary assembly workflows. Availabilityhttps://github.com/NickJD/FrameRate Contactliuwei.wang@fu-berlin.de and nicholas@dimonaco.co.uk
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ganon: precise metagenomics classification against large and up-to-date sets of reference sequences 96%
- panRGP: a pangenome-based method to predict genomic islands and explore their diversity 95%
- No one tool to rule them all: Prokaryotic gene prediction tool performance is highly dependent on the organism of study 95%
Similar papers in this journal
Similar papers in this journal
- HiCBin: Binning metagenomic contigs and recovering metagenome-assembled genomes using Hi-C contact maps 95%
- MetaBinner: a high-performance and stand-alone ensemble binning method to recover individual genomes from complex microbial communities 95%
- Efficient inference of large pangenomes with PanTA 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.