Back

Autometa 2: A versatile tool for recovering genomes from highly-complex metagenomic communities

Rees, E. R.; Uppal, S.; Clark, C.; Lail, A. J.; Waterworth, S. C.; Roesemann, S. D.; Wolf, K. A.; Kwan, J. C.

2023-09-05 bioinformatics
10.1101/2023.09.01.555939 bioRxiv
Show abstract

In 2019, we developed Autometa, an automated binning pipeline that is able to effectively recover metagenome-assembled genomes from complex environmental and non-model host-associated microbial communities. Autometa has gained widespread use in a variety of environments and has been applied in multiple research projects. However, the genome-binning workflow was at times overly complex and computationally demanding. As a consequence of Autometas diverse application, non-technical and technical researchers alike have noted its burdensome installation and inefficient as well as error-prone processes. Moreover its taxon-binning and genome-binning behaviors have remained obscure. For these reasons we set out to improve its accessibility, efficiency and efficacy to further enable the research community during their exploration of Earths environments. The highly augmented Autometa 2 release, which we present here, has vastly simplified installation, a graphical user interface and a refactored workflow for transparency and reproducibility. Furthermore, we conducted a parameter sweep on standardized community datasets to show that it is possible for Autometa to achieve better performance than any other binning pipeline, as judged by Adjusted Rand Index. Improvements in Autometa 2 enhance its accessibility for non-bioinformatic oriented researchers, scalability for large-scale and highly-complex samples and interpretation of recovered microbial communities. Graphical abstractAutometa: An automated taxon binning and genome binning workflow for single sample resolution of metagenomic communities.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.