Back

LEMMIv2: Benchmarking Framework for Metagenomic and 16S Amplicon Profilers with a Catalogue of Evaluated Tools.

Seppey, M.; Benavides, A.; Berkeley, M.; Manni, M.; Zdobnov, E. M.

2025-03-11 genomics
10.1101/2025.03.06.641904 bioRxiv
Show abstract

Sequencing has transformed microbial studies, enabling metagenomic analysis of microbial communities without the need for culturing or prior knowledge of sample composition. The essential analysis of the primary sequencing reads, however, is complex and led to the development of diverse computational strategies. The plethora of available methods and their numerous parameters poses a practical challenge for practitioners and creates a visibility barrier for developers of novel approaches. In addition to technical limitations related to user computing environment, algorithmic solutions, and their scalability, there are critical considerations regarding the reference database, i.e. the knowledge against which the data is interpreted, as well as the target of the analysis, expected sample composition, and peculiarities of read data from different sequencing platforms. To facilitate informed decision-making, we introduced the LEMMI platform for continuous benchmarking of software tools for metagenomic analyses, where developers can receive impartial benchmarks for method publication and users benefit from a standardized and benchmarked catalogue of tools. Here we present developments of LEMMI version 2, including assessments of different target scenarios, long- and short-read sequencing data, alternative taxonomies, and a standalone pipeline (https://lemmi.ezlab.org). In addition to LEMMI, which focuses on shotgun metagenomic profiling, we extended this approach to bacteria profiling with 16S amplicon sequencing (https://lemmi16S.ezlab.org).

Published in Genome Biology (predicted rank #4) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.