Back

CuBi-MeAn Customized Pipeline for Metagenomic Data Analysis

Keshani-Langroodi, S.; Sales, C. M.

2021-09-11 bioinformatics
10.1101/2021.09.10.458355 bioRxiv
Show abstract

1.Whole genome shotgun sequencing is a powerful to study microbial community is a given environment. Metagenomic binning offers a genome centric approach to study microbiomes. There are several tools available to process metagenomic data from raw reads to the interpretation there is still lack of standard approach that can be used to process the metagenomic data step by step. In this study CuBi-MeAn (Customizable Binning and Metagenomic Analysis) create a customizable and flexible processing pipeline, to process the metagenomic data and generate results for further interpretation. This study aims to perform metagenomic binning to enhance taxonomical classification, functional potentials, and interactions among microbial populations in environmental systems. This customized pipeline which is comprised of a series of genomic/metagenomic tools designed to recover better quality results and reliable interpretation of the system dynamics for the given systems. For this reason, a metagenomic data processing pipeline is developed to evaluate metagenomic data from three environmental engineering projects. The use of our pipeline was demonstrated and compared on three different datasets that were of different sizes, from different sequencing platforms, and generated from three different environmental sources. By designing and developing a flexible and customized pipeline, this study has showed how to process large metagenomic data sets with limited resources. This result not only would help to uncover new information from environmental samples, but also, could be applicable to any other metagenomic studies across various disciplines.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.