Back

Recovery of high-qualitied Genomes from a deep-inland Salt Lake Using BASALT

Yu, K.; Qiu, Z.; Mu, R.; Qiao, X.; Zhang, L.; Lian, C.-A.; Deng, C.; Wu, Y.; Xu, Z.; Li, B.; Pan, B.; Zhang, Y.; Fan, L.; Liu, Y.; Cao, H.; Jin, T.; Chen, B.; Wang, F.; Yan, Y.; Xie, L.; Zhou, L.; Yi, S.; Chi, S.; Zhang, T.; Zhuang, W.

2021-03-05 bioinformatics
10.1101/2021.03.05.434042 bioRxiv
Show abstract

Metagenomic binning enables the in-depth characterization of microorganisms. To improve the resolution and efficiency of metagenomic binning, BASALT (Binning Across a Series of AssembLies Toolkit), a novel binning toolkit was present in this study, which recovers, compares and optimizes metagenomic assembled genomes (MAGs) across a series of assemblies from short-read, long-read or hybrid strategies. BASALT incorporates self-designed algorithms which automates the separation of redundant bins, elongate and refine best bins and improve contiguity. Evaluation using mock communities revealed that BASALT auto-binning obtained up to 51% more number of MAGs with up to 10 times better MAG quality from microbial community at low (132 genomes) and medium (596 genomes) complexity, compared to other binners such as DASTool, VAMB and metaWRAP. Using BASALT, a case-study analysis of a Salt Lake sediment microbial community from northwest arid region of China was performed, resulting in 426 non-redundant MAGs, including 352 and 69 bacterial and archaeal MAGs which could not be assigned to any known species from GTDB (ANI < 95%), respectively. In addition, two Lokiarchaeotal MAGs that belong to superphylum Asgardarchaeota were observed from Salt Lake sediment samples. This is the first time that candidate species from phylum Lokiarchaeota was found in the arid and deep-inland environment, filling the current knowledge gap of earth microbiome. Overall, BASALT is proven to be a robust toolkit for metagenomic binning, and more importantly, expand the Tree of Life.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.