Back

A Meta-analysis of the known Global Distribution and Host Range of the Ralstonia Species Complex

Lowe-Power, T.; Chipman, K.

2020-07-13 microbiology
10.1101/2020.07.13.189936 bioRxiv
Show abstract

The Ralstonia species complex is a group of genetically diverse plant wilt pathogens. Our goal is to create a database that contains the reported global distribution and host range of Ralstonia clades (e.g. phylotypes and sequevars). In this fifth release, we have cataloged information from 304 sources that report one or more Ralstonia strains isolated from 107 geographic regions. Metadata for nearly 10,000 strains are available as a supplemental table. The aggregated data suggest that the pandemic brown rot lineage (IIB-1) is the most widely dispersed lineage, and the phylotype I and IIB-4 lineages have the broadest natural host range. Although phylotype III is largely restricted to Africa, one strain collection reports a phylotype III strain isolated from Jamaica in the mid-1900s. In the previous release, we included reported presence of phylotype III strains in Brazil, but closer inspection of those results reveals that the strains were actually phylotype I strains that were mis-identified. Similarly, although phylotype IV is mostly found in East and Southeast Asia, phylotype IV strains are reported to be present in Kenya. Additionally, we have created an open science resource for phylogenomics of the RSSC. We associated strain metadata (host of isolation, location of isolation, and clade) with almost 700 genomes in a public KBase narrative. Our colleagues can use this narrative to identify the phylogenetic position of newly sequenced strains. We further curate a set of 601 high quality genomes based on low contamination and high completeness by CheckM. Our colleagues can use the curated dataset for comparative genomics studies.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.