Shine: To explore specific, sensitive and conserved biomarkers from massive microbial genomic data within intrapopulations.
Ji, C.; Shao, J. J.
Show abstract
BackgroundConcentrations of the pathogenic microorganisms DNA in biological samples are typically low. Therefore, DNA diagnostics of common infections are costly, rarely accurate, and challenging. Limited by failing to cover updated epidemic testing samples, computational services are difficult to implement in clinical applications without complex customized settings. Furthermore, the combined biomarkers used to maintain high conservation may not be cost effective and could cause several experimental errors in many clinical settings. Given the limitations of recent developed technology, 16S rRNA is too conserved to distinguish closely related species, and mosaic plasmids are not effective as well because of their uneven distribution across prokaryotic taxa. ResultsHere, we provide a computational strategy, Shine, that allows extraction of specific, sensitive and well-conserved biomarkers from massive microbial genomic datasets. Distinguished with simple concatenations with blast-based filtering, our method involves a de novo genome alignment-based pipeline to explore the original and specific repetitive biomarkers in the defined population. It can cover all members to detect newly discovered multicopy conserved species-specific or even subspecies-specific target probes and primer sets. The method has been successfully applied to a number of clinical projects and has the overwhelming advantages of automated detection of all pathogenic microorganisms without the limitations of genome annotation and incompletely assembled motifs. Using on our pipeline, users may select different configuration parameters depending on the purpose of the project for routine clinical detection practices on the website https://bioinfo.liferiver.com.cn with easy registration. ConclusionsThe proposed strategy is suitable for identifying shared phylogenetic markers while featuring low rates of false positive or false negative. This technology is suitable for the automatic design of minimal and efficient PCR primers and other types of detection probes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CAIM: Coverage-based Analysis for Identification of Microbiome 96%
- An ultra-sensitive bacterial pathogen and antimicrobial resistance diagnosis workflow using Oxford Nanopore adaptive sampling sequencing method 95%
- MTD: a unique pipeline for host and meta-transcriptome joint and integrative analyses of RNA-seq data 95%
Similar papers in this journal
- BacAnt: A Combination Annotation Server for Bacterial DNA Sequences to Identify Antibiotic Resistance Genes, Integrons, and Transposable Elements. 95%
- Metagenomic association analysis of gut symbiont Lactobacillus reuteri without host-specific genome isolation 95%
- Comparative genomics of three Colletotrichum scovillei strains and genetic analysis revealed genes involved in fungal growth and virulence on chili pepper 94%
Similar papers in this journal
- Phylogeny analysis of whole protein-coding genes in metagenomic data detected an environmental gradient for the microbiota 96%
- Gene co-expression network analysis of the human gut commensal bacterium Faecalibacterium prausnitzii based on WGCNA in R-Shiny 95%
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 94%
Similar papers in this journal
- Amplicon sequencing of single-copy protein-coding genes reveals accurate diversity for sequence-discrete microbiome populations 94%
- Molecular basis and evolutionary cost of a novel phenotype of macrolides/lincosamides resistance in Staphylococcus haemolyticus 94%
- RT-LAMP-CRISPR-Cas13a technology as a promising diagnostic tool for the SARS-CoV-2 virus 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.