Punchline: Identifying and comparing significant Pfam protein domain differences across draft whole genome sequences
Crossman, L. C.
Show abstract
MotivationShort-read draft paired-end Illumina assemblies can be fragmented, contain many contigs and be impacted on by repeat regions, caused by mobile element activity within the genome or inherently repetitive gene structure. Annotating such assemblies for function and analysing gene content can be challenging if predicted genes are fragmented across contigs. Such a case can often occur within specific families of genes such as longer genes with repeating domains, genes specifying several transmembrane domains and of unusual nucleotide content. These genes can often be virulence determinants, therefore losing these specific types of data can seriously impact downstream studies.\n\nResultsRather than studying the predicted gene content of draft genomes, we examined predicted protein content using the Pfam domain complements of predicted proteins. We produced a workflow, Punchline, to study the genetic content of draft contig assemblies by looking at the complement of short domains that are unlikely to be affected. We investigated a dataset of Bacteroides ovatus in terms of a grouping involving the vertebrate host from which the organism was isolated and identified potential host restricted functions and host restricted phylogenetic clustering.\n\nAvailabilityhttps://github.com/LCrossman\n\nContact: seq@sequenceanalysis.co.uk
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Phylum barrier and Escherichia coli intra-species phylogeny drive the acquisition of resistome in E. coli 93%
- Evidence of a novel sublineage of Streptococcus agalactiae in elephants from zoo populations in Germany 93%
- Genomic reconstruction of Bacillus anthracis from complex environmental samples enables high throughput identification and lineage assignment in Pakistan 93%
Similar papers in this journal
- A large-scale genome-based survey of acidophilic Bacteria suggests that genome streamlining is an adaption for life at low pH 93%
- No Assembly Required: Using BTyper3 to Assess the Congruency of a Proposed Taxonomic Framework for the Bacillus cereus group with Historical Typing Methods 93%
- FeGenie: a comprehensive tool for the identification of iron genes and iron gene neighborhoods in genomes and metagenome assemblies 92%
Similar papers in this journal
- No one tool to rule them all: Prokaryotic gene prediction tool performance is highly dependent on the organism of study 95%
- panRGP: a pangenome-based method to predict genomic islands and explore their diversity 94%
- DeNoFo: a file format and toolkit for standardised, comparable de novo gene annotation 93%
Similar papers in this journal
- Deep phylo-taxono genomics reveals Xylella as a variant lineage of plant associated Xanthomonas with Stenotrophomonas and Pseudoxanthomonas as misclassified relatives 94%
- Good host - bad host: molecular and evolutionary basis for survival, its failure, and virulence factors of the zoonotic nematode Anisakis pegreffii 91%
- Genomic architecture of three newly isolated unclassified Butyrivibrio species elucidate their potential role in the rumen ecosystem 90%
Similar papers in this journal
- The global proteome of Streptococcus pneumoniae EF3030 under nutrient-defined in vitro conditions 90%
- Major antigenic differences in Aeromonas salmonicida isolates correlate with the emergence of a new strain causing furunculosis in Chilean salmon farms 90%
- Evidence for the Existence of a Bacterial Etiology for Alzheimers Disease and for a Temporal-Spatial Development of a Pathogenic Microbiome in the Brain 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.