FetchM: Streamlining Genome and Metadata Integration for Microbial Comparative Genomics
Anik, T. A.
Show abstract
FetchM is a Python-based tool for fetching, analyzing, and combining bacterial genomic metadata from the NCBI Genome database and associated sample metadata from NCBI BioSample records. When working with bulk-genome analyses, such as comparative genomics or pangenome studies, you often require a unified dataset that captures the full context of a particular bacterial species population. You can obtain genomic metadata by downloading the ncbi_dataset.tsv file from the NCBI Genome database for a specific bacterial species. However, this file lacks key metadata fields such as Collection Date, Host, Geographic Location, and Isolation Source. FetchM fills this gap by automatically retrieving these missing fields, linking genome accessions to their corresponding BioSample records via the NCBI Entrez API. FetchM not only helps you compile a complete, metadata-rich dataset but also provides visualizations and summaries of both genomic and contextual metadata features. You can filter and download sequences based on specific criteria such as year, host, isolation source, country, continent, and subcontinent, making it a flexible and powerful companion for large-scale genomic studies. FetchM is available as an open-source tool at: https://github.com/Tasnimul-Arabi-Anik/FetchM. It can also be downloaded as a PyPI package.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Density-based binning of gene clusters to infer function or evolutionary history using GeneGrouper 94%
- StrainHub: A phylogenetic tool to construct pathogen transmission networks 93%
- No one tool to rule them all: Prokaryotic gene prediction tool performance is highly dependent on the organism of study 92%
Similar papers in this journal
- Coinfinder: Detecting Significant Associations and Dissociations in Pangenomes 94%
- Bakta: Rapid & standardized annotation of bacterial genomes via alignment-free sequence identification 94%
- FANGORN: A quality-checked and publicly available database of full-length 16S-ITS-23S rRNA operon sequences 94%
Similar papers in this journal
Similar papers in this journal
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 95%
- Extraction of near-complete genomes from metagenomic samples: a new service in PATRIC 94%
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.