Back

SwarmGenomics: A Unified Pipeline for Individual-Based Whole-Genome Analyses

Kylmaenen, A.; Chen, Y.-C.; Javaheri Tehrani, S.; Vellnow, N.; Wilcox, J. J.; Gossmann, T. I.

2025-08-15 genomics
10.1101/2025.08.13.670070 bioRxiv
Show abstract

Advances in sequencing technologies have made whole-genome data widely accessible, enabling research in population genetics, evolutionary biology, and conservation. However, analyzing whole-genome sequencing (WGS) data remains challenging, often requiring multiple specialized tools and substantial bioinformatics expertise. We present SwarmGenomics, a modular, user-friendly command-line pipeline for reference-based genome assembly and individual-based genetic analyses. The pipeline integrates seven modules: heterozygosity estimation, runs of homozygosity detection, Pairwise Sequentially Markovian Coalescent (PSMC) analysis, unmapped reads classification, repeat analysis, mitochondrial genome assembly, and nuclear mitochondrial DNA segment (NUMT) identification. Each module can be run independently or as part of a complete workflow. We demonstrate the pipelines utility with a case study on the giant panda (Ailuropoda melanoleuca), revealing insights into genetic diversity, inbreeding history, historical population size changes, transposable element activity, and microbial contamination. SwarmGenomics lowers the entry barrier for genomic analysis of diploid, non-model species, serving both as a research and teaching tool. The pipeline and documentation are available at https://github.com/AureKylmanen/Swarmgenomics.

Published in Molecular Ecology Resources (predicted rank #1) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.