Back

PopGenPlayground: a population genomics analysis pipeline

Wolfsberger, W. W.; Shchubelka, K.; Oleksyk, O. T.; Hasynets, Y.; Patskun, S.; Vakerych, M.; Kish, R.; Mirutenko, V.; Mirutenko, V.; Cotoraci, C. A.; Pop, C.; Neagu, O.; Balta, C.; Herman, H.; Mare, P.; Dumitra, S.; Papiu, H.; Hermenean, A.; Oleksyk, T. K.

2024-03-02 bioinformatics
10.1101/2024.02.27.582400 bioRxiv
Show abstract

BackgroundPopulation genomic projects are essential in the current drive to map the genome diversity of human populations across the globe. Various barriers persist hindering these efforts, and the lack of bioinformatic expertise and reproducible standardized population-scale analysis is one of the major challenges limiting their discovery potential. Scalable, automated, user-friendly pipelines can help researchers with minimum programming skills to tackle these issues without extensive training. ResultsPopGenPlayground (PGP), is a streamlined, single-command computation pipeline designed for human population genomics analysis based on Snakemake workflow management system. Developed to automate secondary analysis of a previously published national genome project, it leverages the publicly available genomic databases for comparative analysis and annotation of variant calls. ConclusionsPGP presents a multi-platform robust population analysis pipeline, that reduces the time and the expertise levels to perform the main core of population analysis for a national genome project. PGP provides a comprehensive secondary analysis tool and can be used to perform analysis on a personal computer or using a remote high-performance computing platform.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.