Back

DoBSeqWF: A framework for sensitive detection of individual genetic variation in pooled sequencing data

Cort, M.; Hagen, C. M.; Stoltze, U. K.; Hansen, T. v. O.; Nyegaard, M.; Hjalgrim, H.; Baekvad-Hansen, M.; Byrjalsen, A.; Schmiegelow, K.; Wadt, K.; Bybjerg-Grauholm, J.; Rasmusen, S.

2025-04-25 genetic and genomic medicine
10.1101/2025.04.23.25326275 medRxiv
Show abstract

MotivationPopulation screening for rare genetic diseases is limited by the high cost of next- generation sequencing. Double-batched sequencing (DoBSeq) is a cost-effective method for assigning rare variants to individuals using two-dimensional unique double- pooled sequencing. However, this method produces complex, high-depth sequencing data that requires a specialized workflow for efficient and reproducible analysis. ResultsWe developed DoBSeqWF (DoBSeq Workflow), a Nextflow-based pipeline for processing the pooled sequencing data from alignment through variant calling, filtering, and ultimately individual assignment of rare variants. Using separate training and validation datasets with whole genome sequencing as the gold standard, we benchmarked multiple variant callers, and we developed and implemented machine learning filters that improve rare variant calling performance while maintaining high sensitivity. The pipeline enables reproducible analysis and can be easily updated as bioinformatic tools and variant interpretations evolve. Availability and ImplementationDoBSeqWF is freely available at https://github.com/RasmussenLab/DoBSeqWF. Contact: srasmuss@sund.ku.dk

Published in NAR Genomics and Bioinformatics (predicted rank #6) · training set

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.