DoBSeqWF: A framework for sensitive detection of individual genetic variation in pooled sequencing data
Cort, M.; Hagen, C. M.; Stoltze, U. K.; Hansen, T. v. O.; Nyegaard, M.; Hjalgrim, H.; Baekvad-Hansen, M.; Byrjalsen, A.; Schmiegelow, K.; Wadt, K.; Bybjerg-Grauholm, J.; Rasmusen, S.
Show abstract
MotivationPopulation screening for rare genetic diseases is limited by the high cost of next- generation sequencing. Double-batched sequencing (DoBSeq) is a cost-effective method for assigning rare variants to individuals using two-dimensional unique double- pooled sequencing. However, this method produces complex, high-depth sequencing data that requires a specialized workflow for efficient and reproducible analysis. ResultsWe developed DoBSeqWF (DoBSeq Workflow), a Nextflow-based pipeline for processing the pooled sequencing data from alignment through variant calling, filtering, and ultimately individual assignment of rare variants. Using separate training and validation datasets with whole genome sequencing as the gold standard, we benchmarked multiple variant callers, and we developed and implemented machine learning filters that improve rare variant calling performance while maintaining high sensitivity. The pipeline enables reproducible analysis and can be easily updated as bioinformatic tools and variant interpretations evolve. Availability and ImplementationDoBSeqWF is freely available at https://github.com/RasmussenLab/DoBSeqWF. Contact: srasmuss@sund.ku.dk
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 96%
- Concerning the eXclusion in human genomics: The choice of sex chromosome representation in the human genome drastically affects number of identified variants 95%
- GenoTools: An Open-Source Python Package for Efficient Genotype Data Quality Control and Analysis 94%
Similar papers in this journal
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 97%
- Flexible, Production-Scale, Human Whole Genome Sequencing On A Benchtop Sequencer 94%
- Structural variation of the malaria-associated human glycophorin A-B-E region 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.