ContaTester: Fast cross-contamination estimation and identification for large human sequencing cohorts
Delafoy, D.; Mercier, J.; Larsonneur, E.; Wiart, N.; Sandron, F.; Mejean, T.; Meslage, S.; Daian, D.; Olaso, R.; Boland, A.; Deleuze, J.-F.; Meyer, V.
Show abstract
BackgroundInterest in genomic medicine for human health studies and clinical applications is rapidly increasing. Clinical applications require contamination-free samples to avoid misleading results and provide a sound basis for diagnosis. ResultsHere we present ContaTester, a tool which requires only allele balance information gathered from a VCF file to detect cross-contamination in germline human DNA samples. Based on a regression model of allele balance distribution, ContaTester allows fast checking of contamination levels for single samples or large cohorts (less than two minutes per sample). We demonstrate the efficiency of ContaTester using experimental validations: ContaTester shows similar results to methods requiring alignment data but with a significantly reduced storage footprint and less computation time. Additionally, for contamination levels above 5%, ContaTester can identify contaminants across a cohort, providing important clues for troubleshooting and quality assessment. ConclusionsContaTester estimates contamination levels from VCF files generated from whole genome sequencing normal sample and provides reliable contaminant identification for cohorts or experimental batches.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance analysis of conventional and AI-based variant callers using short and long reads 97%
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 95%
- Rare Copy Number Variant analysis in case-control studies using SNP Array Data: a scalable and automated data analysis pipeline 95%
Similar papers in this journal
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 95%
- Flexible, Production-Scale, Human Whole Genome Sequencing On A Benchtop Sequencer 95%
- Visualizing and exploring patterns of large mutational events with SigProfilerMatrixGenerator 95%
Similar papers in this journal
- Optical genome mapping as a next-generation cytogenomic tool for detection of structural and copy number variations for prenatal genomic analyses 94%
- DLO Hi-C Tool for Digestion-Ligation-Only Hi-C Chromosome Conformation Capture Data Analysis 93%
- The FORCE panel: An all-in-one SNP marker set for confirming investigative genetic genealogy leads and for general forensic applications 92%
Similar papers in this journal
- SnpHub: an easy-to-set-up web server framework for exploring large-scale genomic variation data in the post-genomic era with applications in wheat 95%
- A graph clustering algorithm for detection and genotyping of structural variants from long reads 95%
- cfDNA UniFlow: A unified preprocessing pipeline for cell-free DNA data from liquid biopsies 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.