Back

ContaTester: Fast cross-contamination estimation and identification for large human sequencing cohorts

Delafoy, D.; Mercier, J.; Larsonneur, E.; Wiart, N.; Sandron, F.; Mejean, T.; Meslage, S.; Daian, D.; Olaso, R.; Boland, A.; Deleuze, J.-F.; Meyer, V.

2021-10-02 bioinformatics
10.1101/2021.10.01.461647 bioRxiv
Show abstract

BackgroundInterest in genomic medicine for human health studies and clinical applications is rapidly increasing. Clinical applications require contamination-free samples to avoid misleading results and provide a sound basis for diagnosis. ResultsHere we present ContaTester, a tool which requires only allele balance information gathered from a VCF file to detect cross-contamination in germline human DNA samples. Based on a regression model of allele balance distribution, ContaTester allows fast checking of contamination levels for single samples or large cohorts (less than two minutes per sample). We demonstrate the efficiency of ContaTester using experimental validations: ContaTester shows similar results to methods requiring alignment data but with a significantly reduced storage footprint and less computation time. Additionally, for contamination levels above 5%, ContaTester can identify contaminants across a cohort, providing important clues for troubleshooting and quality assessment. ConclusionsContaTester estimates contamination levels from VCF files generated from whole genome sequencing normal sample and provides reliable contaminant identification for cohorts or experimental batches.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.