Back

Synth4bench: a framework for generating synthetic genomics data for the evaluation of tumor-only somatic variant calling algorithms

Fragkouli, S.-C.; Pechlivanis, N.; Anastasiadou, A.; Karakatsoulis, G.; Orfanou, A.; Kollia, P.; Agathangelidis, A.; Psomopoulos, F. E.

2024-03-08 bioinformatics
10.1101/2024.03.07.582313 bioRxiv
Show abstract

MotivationSomatic variant calling is a key activity towards identifying genomic alterations; yet, the evaluation of the respective tools remains challenging due to the scarcity of high quality ground truth datasets. To overcome this limitation, we developed synth4bench, a synthetic data generation pipeline for robust benchmarking. Using a systematic process to create distinct synthetic datasets, we thoroughly evaluated five variant callers (Mutect2, FreeBayes, VarDict, VarScan2 and LoFreq). We compared tool outputs against our synthetic ground truth across key sequencing aspects (such as depth and read length) to assess their capacities and shed light on their underlying algorithmic principles. ResultsSynth4bench is an approach for evaluating tumor-only somatic variant callers that relies on a systematic definition of fully controlled ground-truth datasets. Our analysis revealed significant inconsistencies among the tool outputs and a strong dependence of caller performance on sequencing parameters. Indels remain the hardest-to-call variant type, driven by errors at low allele frequencies. Algorithmic choice is also critical; the most robust callers displayed the highest precision in allele frequency estimation, while the most sensitive caller was best for maximizing true positive recovery. Conversely, the least suitable caller exhibited systematic errors along with the poorest overall performance. These findings indicate that there isnt a one-solution-fit-all; sequencing optimization together with caller selection are necessary to maximize sensitivity and reliability. Furthermore, the pronounced inconsistencies suggest that current algorithms are not yet able to capture all mutational mechanisms adequately, with the modeling of the underlying processes remaining an open challenge. Availabilitycode: https://github.com/sfragkoul/synth4bench/ and data: https://zenodo.org/records/16524193 Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=69 SRC="FIGDIR/small/582313v2_ufig1.gif" ALT="Figure 1"> View larger version (16K): org.highwire.dtl.DTLVardef@1ac0fe5org.highwire.dtl.DTLVardef@1479cddorg.highwire.dtl.DTLVardef@8b7d1borg.highwire.dtl.DTLVardef@1c2a81b_HPS_FORMAT_FIGEXP M_FIG C_FIG

Published in Frontiers in Bioinformatics (predicted rank #18) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.