Back

Complex Indel Detection: A Simulation-Based Framework and Parsing with FreeBayes

Loh, Y. H. E.; Lieber, M. R.; Hsieh, C.-L.; Manojlovic, Z.

2026-05-29 bioinformatics
10.64898/2026.05.26.727999 bioRxiv
Show abstract

In contrast to simple deletions and simple insertions, most complex indels involve both deletions and insertions, often with base changes within a few nucleotides of the indels left and right boundaries. These complex indels often arise from double-strand breaks (DSB), which in normal somatic cells are predominantly repaired by nonhomologous DNA end joining (NHEJ). Such complex indels pose a difficult analytical problem for existing indel callers because the observed VCF representation may be locally shifted, extended with matching flanking bases, or fragmented into several closely spaced calls. To evaluate complex indel representation, we tested six variant calling approaches: FreeBayes, HaplotypeCaller, Mutect2, Strelka2, DRAGEN Germline, and DRAGEN Somatic pipelines. Among the approaches evaluated, FreeBayes most consistently represented simulated complex indels as single nearby variant records. We then developed a parsing workflow that derives effective deleted and inserted sequences from FreeBayes VCF output and enriches for candidate complex indels. This approach supports analysis of naturally occurring DSB repair events in single human colon crypts.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 1%
17.9%
2
BMC Bioinformatics
457 papers in training set
Top 0.2%
17.9%
3
PLOS Computational Biology
1863 papers in training set
Top 2%
14.6%
50% of probability mass above
4
Genome Biology
637 papers in training set
Top 3%
4.2%
5
Nucleic Acids Research
1281 papers in training set
Top 5%
4.2%
6
BioData Mining
22 papers in training set
Top 0.1%
3.9%
7
Bioinformatics Advances
203 papers in training set
Top 2%
3.9%
8
Scientific Reports
3612 papers in training set
Top 30%
3.4%
9
BMC Genomics
406 papers in training set
Top 3%
2.6%
10
PLOS ONE
5266 papers in training set
Top 41%
2.6%
11
Nature Communications
5641 papers in training set
Top 41%
2.3%
12
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
1.8%
13
The American Journal of Human Genetics
234 papers in training set
Top 2%
1.7%
14
GigaScience
212 papers in training set
Top 3%
1.6%
15
Genome Research
468 papers in training set
Top 4%
1.3%
16
Communications Biology
993 papers in training set
Top 24%
1.1%
17
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
1.0%
18
Genetic Epidemiology
55 papers in training set
Top 0.6%
1.0%
19
Briefings in Bioinformatics
354 papers in training set
Top 7%
0.8%
20
PLOS Genetics
862 papers in training set
Top 13%
0.8%
21
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.7%
0.8%
22
Journal of Computational Biology
48 papers in training set
Top 1%
0.8%
23
GENETICS
483 papers in training set
Top 5%
0.8%
24
European Journal of Human Genetics
58 papers in training set
Top 1%
0.6%