Back

Triplex DNA and inverted repeats cause long-read sequencing bias against simple satellite DNA

Carvalho, A. B.; Kim, B. Y.; Uno, F.

2026-07-20 genomics
10.64898/2026.07.13.738322 bioRxiv
Show abstract

We recently showed that Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) have very strong sequencing bias against simple satellites, probably caused by single-stranded DNA folding into non-canonical (non-B) structures during sequencing. Here we extend these observations by computational and experimental approaches in the Drosophila and human genomes. We found that (i) only a small subset of simple satellites cause sequencing bias; many satellites (e.g., (ACTGGG)n) are benign and easily sequenced. (ii) The biases most likely are caused by two distinct non-B DNA structures: triplex DNA formed by some, but not all, AG-rich satellites (only those predicted to form strong mirror repeats), and hairpins formed by some, but not all, AT-rich satellites (only those predicted to form fairly strong inverted repeats). (iii) The correlation between the predicted stability of these non-B structures, and the strength of sequencing bias indicates that non-B DNA is indeed the culprit. (iv) The likely source of these non-B structures is single-stranded DNA formed during ONT and PacBio sequencing, and hence its removal might solve the bias. We tested this by adding single-strand binding protein to ONT sequencing, and found that it irreversibly kills the flow cells. (v) A recent sequencing effort in Drosophila melanogaster using very high depth Ultra-Long ONT sequencing (967x) still failed to assemble many genes located near satellite blocks. Brute-force will not solve the problem; instead, further investment is needed by the sequencing companies to achieve truly unbiased sequencing.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.