Back

MEDUSA: Maintaining Entire DNA Duplexes for Utmost Sequencing Accuracy

Xiong, K.; Li, R.; Li, N.; Liu, R.; Narayan, A.; Rhoades, J.; Song, C.; Lindsley, R. C.; Parsons, H. A.; Makrigiorgos, G. M.; Adalsteinsson, V. A.

2025-11-08 genomics
10.1101/2025.11.07.687202 bioRxiv
Show abstract

Highly accurate DNA sequencing is essential for many applications. Duplex sequencing-based methods achieve unparalleled accuracy by requiring reads from both strands of the original DNA duplex to match. Yet, methods to prepare dsDNA for sequencing may resynthesize portions of each DNA duplex and cause base damage errors on one strand to become indistinguishable from true mutations on both strands. Here, we report MEDUSA (Maintaining Entire DNA Duplexes for Utmost Sequencing Accuracy) which minimizes dsDNA resynthesis to maximize duplex sequencing accuracy and yield. MEDUSA carefully repairs and blunts fragmented dsDNA, then employs apyrase to digest residual dNTPs, followed by restricted dA-tailing to prevent resynthesis. MEDUSA affords full genome coverage in a simplified protocol that is broadly compatible with dsDNA fragmentation and library preparation kits. We benchmarked MEDUSA on sheared genomic DNA and cell-free DNA and found a residual SNV frequency within a median 1.23-fold (range 0.92-1.88; p < 0.001) of what was expected if resynthesis was mostly blocked, but with full genome coverage and duplex yields within a median 1.02-fold (range 0.33 - 1.46; p = 0.258) of traditional methods that do not limit resynthesis. In all, MEDUSA could enable high breadth or depth of duplex sequencing while limiting false discovery. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/687202v1_ufig1.gif" ALT="Figure 1"> View larger version (12K): org.highwire.dtl.DTLVardef@1ddbc77org.highwire.dtl.DTLVardef@802ca9org.highwire.dtl.DTLVardef@f443ecorg.highwire.dtl.DTLVardef@974bdb_HPS_FORMAT_FIGEXP M_FIG C_FIG

Published in Clinical Chemistry (predicted rank #23) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.