Back

Long-read cross-platform validation reveals novel repeat features in myotonic dystrophy type 2

Carlomagno, M.; Suarez Lopez, F. J.; Maestri, S.; Esposito, A.; Obadovic, V.; Visconti, V. V.; Ciabini, D.; Marcolungo, L.; Rossi, N.; Casagrande, M.; Angheben, L.; Spadoni, L.; D Apice, M. R.; Novelli, G.; Delledonne, M.; Botta, A.; Rossato, M.

2026-06-10 genomics
10.64898/2026.06.06.730578 bioRxiv
Show abstract

The broader application of long-read sequencing (LRS) for repeat expansion characterization in myotonic dystrophy type 2 (DM2) and other repeat expansion disorders (REDs) remains limited by the lack of systematic validation and benchmarking of sequencing results and bioinformatic workflows. Here, we performed an orthogonal cross-platform validation of previously generated Oxford Nanopore Technologies (ONT) data by sequencing the same DNA samples with Pacific Biosciences (PacBio) HiFi following amplification-free targeted enrichment in a cohort of 8 DM2 patients. Despite substantial differences in sequencing chemistry and coverage, the two platforms showed high concordance in repeat size estimation, somatic mosaicism, and repeat architecture. This validation confirmed the presence of the (TCTG)n motif and enabled the identification of a previously unreported (CCCG)n motif at the 3' end of expanded alleles, further highlighting the structural complexity of the CNBP expansion. Through this analysis, we also established a bioinformatic workflow that improved ONT-based repeat characterization, addressing limitations in motif resolution and enabling more accurate analysis of CNBP expansions. Overall, this study provides a validated framework for LRS-based CNBP repeat analysis, supporting the integration of these technologies into routine molecular investigation for DM2 and other REDs.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.