Back

Image-based DNA Sequencing Encoding for Detecting Low-Mosaicism Somatic Mobile Element Insertions

Tan, M.; Lin, Z.; Chen, Z.; Zhou, H.; Park, J.; He, Z.; Lee, E. A.; Gao, Z.; Zhu, X.

2025-05-16 genomics
10.1101/2024.11.07.619809 bioRxiv
Show abstract

Active LINE-1 (L1), Alu, and SVA mobile elements in the human genome are capable of retrotransposition, resulting in novel mobile element insertions (MEIs) in both germline and somatic tissues. Detecting MEIs through DNA sequencing relies on supporting reads overlapping MEI junctions; however, artifacts from DNA amplification, sequencing, and alignment errors produce numerous false positives. Systematic detection of somatic MEIs, particularly those with low mosaicism, remains a significant challenge. Previous methods had required a high number of supporting reads which limits the detection sensitivity, or human inspections that are susceptible to biases. Here, we developed RetroNet, an algorithm that encodes MEI-supporting sequencing reads into images, and employs a deep neural network to identify somatic MEIs with as few as two reads. Trained on extensive and diverse datasets and benchmarked across various conditions, RetroNet surpasses previous methods and eliminates the need for extensive manual examinations. The RetroNet analysis on the Illumina sequencing of 161x or 195x of a cancer cell line achieved an average precision of 0.885 and recall of 0.579 for detecting somatic L1 insertions that are present in as few as 1.79% of the cells. Additionally, we demonstrated that RetroNet is effective for analyzing highly degraded DNA, such as circulating tumor DNA. RetroNet is applicable to the rapidly generated short-read sequencing data and has the potential to provide further insights into the functional and pathological implications of somatic retrotranspositions.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.