A phylogeny-guided framework for decoding mechanisms of human endogenous retrovirus regulation in health and disease
Patterson, A.; Duong, B.; Yoon, L.; Foster, M.; MacMullen, L.; Wickramasinghe, J.; Lucas, A.; Srivastava, A.; Jacobson, S.; Murphy, M. E.; Soldan, S.; Lieberman, P. M.; Auslander, N.
Show abstract
Human endogenous retroviruses (HERVs) are remmants of ancient infections which make up to [~]8% of the human genome. Their activity influences development, immunity, and cancer, but studying them has been limited by a key technical challenge: short-read sequencing cannot uniquely assign reads to these highly repetitive elements. Here, we present ERVmancer, a phylogeny-informed method that resolves the read-mapping ambiguity and quantifies HERV expression across scales, from individual loci to entire retroviral clades, depending on mapping confidence. Benchmarking with sample-matched long- and short-read data generated in this study demonsrates that ERVmancer outperforms existing approaches in both sensitivity and specificity. Application of ERVmancer recapitulates known HERV expression patterns in multiple sclerosis and uncovers new biology in breast cancer, including suppression of HERVH-LTR7 by p53. By enabling accurate and scalable quantification of integrated retroviral elements, ERVmancer provides a broadly applicable resource for investigating retroviral mechanisms in health and disease.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DEMINERS enables clinical metagenomics and comparative transcriptomic analysis by increasing throughput and accuracy of nanopore direct RNA sequencing 96%
- Enhancer regulatory networks globally connect non-coding breast cancer loci to cancer genes 95%
- RATTLE: Reference-free reconstruction and quantification of transcriptomes from Nanopore sequencing 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.