A comprehensive workflow for target adaptive sampling long-read sequencing applied to hereditary cancer patient genomes
Nakamura, W.; Hirata, M.; Oda, S.; Chiba, K.; Okada, A.; Mateos, R. N.; Sugawa, M.; Iida, N.; Ushiama, M.; Tanabe, N.; Sakamoto, H.; Kawai, Y.; Tokunaga, K.; NCBN Controls WGS Consortium, ; Tsujimoto, S.; Shiba, N.; Ito, S.; Yoshida, T.; Shiraishi, Y.
Show abstract
Innovations in sequencing technology have led to the discovery of novel mutations that cause inherited diseases. However, many patients with suspected genetic diseases remain undiagnosed. Long-read sequencing technologies are expected to significantly improve the diagnostic rate by overcoming the limitations of short-read sequencing. In addition, Oxford Nanopore Technologies (ONT) offers a computationally-driven target enrichment technology, adaptive sampling, which enables intensive analysis of targeted gene regions at low cost. In this study, we developed an efficient computational workflow for target adaptive sampling long-read sequencing (TAS-LRS) and evaluated it through application to 33 genomes collected from suspected hereditary cancer patients. Our workflow can identify single nucleotide variants with nearly the same accuracy as the short-read platform and elucidate complex forms of structural variations. We also newly identified SVAs affecting the APC gene in two patients with familial adenomatous polyposis, as well as their sites of origin. In addition, we demonstrated that off-target reads from adaptive sampling, which are typically discarded, can be effectively used to accurately genotype common SNPs across the entire genome, enabling the calculation of a polygenic risk score. Furthermore, we identified allele-specific MLH1 promoter hypermethylation in a Lynch syndrome patient. In summary, our workflow with TAS-LRS can simultaneously capture monogenic risk variants including complex structural variations, polygenic background as well as epigenetic alterations, and will be an efficient platform for genetic disease research and diagnosis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 95%
- Multi-modal investigation of the schizophrenia-associated 3q29 genomic interval reveals global genetic diversity with unique haplotypes and segments that increase the risk for non-allelic homologous recombination 95%
- A multilayered post-GWAS assessment on genetic susceptibility to pancreatic cancer 94%
Similar papers in this journal
- Large scale genome-wide association study in a Japanese population identified 45 novel susceptibility loci for 22 diseases 95%
- Long read sequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other traits 95%
- Saturation Genome Editing Resolves the Functional Spectrum of Pathogenic VHL Alleles 95%
Similar papers in this journal
- Utilizing Non-Invasive Prenatal Test Sequencing Data Resource for Human Genetic Investigation 95%
- Joint estimation and imputation of variant functional effects using high throughput assay data 95%
- Long-read sequencing of diagnosis and post-therapy medulloblastoma reveals complex rearrangement patterns and epigenetic signatures 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.