CNV-Finder: Streamlining Copy Number Variation Discovery
Kuznetsov, N.; Daida, K.; Makarious, M. B.; Al-Mubarak, B.; Atterling Brolin, K.; Malik, L.; Kouam, C.; Baker, B.; Real, R.; Step, K.; Lange, L. M.; Wu, L.; Ostrozovicova, M.; Andersh, K. M.; Kung, P.-J.; Mecheri, Y.; Tay, Y.-W.; Soundous Malek, B.; Al Tassan, N.; Teresa Perinan, M.; Hong, S.; Koretsky, M. J.; Sargeant, L.; Levine, K.; Blauwendraat, C.; Billingsley, K. J.; Bandres-Ciga, S.; Leonard, H. L.; Bardien, S.; Morris, H. R.; Singleton, A. B.; Nalls, M. A.; Vitale, D.; The Global Parkinson's Genetics Program,
Show abstract
Copy Number Variations (CNVs) play pivotal roles in the etiology of complex diseases and are variable across diverse populations. Understanding the association between CNVs and disease susceptibility is significant in disease genetics research and often requires analysis of large sample sizes. One of the most cost-effective and scalable methods for detecting CNVs is based on normalized signal intensity values, such as Log R Ratio (LRR) and B Allele Frequency (BAF), from Illumina genotyping arrays. In this study, we present CNV-Finder, a novel pipeline integrating deep learning techniques on array data, specifically a Long Short-Term Memory (LSTM) network, to expedite the large-scale identification of CNVs within predefined genomic regions. This facilitates efficient prioritization of samples for time-consuming or costly subsequent analyses such as Multiplex Ligation-dependent Probe Amplification (MLPA), short-read, and long-read whole genome sequencing. We incorporate four genes to establish our methods--Parkin (PRKN), Leucine Rich Repeat And Ig Domain Containing 2 (LINGO2), Microtubule Associated Protein Tau (MAPT), and alpha-Synuclein (SNCA)--which may be relevant to neurological diseases such as Alzheimers disease (AD), Parkinsons disease (PD), Progressive Supranuclear Palsy (PSP), or related disorders such as essential tremor (ET). By training our models on expert-annotated samples and validating them across diverse cohorts, including those from the Global Parkinsons Genetics Program (GP2) and additional dementia-specific databases, we demonstrate the efficacy of CNV-Finder in accurately detecting deletions and duplications. Our pipeline outputs app-compatible files for visualization within CNV-Finders interactive web application. This interface enables researchers to review predictions and filter displayed samples by model prediction values, LRR range, and variant count in order to explore or confirm results. Our pipeline integrates this human feedback to enhance model performance and reduce false positive rates. Through a series of comprehensive analyses and validations using visual inspection, MLPA, short-read, and long-read sequencing data, we demonstrate the robustness and adaptability of CNV-Finder in identifying CNVs with regions of varied size, probe density, and noise. Our findings highlight the significance of contextual understanding and human expertise in enhancing the precision of CNV identification, particularly in complex genomic regions like 17q21.31. The CNV-Finder pipeline is a scalable, publicly available resource for the scientific community, available on GitHub (https://github.com/GP2code/CNV-Finder; DOI 10.5281/zenodo.14182563). CNV-Finder not only expedites accurate candidate identification but also significantly reduces the manual workload for researchers, enabling future targeted validation and downstream analyses in regions or phenotypes of interest.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 95%
- Genome-Wide Sequencing as a First-Tier Screening Test for Short Tandem Repeat Expansions 95%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 94%
Similar papers in this journal
- GeneTerpret: a customizable multilayer approach to genomic variant prioritization and interpretation 94%
- Identification of single nucleotide variants using position-specific error estimation in deep sequencing data 93%
- Identification of allele-specific KIV-2 repeats and impact on Lp(a) measurements for cardiovascular disease risk 93%
Similar papers in this journal
- Exome copy number variant detection, analysis and classification in a large cohort of families with undiagnosed rare genetic disease 95%
- HiFi long-read genomes for difficult-to-detect clinically relevant variants 95%
- The impact of 22q11.2 copy number variants on human traits in the general population 94%
Similar papers in this journal
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 93%
- MitoDelta: identifying mitochondrial DNA deletions at cell-type resolution from single-cell RNA sequencing data 93%
- Precise Exome Analysis Of Blastocyst Biopsy Scale Samples Using Primary Template-Directed Amplification 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.