CNV-Profile Regression: A New Approach for Copy Number Variant Association Analysis in Whole Genome Sequencing Data
Si, Y.; Lu, W.; Holloway, S. T.; Wang, H.; Tucci, A. A.; Brucker, A.; Cheng, Y.; Wang, L.-S.; Schellenberg, G. D.; Lee, W.-P.; Tzeng, J.-Y.
Show abstract
Copy number variants (CNVs) are DNA gains or losses involving >50 base pairs. Assessing CNV effects on disease risk requires consideration of several factors. First, there are no natural definitions for CNV loci. Second, CNV effects can depend on dosage and length. Third, CNV effects can be more accurately estimated when all CNV events in a genomic region are analyzed together to assess their joint effects. We propose a new framework for association analysis that directly models an individuals entire CNV profile within a genomic region. This framework represents an individuals CNVs using a CNV profile curve to capture variations in CNV length and dosage and to bypass the need to predefine CNV loci. CNV effects are estimated at each genome position, making the results comparable across different studies. To jointly estimate the effects of all CNVs, we use a Lasso penalty to select CNVs associated with the trait and integrate a weighted L2-fusion penalty to encourage similar effects of adjacent CNVs when supported by the data. Simulations show that the proposed model can more effectively identify causal CNVs while maintaining false positive rates comparable to baseline methods and yield more precise effect-size estimates across different settings. When applied to CNV derived from whole genome sequencing data of the Alzheimers Disease Sequencing Project, the proposed methods identify additional CNVs associated with Alzheimers Disease (AD). These identified CNVs overlap with several known AD-risk genes and are significantly enriched by biological processes related to neuron structures and functions crucial in AD development.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bayesian transcriptome-wide association study method leveraging both cis- and trans- eQTL information through summary statistics 96%
- UK-Biobank Whole Exome Sequence Binary Phenome Analysis with Robust Region-based Rare Variant Test 95%
- Significance tests for R2 of out-of-sample prediction using polygenic scores 95%
Similar papers in this journal
- An exact, unifying framework for region-based association testing in family-based designs, including higher criticism approaches, SKATs, multivariate and burden tests 95%
- Incorporating family disease history and controlling case-control imbalance for population based genetic association studies 94%
- CoMM-S2: a collaborative mixed model using summary statistics in transcriptome-wide association studies 94%
Similar papers in this journal
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 95%
- BayesKAT: Bayesian Optimal Kernel-based Test for genetic association studies reveals joint genetic effects in complex diseases 95%
- WEVar: a novel statistical learning framework for predicting noncoding regulatory variants 94%
Similar papers in this journal
Similar papers in this journal
- CNValidatron: Accurate And Efficient Validation of PennCNV Calls Using Computer Vision 95%
- Identification and Utilization of Copy Number Information for Correcting Hi-C Contact Map of Cancer Cell Line 93%
- DMRscaler: A Scale-Aware Method to Identify Regions of Differential DNA Methylation Spanning Basepair to Multi-Megabase Features 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.