Detecting CYP2C19 deletions from genotyping array signals using neural networks
Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.
Show abstract
Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Boosting variant-calling performance with multi-platform sequencing data using Clair3-MP 93%
- Rare Copy Number Variant analysis in case-control studies using SNP Array Data: a scalable and automated data analysis pipeline 92%
- MADloy: Robust detection of mosaic loss of chromosome Y from genotype-array-intensity data 92%
Similar papers in this journal
- labelSeg: segment annotation for tumor copy number alteration profiles 93%
- WEVar: a novel statistical learning framework for predicting noncoding regulatory variants 92%
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 92%
Similar papers in this journal
- Association Tests Using Copy Number Profile Curves (CONCUR) Enhances Power in Rare Copy Number Variant Analysis 94%
- Explainable deep transfer learning model for disease risk prediction using high-dimensional genomic data 92%
- Variant calling tool evaluation for variable size indel calling from next generation whole genome and targeted sequencing data 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.