PGViS: Personal Genome Variant interpretation Score for lung cancer genomes
Surana, P.; Dutta, P.; Boffetta, P.; Davuluri, R. V.
Show abstract
Inherited lung cancer risk arises from both protein-coding and non-coding germline variants, but the functional non-coding component is largely uncharacterized. Genome-wide association studies and polygenic risk scores identify tag variants, not causal ones. Neither resolves which regulatory element is perturbed. DNA foundation models such as DNABERT decode non-coding variant effects directly from sequence, without a large GWAS cohort. What is missing is a patient-level framework linking these predictions to population-level variant prevalence. We present PGViS (Personal Genome Variant interpretation Score), a statistical framework that quantifies individual non-coding germline regulatory risk in non-small cell lung cancer (NSCLC). PGViS integrates three variant-level signals: DNABERT-predicted disruption at transcription factor binding and splice sites, the cancer v/s reference alternate allele frequency shift, and a regulatory interaction term derived from cancer-to-reference allele frequency ratios. Each signal is weighted by cohort prevalence which are aggregated into a single ancestry-matched, reference-normalized score per patient. We applied PGViS to germline whole-genome sequencing from 1,102 TCGA and CPTAC patients, using the 1000 Genomes Project (n = 2,504 individuals) as the normal population reference. PGViS separated adenocarcinoma (AD) and squamous cell carcinoma from controls in European ancestry and East Asian AD. Genes at contributing loci were enriched for PI3K-Akt, Wnt, DNA damage response, and epithelial-mesenchymal transition programs. Smoking-stratified analysis concentrated this signal on canonical NSCLC driver pathways. PGViS is modular: it accommodates cohorts with broader ancestral representation and can adapt to other solid tumors, offering a cost-effective route to personal-genome risk assessment from germline variants alone.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Discovering genetic interactions bridging pathways in genome-wide association studies 95%
- Prediction and functional interpretation of inter-chromosomal genome architecture from DNA sequence with TwinC 94%
- Chromatin-informed inference of transcriptional programs in gynecologic and basal breast cancers 94%
Similar papers in this journal
Similar papers in this journal
- Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data 94%
- BASCULE: Bayesian inference and clustering of mutational signatures leveraging biological priors 94%
- Juggling offsets unlocks RNA-seq tools for fast scalable differential usage, aberrant splicing and expression analyses. 93%
Similar papers in this journal
- Allele-specific genomics decodes gene targets and mechanisms of the non-coding genome 94%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 94%
- Whole genome base-wise aggregation and functional prediction for human non-coding regulatory variants 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.