Back

Ancestry-Aware Modeling of Dark Cyclobutane Pyrimidine Dimer Formation Integrating GTEx Skin Transcriptomes and Evolutionary Genomics

Essel Arthur, K.

2025-10-28 genomics
10.1101/2025.10.25.684582 bioRxiv
Show abstract

Ultraviolet (UV) radiation induces cyclobutane pyrimidine dimers (CPDs) in DNA, initiating mutagenic cascades that underlie photocarcinogenesis. Even hours after irradiation, dark-CPDs photoproducts generated via melanin-mediated chemiexcitation continue to form in melanocytes (4-6). While melanin confers photoprotection, its oxidative by-products can paradoxically extend DNA damage. Here we empirically validate an ancestry-aware computational framework coupling population pigmentation genetics with transcriptional regulation of DNA-repair pathways using Genotype-Tissue Expression (GTEx v9) skin RNA-seq data. We analyzed 604 GTEx donors from sun-exposed and non-exposed skin (lower leg, suprapubic) across inferred ancestry axes. Expression modules for melanin synthesis (TYR, TYRP1, SLC24A5, MC1R) and nucleotide-excision/oxidative-repair (XPC, DDB2, POLH, OGG1) were examined through differential expression, random-forest modeling, and 1,000-fold bootstrap uncertainty quantification. POLH and DDB2 were significantly upregulated in sun-exposed tissue (log2FC = 0.88 +/- 0.12 and 0.64 +/- 0.18; FDR < 0.05), whereas SLC24A5 and TYR displayed ancestry-linked gradients consistent with prior GWAS (8-10, 15, 16). Predictive modeling of a composite dark-CPD index achieved mean R^2 = 0.62 +/- 0.04 and RMSE = 0.21 +/- 0.03 (95 % CI), highlighting SLC24A5 (27 %) and XPC (19 %) as major contributors. These results empirically demonstrate co-regulation between pigmentation and repair pathways within realistic transcriptomic uncertainty bounds. Our integrative approach provides a reproducible, ancestry-aware platform for equitable dermatogenomic risk assessment and mechanistic insight into delayed UV mutagenesis.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.