Back

preSCRIPT: Large-scale prescription search and annotation engine for pharmacogenomic studies

Pieczarka, M.; Pienkowski, P.; Konowalska, P.; Grubarek, S.; Hajto, J.; Hoinkis, D.; Piechota, M.; Borczyk, M.; Korostynski, M.

2026-04-29 genetic and genomic medicine
10.64898/2026.04.28.26351989 medRxiv
Show abstract

Pharmacogenetics (PGx) has traditionally focused on a small number of high-impact variants affecting drug response due to the fact that PGx studies are labor-intensive and therefore low-throughput. Population biobanks linked to electronic health records (EHRs), including the UK Biobank (UKB) with prescription data for [~]230,000 individuals offer opportunities to scale PGx research. This, however, comes with a challenge as EHRs do not provide direct treatment response outcomes. One way to overcome this is to draw indirect drug response phenotypes from prescription records. Here, we propose preSCRIPT, a framework to filter and annotate raw prescriptions from the UKB to derive phenotypes for analyses which includes an algorithm to distinguish short prescription gaps from true dose changes. As a proof of concept, we applied preSCRIPT to warfarin, paracetamol, codeine, amitriptyline, simvastatin, aspirin, and amlodipine and derived therapy length and median daily doses. We tested associations for those seven drugs and two phenotypes across SNPs, cytochrome P450 (CYP) genes, and HLA alleles. We replicated known associations such as CYP2D6 variants with amitriptyline therapy length and dose, CYP2C9/CYP4F2/CYP2C19 with warfarin dose, and CYP2D6 with codeine dose. For drugs without formal PGx guidelines, we identified an association between CYP2D6 activity and aspirin therapy length and several SNPs, including rs62471929 (CYP3A5), a variant for amlodipine dose, replicated in an independent hold-out set. Overall, our study shows that preSCRIPT can recover established PGx associations, prioritize exploratory novel candidate loci, and may serve as a tool for large-scale pharmacogenomics.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 3%
9.7%
2
The Pharmacogenomics Journal
11 papers in training set
Top 0.1%
6.7%
3
Genome Medicine
183 papers in training set
Top 0.4%
6.7%
4
Nature Communications
5641 papers in training set
Top 29%
4.8%
5
Communications Medicine
113 papers in training set
Top 0.5%
4.8%
6
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.8%
4.0%
7
Bioinformatics Advances
203 papers in training set
Top 2%
3.5%
8
PLOS ONE
5266 papers in training set
Top 39%
3.2%
9
Scientific Reports
3612 papers in training set
Top 42%
2.4%
10
Clinical Pharmacology & Therapeutics
25 papers in training set
Top 0.1%
2.4%
11
Nature Genetics
286 papers in training set
Top 2%
2.4%
50% of probability mass above
12
PLOS Computational Biology
1863 papers in training set
Top 13%
2.1%
13
Cell Genomics
172 papers in training set
Top 2%
2.1%
14
eLife
5828 papers in training set
Top 44%
2.1%
15
The American Journal of Human Genetics
234 papers in training set
Top 2%
1.9%
16
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
1.7%
17
BMC Genomics
406 papers in training set
Top 4%
1.7%
18
GigaScience
212 papers in training set
Top 2%
1.7%
19
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.6%
1.7%
20
JAMIA Open
42 papers in training set
Top 1%
1.4%
21
npj Digital Medicine
118 papers in training set
Top 3%
1.3%
22
Wellcome Open Research
67 papers in training set
Top 0.9%
1.3%
23
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.1%
24
Genetic Epidemiology
55 papers in training set
Top 0.5%
1.1%
25
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.2%
1.1%
26
iScience
1154 papers in training set
Top 30%
1.0%
27
Frontiers in Genetics
230 papers in training set
Top 5%
1.0%
28
Genetics in Medicine
78 papers in training set
Top 0.9%
1.0%
29
BioData Mining
22 papers in training set
Top 0.6%
1.0%
30
Trials
29 papers in training set
Top 0.9%
1.0%