Back

Developing an OMOP-Standardized Prostate Cancer Database and Improving Data Quality Using NLP and PSA-Based Algorithms

Wang, J.; Jackson, J. C.; Garza, A.; Nalla, S.; Ninnemann, T.; Zhang, Y.; Kuo, Y.-F.

2026-07-02 health informatics
10.64898/2026.06.30.26356984 medRxiv
Show abstract

Objective: To develop and evaluate an Observational Medical Outcomes Partnership (OMOP) standardized prostate cancer database from the University of Texas Medical Branch (UTMB) Epic Electronic Health Record (EHR) and improve data quality using natural language processing (NLP) and prostate-specific antigen (PSA) based algorithms. Materials and Methods: We built a data pipeline to transform UTMB Epic EHR data from 2010 to 2021 into OMOP Common Data Model (CDM) v5.4. Data quality was assessed by comparing the OMOP-standardized data with Galveston Cancer Registry data using availability agreement, Cohen's kappa, and Intraclass Correlation Coefficient. NLP was used to extract PSA, Gleason score, and cancer stage from clinical text, and PSA-based algorithms were used to identify missing treatment and biochemical recurrence. Results: We extracted 815 analytic cases from UTMB EHR. Among them, 700, or 85.9%, were complete and concordant with the cancer registry. PSA showed excellent value agreement. Structured Gleason score and stage data were sparse, with fewer than 20 cases, but NLP greatly improved capture. Treatment agreement was good compared with the cancer registry and improved slightly for radical prostatectomy after applying a PSA-based algorithm. Using PSA trajectories, we identified 60 cases of biochemical recurrence. Discussion: The OMOP-standardized data from UTMB showed good agreement with the cancer registry. However, structured EHR fields incompletely captured diagnosis, pathology, and treatment details. NLP and PSA-based algorithms substantially improved data capture. Manual review also revealed errors in registry data, showing that OMOP-standardized EHR data can complement and help improve cancer registry quality. Conclusion: OMOP standardization combined with NLP and PSA-based algorithms improved prostate cancer data quality and research readiness.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.1%
21.7%
2
JAMIA Open
42 papers in training set
Top 0.1%
11.8%
3
PLOS ONE
5266 papers in training set
Top 25%
6.7%
4
BMJ Open
601 papers in training set
Top 3%
6.7%
5
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.4%
5.4%
50% of probability mass above
6
The Prostate
11 papers in training set
Top 0.1%
4.0%
7
BMJ Health & Care Informatics
15 papers in training set
Top 0.2%
4.0%
8
JMIR Public Health and Surveillance
45 papers in training set
Top 0.1%
4.0%
9
Scientific Reports
3612 papers in training set
Top 34%
3.2%
10
International Journal of Medical Informatics
26 papers in training set
Top 0.4%
3.2%
11
Annals of Internal Medicine
28 papers in training set
Top 0.1%
3.1%
12
BMC Medical Research Methodology
47 papers in training set
Top 0.6%
2.1%
13
International Journal of Radiation Oncology*Biology*Physics
25 papers in training set
Top 0.3%
1.9%
14
Modern Pathology
22 papers in training set
Top 0.2%
1.9%
15
JMIR Medical Informatics
18 papers in training set
Top 0.7%
1.1%
16
Scientific Data
209 papers in training set
Top 2%
1.0%
17
Bioinformatics
1204 papers in training set
Top 8%
1.0%
18
npj Digital Medicine
118 papers in training set
Top 3%
1.0%
19
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.7%
0.8%
20
BMC Cancer
67 papers in training set
Top 2%
0.8%
21
Biology Methods and Protocols
61 papers in training set
Top 3%
0.6%
22
Journal of Biomedical Informatics
47 papers in training set
Top 1%
0.6%
23
Cancer Medicine
26 papers in training set
Top 1%
0.6%
24
JAMA Network Open
130 papers in training set
Top 5%
0.6%
25
Communications Medicine
113 papers in training set
Top 6%
0.6%