Feasibility of converting Japanese oncology electronic medical records into the Observational Medical Outcomes Partnership Common Data Model and data quality assessment
Aoyagi, Y.; Terao, S.; Masahiro, B.; Nomura, K.; Ikeda, Y.; Sato, A.
Show abstract
The potential of utilizing Japanese electronic medical record (EMR) data in global observational research is significant because of high EMR adoption and universal health insurance. However, a few studies have addressed the conversion of Japanese EMR data to the Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) standard, which regulates EMR data for global observational research. In this study, we investigated the feasibility of converting Japanese oncology EMR data to the OMOP CDM and applying the Observational Health Data Sciences and Informatics (OHDSI) tools for analysis. We focused on data from the National Cancer Center Hospital East, encompassing 8,447 patients with breast cancer between January 2015 and November 2023. The main objectives included vocabulary standardization and data structure standardization. The anonymized dataset included clinical information such as patient demographics, diagnoses, treatments, and laboratory results. A total of 3,697 unique disease names, 987 specimen test result terms, and 1,144 drug terms were successfully mapped to OMOP CDM standards, with IC-10 terms showing the highest success rate for disease names. A total of 90% of clinical terms were successfully mapped to OMOP CDM standards, with 80% of source data fully integrated. However, only 32 surgical terms were identified. The feasibility of converting EMR data to OMOP CDM was evaluated by mapping source terms, comparing local raw datasets, and conducting a comprehensive quality assessment using a Data Quality Dashboard. A total of 1,991 validation checks were performed to evaluate the validity of data, suitability, and completeness. The results revealed 24 checks flagged as FAIL or ERROR, with the most frequent issues in the measurement table (10 errors). Despite these issues, the conversion process demonstrated high feasibility. Overall, this study positions Japan as a key player in international observational oncology research, enhancing the global understanding of treatment effectiveness and patient outcomes in real-world settings.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An Ontology-based Approach to Guide and Document Variable and Data Source Selection and Data Integration Process to Support Integrative Data Analysis in Cancer Outcomes Research 94%
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 94%
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 94%
Similar papers in this journal
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 93%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 93%
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 93%
Similar papers in this journal
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 96%
- Development of a COVID-19 Application Ontology for the ACT Network 94%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 93%
Similar papers in this journal
- An interactive retrieval system for clinical trial studies with context-dependent protocol elements 96%
- Postoperative mortality analysis of national Japanese Diagnosis Procedure Combination database with a focus on regional comparisons and changes over time 94%
- A clinical specific BERT developed with huge size of Japanese clinical narrative 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.