A Neo4j-Based Framework for Integrating Clinical Data with Medical Ontologies: Performance Optimization and Quality Measure Applications in Healthcare
Jeon, S.
Show abstract
BackgroundElectronic Health Records face a fundamental challenge: the semantic gap between relational data storage and clinical reasoning patterns. Traditional databases struggle with complex healthcare queries requiring multiple joins and temporal analysis, creating performance bottlenecks that limit real-time clinical applications. MethodsWe developed a Neo4j-based framework integrating MIMIC-IV clinical data (1,504 patients, 4,967 admissions) with SNOMED CT medical ontology through ICD-10-CM mappings. The implementation created a unified graph comprising 625,708 nodes and 2,189,093 relationships, with systematic preservation of temporal and semantic connections. ResultsPerformance analysis demonstrated substantial improvements over PostgreSQL across five query types, with Neo4j showing 5.4x to 48.4x faster execution times. The framework successfully enabled three clinical applications: ventilator-associated pneumonia temporal analysis (revealing 47.79% pneumonia rates among ventilated ICU stays), hypertension semantic network mapping through multi-level SNOMED-CT relationships, and Medicare Part D quality measure monitoring. Notably, the system identified that 96.7% of eligible diabetic patients lacked statin prescriptions, demonstrating practical utility for healthcare quality improvement initiatives. ConclusionThis graph-based approach provides a robust foundation for next-generation clinical decision support systems by bridging the gap between fragmented clinical data and integrated patient-centric analysis. The frameworks demonstrated performance advantages and practical applications in quality measure monitoring establish its potential for addressing real-world healthcare challenges while supporting the transition toward more effective, evidence-based patient care.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 95%
- Temporal Relationship of Computed and Structured Diagnoses in Electronic Health Record Data 95%
- An Ontology-based Approach to Guide and Document Variable and Data Source Selection and Data Integration Process to Support Integrative Data Analysis in Cancer Outcomes Research 94%
Similar papers in this journal
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 96%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 96%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 95%
Similar papers in this journal
- Development of a COVID-19 Application Ontology for the ACT Network 97%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 97%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 94%
Similar papers in this journal
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 97%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 95%
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 95%
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 96%
- De-novo FAIRification via an Electronic Data Capture system by automated transformation of filled electronic Case Report Forms into machine-readable data 94%
- Signal from the Noise: A Mixed Methods Process Mining Approach to Evaluate Care Pathways. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.