JMIR Medical Informatics
◐ JMIR Publications Inc.
All preprints, ranked by how well they match JMIR Medical Informatics's content profile, based on 18 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Johnson, A. M.; Adibuzzaman, M.; Griffin, P.; Bikak, M.
Show abstract
ObjectivesThe aim of this research was to develop data-driven models using electronic health records (EHRs) to conduct clinical studies for predicting clinical outcomes through probabilistic analysis that considers temporal aspects of clinical data. We assess the efficacy of antibiotics treatment and the optimal time of initiation for in-hospitalized diagnosed with acute exacerbation of COPD (AECOPD) as an application to probabilistic modeling. Materials and MethodsWe developed a semi-automatic Markov Chain Monte Carlo (MCMC) modeling and simulation approach that encodes clinical conditions as computable definitions of health states and exact time duration as input for parameter estimations using raw EHR data. We applied the MCMC approach to the MIMIC-III clinical database, where ICD-9 diagnosis codes (491.21, 491.22, and 494.1) were used to identify data for 697 AECOPD patients of which 25.9% were administered antibiotics. ResultsThe average time to antibiotic administration was 27 hours, and 32% of patients were administered vancomycin as the initial antibiotic. The model simulations showed a 50% decrease in mortality rate as the number of patients administered antibiotics increased. There was an estimated 5.5% mortality rate when antibiotics were initially administrated after 48 hours vs 1.8% when antibiotics were initially administrated between 24 and 48 hours. DiscussionOur findings suggest that there may be a mortality benefit in initiation of antibiotics early in patient with severe respiratory failure in settings of COPD exacerbations warranting an ICU admission. ConclusionProbabilistic modeling and simulation methods that considers temporal aspects of raw clinical patient data can be used to adequately generate evidence for clinical guidelines.
Sharma, R.
Show abstract
This paper re-imagines a world of abundance in the treatment of chronic diseases such as Tpe 2 Diabetes. It asks: what if preventive and diagnostic remedies were widely made available across the world, informed by the latest medical research? As Proof-of-Concept of a proposed solution, the paper describes the development and validation of a local Large Language Models (local-LLMs) based on Graph-based Retrieval-Augmented Generation (GraphRAG) for managing Gestational Diabetes Mellitus (GDM). The research thus seeks new insights into optimizing GDM treatment through a knowledge graph architecture, contributing to a deeper understanding of how artificial intelligence can extend medical expertise to underserved populations globally. The study employs an agile, prototyping approach utilizing GraphRAG to enhance knowledge graphs by integrating retrieval-based and generative artificial intelligence techniques. Training data was from academic papers published between January 2000 and May 2024 using the Semantic Scholar API and analyzed by mapping complex associations within GDM management to create a comprehensive knowledge graph architecture. It is categorically stated that, since the primary research objective was to establish the feasibility of a GraphRAG local-LLM PoC, no human subjects nor actual patient datasets were used. Empirical results indicate that the GraphRAG-based Proof of Concept outperforms open-source LLMs such as ChatGPT, Claude, and BioMistral across key evaluation metrics. Specifically, GraphRAG achieves superior accuracy with BLEU scores of 0.99, Jaccard similarity of 0.98, and BERT scores of 0.98, offering significant implications for personalized medical insights that enhance diagnostic accuracy and treatment efficacy. This research offers a novel perspective on applying GraphRAG-enabled LLM technologies to GDM management, providing valuable insights that extend current understanding of AI applications in healthcare. The studys findings contribute to advancing the feasibility of GenAI for proactive GDM treatment and extending medical expertise to underserved populations globally.
Mondal, A.; Naskar, A.
Show abstract
BackgroundThe escalating global burden of diabetes necessitates innovative management strategies. Artificial intelligence, particularly large language models like GPT-4, presents a promising avenue for improving guideline adherence in diabetes care. Such technologies could revolutionize patient management by offering personalized, evidence-based treatment recommendations. MethodsA comparative, blinded design was employed, involving 50 hypothetical diabetes mellitus case summaries, emphasizing varied aspects of diabetes management. GPT-4 evaluated each summary for guideline adherence, classifying them as compliant or non-compliant, based on the ADA guidelines. A medical expert, blinded to GPT-4s assessments, independently reviewed the summaries. Concordance between GPT-4 and the experts evaluations was statistically analyzed, including calculating Cohens kappa for agreement. ResultsGPT-4 labelled 30 summaries as compliant and 20 as non-compliant, while the expert identified 28 as compliant and 22 as non-compliant. Agreement was reached on 46 of the 50 cases, yielding a Cohens kappa of 0.84, indicating near-perfect agreement. GPT-4 demonstrated a 92% accuracy, with a sensitivity of 86.4% and a specificity of 96.4%. Discrepancies in four cases highlighted challenges in AIs understanding of complex clinical judgments related to medication adjustments and treatment modifications. ConclusionGPT-4 exhibits promising potential to support health-care professionals in reviewing diabetes management plans for guideline adherence. Despite high concordance with expert assessments, instances of non-agreement underscore the need for AI refinement in complex clinical scenarios. Future research should aim at enhancing AIs clinical reasoning capabilities and exploring its integration with other technologies for improved healthcare delivery.
Dominguez, J.; Prociuk, D.; Marovic, B.; Cyras, K.; Cocorascu, O.; Ruiz, F.; Mi, E.; Mi, E.; Ramtale, C.; Rago, A.; Darzi, A.; Toni, F.; Curcin, V.; Delaney, B. C.
Show abstract
I.A. ObjectiveClinical Decision Support (CDS) systems (CDSSs) that integrate clinical guidelines need to reflect real-world co-morbidity. In patient-specific clinical contexts, transparent recommendations that allow for contraindications and other conflicts arising from co-morbidity are a requirement. We aimed to develop and evaluate a non-proprietary, standards-based approach to the deployment of computable guidelines with explainable argumentation, integrated with a commercial Electronic Health Record (EHR) system in a middle-income country. B. Materials and MethodsWe used an ontological framework, the Transition-based Medical Recommendation (TMR) model, to represent, and reason about, guideline concepts, and chose the 2017 International Global Initiative for Chronic Obstructive Lung Disease (GOLD) guideline and a Serbian hospital as the deployment and evaluation site, respectively. To mitigate potential guideline conflicts, we used a TMR-based implementation of the Assumptions-Based Argumentation framework extended with preferences and Goals (ABA+G). Remote EHR integration of computable guidelines was via a microservice architecture based on HL7 FHIR and CDS Hooks. A prototype integration was developed to manage COPD with comorbid cardiovascular or chronic kidney diseases, and a mixed-methods evaluation was conducted with 20 simulated cases and five pulmonologists. C. ResultsPulmonologists agreed 97% of the time with the GOLD-based COPD symptom severity assessment assigned to each patient by the CDSS, and 98% of the time with one of the proposed COPD care plans. Comments were favourable on the principles of explainable argumentation; inclusion of additional co-morbidities were suggested in the future along with customisation of the level of explanation with expertise. D. ConclusionAn ontological model provided a flexible means of providing argumentation and explainable artificial intelligence for a long-term condition. Extension to other guidelines and multiple co-morbidities is needed to test the approach further. E. FundingThe project was funded by the British government through the Engineering and Physical Sciences Research Council (EPSRC) - Global Challenges Research Fund.1
Ma, Y.; Wang, C.; Cui, G.; Li, Y.; Yue, C.; Wang, W.
Show abstract
BackgroundPelvic fractures have consistently been a focal point in orthopedic research. This study aims to provide a comprehensive analysis of the literature on pel-vic fractures published between 1983 and 2023, revealing research trends, hotspots, and frontiers in this field. ObjectiveThis study aimed to provide a comprehensive bibliometric and knowledge graph-based analysis of pelvic fracture literature published between 1983 and 2023, identifying research trends, hotspots, and emerging frontiers in this field. MethodsWe searched the Web of Science database using a predefined strategy restricted to review articles and original research articles, excluding studies outside orthopedics and surgery. Medical entities and relationships were extracted to construct a comprehensive knowledge graph. Entity recognition, relationship extraction, and network topology analyses were performed to map research evolution and collaboration networks. ResultsA total of 5248 articles were included for analysis. The results show a steady increase in the annual publication of pelvic fracture research, particularly after 2005. The United States, Germany, and China are the top three coun-tries in terms of the number of publications, with the University of Wash-ington, University of California, and University of San Francisco ranking the top three regions. Pohlmann T published the most significant number of ar-ticles, and Vaidya R was the strongest citation bursts author. Research on pelvic fractures has made significant progress over the past forty years, espe-cially in treatment techniques and methods. Bibliometric analysis reveals re-search hotspots in this field, such as hemostasis control, fracture fixation techniques, and osteoporotic fractures. Conclusions: This study employs bibliometrics to quantify and delineate the contemporary research landscape and trends in pelvic fracture research, aspiring to provide scholars with a compass for navigating the realm of pelvic fracture-related research.
Mondal, A.; Naskar, A.; Roy Choudhury, B.; Chakraborty, S.; Biswas, T.; Sinha, S.
Show abstract
BackgroundThe integration of large language models (LLMs) such as GPT-4 into healthcare presents potential benefits and challenges. While LLMs have shown promise in applications ranging from scientific writing to personalized medicine, their practical utility and safety in clinical settings remain under scrutiny. Concerns about accuracy, ethical considerations and bias necessitate rigorous evaluation of these technologies against established medical standards. ObjectiveTo compare the completeness, necessity, dosage accuracy and overall safety of type 2 diabetes management plans created by GPT-4 with those devised by medical experts. MethodsThis study involved a comparative analysis using anonymized patient records from a healthcare setting in West Bengal, India. Management plans for 50 Type 2 diabetes patients were generated by GPT-4 and three blinded medical experts. These plans were evaluated against a reference management plan based on American Diabetes Society guidelines. Completeness, necessity and dosage accuracy were quantified and an error score was devised to assess the quality of the generated management plans. The safety of the management plans generated by GPT-4 was also assessed. ResultsResults indicated that medical experts management plans had fewer missing medications compared to those generated by GPT-4 (p=0.008). However, GPT-4 generated management plans included fewer unnecessary medications (p=0.003). No significant difference was observed in the accuracy of drug dosages (p=0.975). The overall error scores were comparable between human experts and GPT-4 (p=0.301). Safety issues were noted in 16% of the plans generated by GPT-4, highlighting potential risks associated with AI-generated management plans. ConclusionThe study demonstrates that while GPT-4 can effectively reduce unnecessary drug prescriptions, it does not yet match the performance of medical experts in terms of plan completeness and safety. The findings support the use of LLMs as supplementary tools in healthcare, underscoring the need for enhanced algorithms and continuous human oversight to ensure the efficacy and safety of AI applications in clinical settings. Further research is necessary to improve the integration of LLMs into complex healthcare environments.
Frexia, F.; Mascia, C.; Lianas, L.; Delussu, G.; Sulis, A.; Meloni, V.; Del Rio, M.; Zanetti, G.
Show abstract
The FAIR Principles are a set of recommendations that aim to underpin knowledge discovery and integration by making the research outcomes Findable, Accessible, Interoperable and Reusable. These guidelines encourage the accurate recording and exchange of structured data, coupled with contextual information about their creation, expressed in domain-specific standards and machine readable formats. This paper analyses the potential support to FAIRness of the openEHR e-health standard, by theoretically assessing the compliance with each of the 15 FAIR principles of a hypothetical Clinical Data Repository (CDR) developed according to the openEHR specifications. Our study highlights how the openEHR approach, thanks to its computable semantics-oriented design, is inherently FAIR-enabling and is a promising implementation strategy for creating FAIR-compliant CDRs.
Lim, H.; Yi, H.; Yoon, J. Y.; Kwon, H.; Lee, D.; Kim, N.
Show abstract
Diagnostic errors, including misdiagnoses and delayed clinical diagnoses, could affect outcomes of a significant patient population, particularly individuals presenting with rare diseases or non-specific symptoms. From rule-based diagnostic decision supporting systems (DDSS) to large language model (LLM) based tools for clinical reasoning have been developed to address these limitations. However, existing DDSS are often proprietary and difficult to integrate, and recent LLM-based tools remain hindered by operational challenges such as cost, resources constraint, and privacy concerns. Moreover, existing systems interpret electronic medical records (EMR) and generate diagnoses separately, limiting continuous evidence-based analysis and imposing repeated clinician involvement. In this paper, we present DDx-Finder, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns. A clinical case study demonstrates the systems feasibility and its potential to provide accessible, transparent, and systematic differential diagnostic support for complex cases.
Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.
Show abstract
The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.
li, j.; Gao, X.; Dou, T.; Gao, Y.; Zhu, W.
Show abstract
BackgroundLarge Language Models (LLMs) like GPT-4 demonstrate potential applications in diverse areas, including healthcare and patient education. This study evaluates GPT-4s competency against osteoarthritis (OA) treatment guidelines from the United States and China and assesses its ability in diagnosing and treating orthopedic diseases. MethodsData sources included OA management guidelines and orthopedic examination case questions. Queries were directed to GPT-4 based on these resources, and its responses were compared with the established guidelines and cases. The accuracy and completeness of GPT-4s responses were evaluated using Likert scales, while case inquiries were stratified into four tiers of correctness and completeness. ResultsGPT-4 exhibited strong performance in providing accurate and complete responses to OA management recommendations from both the American and Chinese guidelines, with high Likert scale scores for accuracy and completeness. It demonstrated proficiency in handling clinical cases, making accurate diagnoses, suggesting appropriate tests, and proposing treatment plans. Few errors were noted in specific complex cases. ConclusionsGPT-4 exhibits potential as an auxiliary tool in orthopedic clinical practice and patient education, demonstrating high accuracy and completeness in interpreting OA treatment guidelines and analyzing clinical cases. Further validation of its capabilities in real-world clinical scenarios is needed.
Ohno, K.; Hashimoto, S.
Show abstract
Background: Japan faces critical challenges in medical data interoperability, conceptualized as the "Ten Walls" obstructing the Japanese Health Data Space (JHDS) [1]. The Japanese Intensive Care Patient Database (JIPAD) - Japan's largest national ICU registry with 151 participating facilities - represents a high-quality critical care dataset that remains isolated from international data ecosystems. Objective: To develop a formal mapping of all 122 JIPAD variables to HL7 FHIR R4, characterize the nature and magnitude of semantic gaps, and assess the feasibility of JIPAD integration into the JHDS. Methods: All 122 JIPAD variables (Data Dictionary v3.7.2; Linkage Items List 20231020) were evaluated using ISO 21564 [8]-based semantic equivalence scoring across three tiers: High (direct FHIR R4 Core mapping), Partial (mapping via JP-Core Implementation Guide extensions [3]), and Low/No Equivalence (structural institutional gap). Semantically identical multi-instance fields (e.g., secondary disease codes x5) were consolidated into single mapping entries, yielding 114 mapping entries. Pseudonymization architecture was characterized from primary documentation. Results: Of 114 mapping entries representing the 122 JIPAD variables, 97 (85.1%) achieved High Equivalence via LOINC/SNOMED CT, and 12 (10.5%) achieved Partial Equivalence via JP-Core extensions, value-set translation, or FHIR R4 Core extension mechanisms - yielding a combined technical feasibility of 95.6% (109/114). Only 5 entries (4.4%) were classified as Low/No Equivalence, all attributable to Japan's proprietary disease classification system (288 adult codes; 165 pediatric codes) embedded in the DPC reimbursement framework, plus one Japan-specific procedure (PMX endotoxin adsorption) absent from international terminology systems. Variable-level mapping details are provided in Supplementary Table S1. Critically, JIPAD employs pseudonymization with record-linkage capability, enabling 99% DPC data matching - demonstrating that technical and design-level barriers to FHIR integration have already been resolved. Conclusion: JIPAD is technically and architecturally ready for FHIR integration at a 95.6% level. The remaining 4.4% barrier is exclusively institutional - rooted in MHLW policy frameworks governing the DPC disease classification system [6] - rather than technical. FHIR integration would further unlock pharmacoepidemiological and social epidemiological research currently inaccessible due to data isolation. As the sole national ICU registry providing high-acuity anchor data unavailable in general health records, JIPAD integration is essential for a clinically meaningful JHDS by 2027.
Schoenthaler, M.; Hempen, N.; Weymann, M.; von Bargen, M. F.; Glienke, M.; Elsaesser, A.; Behrens, M.; Binder, H.; Binder, N.
Show abstract
BackgroundTo provide more evidence in urolithiasis research, we have established the German Nationwide Register for RECurrent URolithiasis (RECUR) using local clinical data warehouses (CDWH). For RECUR and other registers relying on digitalized clinical data, it is crucial to ensure the datas reliability for answering scientific questions. In this work, we aim to compare the results of different CDWH-based queries on urolithiasis cases next to manual case extraction from the primary source. MethodsSources for data extraction included the Medical Center University of Freiburg (MCUF) hospital information system (HIS), MCUF performance data (a clinical data set with merged data from patients including data from various time points throughout their treatment), and MCUF reimbursement data. We extracted data on caseloads in urolithiasis algorithmically (performance and reimbursement data) and compared those to a reference group compiled of manually extracted data from the local HIS and algorithmically extracted data. ResultsAlgorithmic extraction based on performance data resulted in correct and complete case identification as compared to the reference group. The case numbers from manual extraction from HIS data and algorithmic extraction from reimbursement data differed by 14% and 12%, respectively. The reasons for deviations in HIS data included human errors and a lack of data availability from different wards. Deviations in reimbursement data arose primarily due to the merging of cases in the context of reimbursement mechanisms. As the CDWH at MCUF is part of the German Medical Informatics Initiative (MII), the results can be transferred to other medical centers with similar CDWH structure. ConclusionsThe current study provides firm evidence of the importance of clearly defining a studys target variable, e.g., urolithiasis cases, and a thorough understanding of the data sources and modes used to extract the target data. Our work clearly shows that, depending on various data sources, a case is not a case is not a case.
Simmons, A.; Takkavatakarn, K.; McDougal, M.; Dilcher, B.; Pincavitch, J.; Meadows, L.; Kauffman, J.; Klang, E.; Wig, R.; Smith, G.; Soroush, A.; Freeman, R.; Apakama, D. J.; Charney, A.; Kohli-Seth, R.; Nadkarni, G.; Sakhuja, A.
Show abstract
BackgroundHealthcare reimbursement and coding is dependent on accurate extraction of International Classification of Diseases-tenth revision - clinical modification (ICD-10-CM) codes from clinical documentation. Attempts to automate this task have had limited success. This study aimed to evaluate the performance of large language models (LLMs) in extracting ICD-10-CM codes from unstructured inpatient notes and benchmark them against human coder. MethodsThis study compared performance of GPT-3.5, GPT4, Claude 2.1, Claude 3, Gemini Advanced, and Llama 2-70b in extracting ICD-10-CM codes from unstructured inpatient notes against a human coder. We presented deidentified inpatient notes from American Health Information Management Association Vlab authentic patient cases to LLMs and human coder for extraction of ICD-10-CM codes. We used a standard prompt for extracting ICD-10-CM codes. The human coder analyzed the same notes using 3M Encoder, adhering to the 2022-ICD-10-CM Coding Guidelines. ResultsIn this study, we analyzed 50 inpatient notes, comprising of 23 history and physicals and 27 progress notes. The human coder identified 165 unique codes with a median of 4 codes per note. The LLMs extracted varying numbers of median codes per note: GPT 3.5: 7, GPT4: 6, Claude 2.1: 6, Claude 3: 8, Gemini Advanced: 5, and Llama 2-70b:11. GPT 4 had the best performance though the agreement with human coder was poor at 15.2% for overall extraction of ICD-10-CM codes and 26.4% for extraction of category ICD-10-CM codes. ConclusionCurrent LLMs have poor performance in extraction of ICD-10-CM codes from inpatient notes when compared against a human coder.
Jeon, S.
Show abstract
BackgroundElectronic Health Records face a fundamental challenge: the semantic gap between relational data storage and clinical reasoning patterns. Traditional databases struggle with complex healthcare queries requiring multiple joins and temporal analysis, creating performance bottlenecks that limit real-time clinical applications. MethodsWe developed a Neo4j-based framework integrating MIMIC-IV clinical data (1,504 patients, 4,967 admissions) with SNOMED CT medical ontology through ICD-10-CM mappings. The implementation created a unified graph comprising 625,708 nodes and 2,189,093 relationships, with systematic preservation of temporal and semantic connections. ResultsPerformance analysis demonstrated substantial improvements over PostgreSQL across five query types, with Neo4j showing 5.4x to 48.4x faster execution times. The framework successfully enabled three clinical applications: ventilator-associated pneumonia temporal analysis (revealing 47.79% pneumonia rates among ventilated ICU stays), hypertension semantic network mapping through multi-level SNOMED-CT relationships, and Medicare Part D quality measure monitoring. Notably, the system identified that 96.7% of eligible diabetic patients lacked statin prescriptions, demonstrating practical utility for healthcare quality improvement initiatives. ConclusionThis graph-based approach provides a robust foundation for next-generation clinical decision support systems by bridging the gap between fragmented clinical data and integrated patient-centric analysis. The frameworks demonstrated performance advantages and practical applications in quality measure monitoring establish its potential for addressing real-world healthcare challenges while supporting the transition toward more effective, evidence-based patient care.
Harper, A.; Monks, T.; Wilson, R.; Redaniel, M. T.; Eyles, E.; Jones, T.; Penfold, C.; Elliott, A.; Keen, T.; Pitt, M.; Blom, A.; Whitehouse, M.; Judge, A.
Show abstract
ObjectivesTo develop a simulation model to support orthopaedic elective capacity planning. MethodsAn open-source, generalisable discrete-event simulation was developed, including a web-based application. The model used anonymised patient records between 2016-2019 of elective orthopaedic procedures from an NHS Trust in England. In this paper, it is used to investigate scenarios including resourcing (beds and theatres) and productivity (lengths-of-stay, delayed discharges, theatre activity) to support planning for meeting new NHS targets aimed at reducing elective orthopaedic surgical backlogs in a proposed ring fenced orthopaedic surgical facility. The simulation is interactive and intended for use by health service planners and clinicians. ResultsA higher number of beds (65-70) than the proposed number (40 beds) will be required if lengths-of-stay and delayed discharge rates remain unchanged. Reducing lengths-of-stay in line with national benchmarks reduces bed utilisation to an estimated 60%, allowing for additional theatre activity such as weekend working. Further, reducing the proportion of patients with a delayed discharge by 75% reduces bed utilisation to below 40%, even with weekend working. A range of other scenarios can also be investigated directly by NHS planners using the interactive web app. ConclusionsThe simulation model is intended to support capacity planning of orthopaedic elective services by identifying a balance of capacity across theatres and beds and predicting the impact of productivity measures on capacity requirements. It is applicable beyond the study site and can be adapted for other specialties. Strengths and Limitations of this studyO_LIThe simulation model provides rapid quantitative estimates to support post-COVID elective services recovery toward medium-term elective targets. C_LIO_LIParameter combinations include changes to both resourcing and productivity. C_LIO_LIThe interactive web app enables intuitive parameter changes by users while underlying source code can be adapted or re-used for similar applications. C_LIO_LIPatient attributes such as complexity are not included in the model but are reflected in variables such as length-of-stay and delayed discharge rates. C_LIO_LITheatre schedules are simplified, incorporating the five key orthopaedic elective surgical procedures. C_LI
Mazzucato, S.; Leeuwenberg, A.; van Doorn, S.; van Rosmalen, J.; Slurink, I. A. L.
Show abstract
Extracting clinical information from Dutch free-text medical notes requires language-specific annotation resources, yet Dutch primary care lacks a reusable event-annotation framework for infections, post-acute infection syndromes (PAIS), and related symptoms. We adapted the COVID-19 Annotated Clinical Text (CACT) framework to Dutch and applied it to GP notes for PAIS event extraction. The framework has three annotation layers: a DiagnosticExpression typology covering acute infections, post-acute syndromes, and relevant comorbidities; an eleven-subtype Evidence inventory grounded in Dutch primary-care testing practice; and explicit decision rules for the SOEP structure of Dutch general practitioner (GP) notes (Subjective, Objective, Evaluation, Plan), including the distinction between clinician hedging and patient-side hypotheticals. On a 200-note pilot, span-level F1 under the Lybarger criterion reached 0.51 [95% CI: 0.47, 0.55] across six core entities; restricted to spans both annotators noticed, conditional F1 reached 0.78 [0.75, 0.80], indicating that most disagreement stems from annotation coverage rather than label assignment. The adaptation illustrates how an English event-based clinical annotation framework can be extended to a new language and clinical setting, yielding a reusable resource for Dutch clinical NLP; which steps generalise beyond this case (CACT to Dutch primary care) and which are specific to Dutch or PAIS remain to be tested.
Ytsma, C. R.; Torralbo, A.; Fitzpatrick, N. K.; Pietzner, M.; Louloudis, I.; Nguyen, D.; Ansarey, S.; Denaxas, S.
Show abstract
ObjectiveThe aim of this study was to develop and validate an automated, scalable framework to harmonise fragmented UK primary care prescription records into a research-ready dataset by mapping four diverse medical ontologies to a unified, historically comprehensive reference standard. Materials and MethodsWe used raw prescription records for consented participants in the UK Biobank, in which participants are uniquely characterized by multiple data modalities. Primary care data were preprocessed by selecting one drug code if multiple were recorded, cleaning codes to match reference presentations, expanding code granularity based on drug descriptions, and updating outdated codes to a single reference version. Harmonisation entailed mapping British National Formulary (BNF) and Read2 codes to dm+d, the universal NHS standard vocabulary for uniquely identifying and prescribing medicines. Harmonised dm+d records were then homogenised to a single concept granularity, the Virtual Medicinal Product (VMP). We validated our methods by creating medication profiles mapping contemporary drug prescribing patterns in 312 physical and mental health conditions. ResultsWe preprocessed 57,659,844 records (100%) from 221,868 participants (100%). Of those, 48,950 records were dropped due to lack of drug code. 7,357,572 records (13%) used multiple ontologies. Most (76%) records were encoded in BNF and most had the code granularity expanded via the drug description (N=28,034,282; 49%). 41,244,315 records (72%) were harmonised to dm+d and 99.98% of these were converted to VMP as a homogeneous dataset. Across 312 diseases, we identified 23,352 disease-drug associations with 237 medications (represented as BNF subparagraphs) that survived statistical correction of which most resembled drug - indication pairs. ConclusionOur methodology converts highly fragmented and raw prescription records with inconsistent data quality into a streamlined, enriched dataset at a single reference, version, and granularity of information. Harmonised prescription records can be easily utilised by researchers to perform large-scale analyses in research.
Ostropolets, A.; Hripcsak, G.; Husain, S. A.; Richter, L. R.; Spotnitz, M.; Elhussein, A.; Ryan, P. B.
Show abstract
ObjectiveChart review as the current gold standard for phenotype evaluation cannot support observational research at scale. It is expensive, time-consuming, and variable. We aimed to evaluate the ability of structured data to support efficient patient status ascertainment and develop a standardized and scalable alternative to chart review. MethodsWe developed Knowledge-Enhanced Electronic Patient Profile Review system (KEEPER) that extracts a patients structured data elements relevant to a given phenotype and presents them in a standardized fashion that follows clinical reasoning principles. We evaluated its performance compared to manual chart review for four conditions (diabetes type I, acute appendicitis, end stage renal disease and chronic obstructive lung disease) using randomized two-period, two-sequence crossover design. Inter-method agreement, inter-rater agreement, accuracy, and review duration were measured. ResultsAscertaining patient status with KEEPER was twice as fast compared to manual chart review. 88.1% of the patients were classified concordantly using full chart and KEEPER, but agreement varied depending on the condition. Pairs of clinicians agreed in classification of patient status in 91.2% of the cases when using KEEPER compared to 76.3% when using full chart. Patient classification aligned with the gold standard in 88.1% and 86.9% of the cases respectively. ConclusionThis proof-of-concept study demonstrated that structured data can be used for efficient patient ascertainment if are limited to only relevant subset and organized according to the clinical reasoning principles. A system that implements these principles can achieve similar accuracy and higher inter-rater reliability compared to chart review at a fraction of time.
Schubert, M. C.; Wick, W.; Venkataramani, V.
Show abstract
Large Language Models (LLMs) offer potential in healthcare, especially in the evaluation of medical documents. This research introduces MedCheckLLM, a multi-step framework designed for the systematic assessment of medical records against established evidence-based guidelines, a process termed guideline-in-the-loop. By keeping the guidelines separate from the LLMs training data, this approach emphasizes validity, flexibility, and interpretability. Suggested evidence-based guidelines are externally accessed and fed back into the LLM for a evaluation. The method enables implementation of guideline updates and personalized protocols for specific patient groups without retraining. We applied MedCheckLLM to expert-validated simulated medical reports, focusing on headache diagnoses following International Headache Society guidelines. Findings revealed MedCheckLLM correctly extracted diagnoses, suggested appropriate guidelines, and accurately evaluated 87% of checklist items, with its evaluations aligning significantly with expert opinions. The system not only enhances healthcare quality assurance but also introduces a transparent and efficient means of applying LLMs in clinical settings. Future considerations must address privacy and ethical concerns in actual clinical scenarios.
Lichtner, G.; Haese, T.; Brose, S.; Roehrig, L.; Lysyakova, L.; Rudolph, S.; Uebe, M.; Sass, J.; Bartschke, A.; Hillus, D.; Kurth, F.; Sander, L. E.; Eckart, F.; Toepfner, N.; Berner, R.; Frey, A.; Doerr, M.; Vehreschild, J. J.; von Kalle, C.; Thun, S.
Show abstract
BackgroundThe COVID-19 pandemic has spurred large-scale, inter-institutional research efforts. To enable these efforts, researchers must agree on dataset definitions that not only cover all elements relevant to the respective medical specialty but that are also syntactically and semantically interoperable. Following such an effort, the German Corona Consensus (GECCO) dataset has been developed previously as a harmonized, interoperable collection of the most relevant data elements for COVID-19-related patient research. As GECCO has been developed as a compact core dataset across all medical fields, the focused research within particular medical domains demands the definition of extension modules that include those data elements that are most relevant to the research performed in these individual medical specialties. ObjectiveTo (i) specify a workflow for the development of interoperable dataset definitions that involves a close collaboration between medical experts and information scientists and to (ii) apply the workflow to develop dataset definitions that include data elements most relevant to COVID-19-related patient research in immunization, pediatrics, and cardiology. MethodsWe developed a workflow to create dataset definitions that are (i) content-wise as relevant as possible to a specific field of study and (ii) universally usable across computer systems, institutions, and countries, i.e., interoperable. We then gathered medical experts from three specialties (immunization, pediatrics, and cardiology) to the select data elements most relevant to COVID-19-related patient research in the respective specialty. We mapped the data elements to international standardized vocabularies and created data exchange specifications using HL7 FHIR. All steps were performed in close interdisciplinary collaboration between medical domain experts and medical information scientists. The profiles and vocabulary mappings were syntactically and semantically validated in a two-stage process. ResultsWe created GECCO extension modules for the immunization, pediatrics, and cardiology domains with respect to the pandemic requests. The data elements included in each of these modules were selected according to the here developed consensus-based workflow by medical experts from the respective specialty to ensure that the contents are aligned with the respective research needs. We defined dataset specifications for a total number of 48 (immunization), 150 (pediatrics), and 52 (cardiology) data elements that complement the GECCO core dataset. We created and published implementation guides and example implementations as well as dataset annotations for each extension module. ConclusionsThese here presented GECCO extension modules, which contain data elements most relevant to COVID-19-related patient research in immunization, pediatrics and cardiology, were defined in an interdisciplinary, iterative, consensus-based workflow that may serve as a blueprint for the development of further dataset definitions. The GECCO extension modules provide a standardized and harmonized definition of specialty-related datasets that can help to enable inter-institutional and cross-country COVID-19 research in these specialties.