Human-supervised, large language model-based clinical decision support aligned to national newborn protocols in Kenya: a pragmatic, early-stage evaluation
Kuria, T.; Kamau, G.; Makokha, F.; Omondi, P.; Mbugua, G.; David, K.; Mbugua, S.; Gitaka, J.
Show abstract
Introduction: Timely, protocol-adherent clinical decisions are crucial for reducing neonatal mortality in low-resource settings. Translating extensive national guidelines into bedside practice remains challenging. Objective: We developed and evaluated AIFYA, a human-supervised, large language model LLM based clinical decision support system CDSS aligned with Kenya's national newborn care protocols. Methods: This prospective mixed methods early stage evaluation guided by the DECIDE-AI framework embedded AIFYA into routine workflows at two public health facilities Level 5 and Level 4 in Bungoma County Kenya from September 2024 to June 2025. Primary outcomes were adoption measured by cumulative neonatal cases managed training reach assessed by credentialed healthcare workers HCWs and guideline and citation concordance evaluated through blinded review of 118 AI generated recommendations by two neonatologists with adjudication by a third. Secondary outcomes included protocol adherence and triage to decision time. Results: A total of 50 HCWs were trained and 550 neonatal cases were managed over 10 months. Among surveyed HCWs n equals 33, 76 percent were female with mean age 32.1 years. Expert review found 75 percent of recommendations were correct and 15 percent partially correct with strong inter rater reliability weighted Cohen's kappa 0.85 and 95 percent CI 0.79 to 0.91. Citation accuracy was 96 percent. In 40 complex dosing scenarios 75 percent of outputs were rated correct. The median triage to decision time was 23 minutes with interquartile range 18 to 31. Implementation was supported by an offline first architecture and a facility based coaching model sustaining engagement despite staff turnover. Conclusion: A human supervised AI CDSS directly and transparently anchored to national clinical guidelines can be successfully implemented in routine low resource neonatal care settings. The system demonstrated high user adoption and strong expert rated concordance. High citation accuracy builds clinical trust ensuring safety and enabling auditable AI. These findings support progression to controlled multi site trials to evaluate clinical effectiveness. Keywords: Neonatal care Clinical decision support system Large language model Artificial intelligence Human supervised Low resource settings Guideline adherence Digital health Kenya
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- ePOCT+ and the medAL-suite: Development of an electronic clinical decision support algorithm and digital platform for pediatric outpatients in low- and middle-income countries 95%
- Implementation of Smart Triage combined with a quality improvement program for children presenting to facilities in Kenya and Uganda: An interrupted time series analysis 95%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 95%
Similar papers in this journal
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 95%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 92%
Similar papers in this journal
- Identifying clinical skill gaps of healthcare workers using a digital clinical decision support algorithm during outpatient pediatric consultations in primary health centers in Rwanda 95%
- “It reminds me and motivates me” : Human-centered design and implementation of an interactive, SMS-based digital intervention to improve early retention on antiretroviral therapy: usability and acceptability among new initiates in a high-volume, public clinic in Malawi 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.