Healthcare
○ MDPI AG
Preprints posted in the last 30 days, ranked by how well they match Healthcare's content profile, based on 17 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Isaiev, B.; Stukalova, I.
Show abstract
Background: The growing burden of lifestyle-related chronic diseases has increased the need for clinically interpretable decision-support tools capable of integrating artificial intelligence with evidence-based preventive nutrition. Although machine learning has shown considerable potential for health risk prediction, most existing approaches remain limited to isolated predictive models or conventional nutritional software, with little integration of multidimensional clinical assessment and personalized recommendations. Objective: To develop and internally validate NutrIA, a hybrid web-based Clinical Decision Support System (CDSS) that combines machine learning, validated clinical assessment, structured clinical reasoning and personalized nutritional recommendations for preventive medicine. Methods: NutrIA was developed using harmonized data from the National Health and Nutrition Examination Survey (NHANES, 1988 to 2018). A supervised machine learning model was trained to estimate 5-, 10- and 20-year all-cause mortality risk and subsequently integrated with an adaptive clinical questionnaire, validated screening instruments, nutritional indicators, dietary clustering, clinical phenotyping and a transparent rule-based recommendation engine within a unified web-based platform. Results: The predictive model achieved ROC-AUC values of 0.894, 0.914 and 0.923 for 5-, 10- and 20-year mortality prediction, respectively. The implemented CDSS incorporates an adaptive questionnaire (151 items), 39 validated clinical assessment instruments, 17 clinical phenotypes and 31 dietary clustering modules to generate individualized nutritional and lifestyle recommendations together with an automated clinical report. The integrated framework translates probabilistic risk estimates into clinically interpretable decision support for personalized preventive nutrition. Conclusions: NutrIA demonstrates the technical feasibility of integrating machine learning with knowledge-based clinical reasoning within a single web-based CDSS for preventive nutrition. Although external validation and prospective clinical evaluation are required before routine implementation, the proposed architecture represents a promising step toward clinically interpretable artificial intelligence for personalized nutritional care.
Rotenberg, S.; Chilufya-Moyo, M.; Valentine, A.; Smythe, T.; Forde, I.; Mitra, M.; Kuper, H.
Show abstract
Background: Reducing maternal mortality and improving newborn and child outcomes are targets of the Sustainable Development Goals. Evidence on how these efforts are reaching women with disabilities is lacking. Objectives: to estimate global, relative inequalities in stillbirth, neonatal, infant, and maternal mortality for women with disabilities compared to women without disabilities. Methods: We searched MEDLINE, Global Health, PsycINFO, and Embase from 1 January 2015 to 28 January 2026, to identify articles on disability and stillbirth, neonatal, infant, and maternal mortality. We included studies that had a recognised measure of disability as an exposure, a control group of women without disabilities and at least one of the four outcomes. A pooled estimate for each outcome was done using a random-effects meta-analysis of the minimally adjusted results. Results: We identified over 4,300 titles, of which 17 papers were eligible for inclusion. Almost all data came from nationally-representative data sources in high-income countries. We found that women with disabilities were 4.66 times more likely (95% C.I. 1.70-12.75) to experience maternal mortality compared to women without disabilities. Women with disabilities were also 37% more likely to have a stillbirth and 27% and 48% more likely to have a neonatal or infant death compared to women without disabilities, respectively. Conclusions: Women with disabilities consistently have higher incidence of maternal, stillbirth, neonatal, and infant mortality, even in countries that have relatively low incidence of these outcomes. There is a lack of evidence globally and particularly from LMICs, and on effective interventions to improve maternal and infant outcomes for women with disabilities.
Fu, Z.
Show abstract
Against the backdrop of nationwide inclusive education promotion, children with autism spectrum disorder, intellectual disabilities and other special educational needs (SEN) in Jiangxi Province have raised growing demands for equitable schooling. Paraeducators serve as a critical on-site support mechanism enabling SEN children to access mainstream classrooms; the adequacy of shadow teacher service provision and the maturity of corresponding multi-stakeholder support systems jointly determine the overall quality of local inclusive education. This study adopted mixed quantitative-qualified methods, including questionnaire surveys and semi-structured interviews, to investigate SEN children, paraeducators, general and special education teachers, as well as SEN caregivers across multiple prefecture-level cities in Jiangxi. Grounded in provincial special education policies and local frontline inclusive education practices, we systematically unpacked multidimensional service demands from four core stakeholder groups, diagnosed prominent practical bottlenecks restricting the sustainable operation of shadow teacher services, and constructed regionally tailored multi-layered support strategies aligned with Jiangxi's educational realities. The findings of this research offer empirical evidence and actionable policy references to advance high-quality inclusive education for SEN children across central China's Jiangxi Province.
Yoshimasu, T.; Abe, K.; Sato, M.; Ohashi, K.; Inao, T.; Ono, S.; Yokota, I.; Ogasawara, K.
Show abstract
Aim: Little is known about patient safety in a less consolidated obstetric system where various facilities, such as perinatal medical centers (PMCs), general hospitals, and clinics, collaborate under risk-based role differentiation. We aimed to compare maternal complications after cesarean section by facility type across area types (rural, provincial, and metropolitan) in Hokkaido, Japan. Methods: This retrospective cohort study used insurance claims data from Hokkaido (2018-2025). Two comparisons were conducted for a composite outcome of postpartum hemorrhage, infection, and thrombosis: a two-category comparison (PMC vs. non-PMC, combining general hospitals and clinics) across all areas, with an interaction term between facility and area type; a three-category comparison (PMC vs. general hospital vs. clinic) restricted to metropolitan areas. Generalized estimating equations with a Poisson distribution, accounting for clustering within facilities, were applied to estimate risk ratios. Results: A total of 1,822 participants underwent cesarean section. PMCs were associated with lower maternal complication rates compared to non-PMC facilities (adjusted RR 0.37, 95% CI 0.16-0.88 in rural areas; adjusted RR 0.20, 95% CI 0.11-0.37 in provincial areas). In metropolitan areas, PMCs and general hospitals were associated with lower maternal complication rates compared to clinics (PMC vs clinic: adjusted RR 0.42, 95% CI 0.18-0.97; general hospital vs clinic: adjusted RR 0.25, 95% CI 0.09-0.70). Conclusions: Higher-level facilities were associated with lower maternal complication rates after cesarean section in Japan. These findings provide important evidence for regional consolidation of obstetric care.
Sugawara, H.
Show abstract
Background: Whether corrective actions documented in medical safety incident reports rely on individual vigilance ("Safety-I") or on structural, system-level intervention ("Safety-II") has not been quantitatively evaluated on a national scale in Japan. We developed an automated classification pipeline to assign corrective-action free-text to a 7-level maturity scale (L0-L6) and computed two summary indices: the Safety Measure Quality Profile (SMQP), the full L0-L6 distribution, and the System-based Safety Measure Rate (SSMR), the proportion of non-L0 records classified L3-L6. Methods: We analyzed all 11,507 corrective-action free-text entries from the 2010 release of Japan's national medical accident and near-miss reporting database (Japan Council for Quality Health Care, JCQHC), comprising 8,804 near-miss (Hiyari-Hatto) and 2,703 accident (Jiko) reports. Records were classified using a five-stage hybrid pipeline: an expert-developed rule dictionary, TF-IDF + k-nearest-neighbor matching, cosine-similarity matching, a two-tier large-language-model (LLM) classifier, and a conservative priority-cascade fallback. SSMR was compared between near-miss and accident reports using a chi-square test, Wilson 95% confidence intervals, Cramer's V, and the risk difference (RD), against pre-specified minimal clinically important difference (MCID) criteria of RD >= 2 percentage points and Cramer's V >= 0.10. Results: Every record received a definitive L0-L6 label (0% unresolved). Overall, 16.6% of records were unclassifiable (L0); among the 9,599 classifiable (non-L0) records, individual-vigilance actions (L1) predominated (54.6% of all records), and only 11.82% (95% CI, 11.19-12.49%) met the SSMR criterion (L3-L6). SSMR was higher for accident reports than for near-miss reports (18.12% [95% CI, 16.69-19.64%] vs. 9.[95% CI,46% [95% CI, 8.80-10.17%]; RD = 8.66 percentage points; Cramer's V = 0.120; chi-square(1) = 136.97001), exceeding both pre-specified MCID thresholds. Conclusions: In this interim single-year analysis, the large majority of documented corrective actions in Japanese medical safety reports remained individual-vigilance-based rather than system-based, with accident reports showing a substantively, rather than merely statistically, higher proportion of system-based actions than near-miss reports. These findings support the feasibility of large-scale automated assessment of corrective-action quality and provide the rationale for the planned 16-year longitudinal analysis.
Sadeghi Naieni Fard, F.; Oppong, J. R.; Tiwari, C.; Boakye, K.; Fard, F.
Show abstract
Cancer prevalence is distributed unevenly across regions and caused by the interaction of multiple risk factors. Previous studies focused on the use of global modeling techniques to predict cancer at the county level that overlooks important spatial differences. This study aims to develop geographically weighted machine learning models to predict cancer prevalence at the census tract level in the United States and identify local determinants of cancer burden. First, a scoping review was conducted to find a list of measurable drivers of cancer in the United States. Using this list, the data of these variables for 84415 census tracts were obtained from the Center for Disease Control and Prevention PLACES dataset and other publicly accessible resources. Then, several predictive models, including Ordinary Least Squares (OLS) and Geographically Weighted Regression (GWR), as well as Random Forest, XGBoost, and Deep Neural Network and their geographically weighted counterparts, were developed and compared using the Coefficient of Determination, Root Mean Square Error, and Absolute Error. Results presented that geographically weighted models outperformed other methods, and geographically weighted XGBoost achieved the strongest and most consistent overall performance with pseudo-R2 ranging between 0.89 and 0.98. Feature importance analysis of this model illustrated that most important cancer drivers changed location by location. Aged people, racial composition, preventative behaviors, and metabolic conditions such as diabetes, hypertension, and high cholesterol were determined as influential predictors, although their relative importance varied across regions. These findings revealed the value of localized models at a small geographic scale to identify regional cancer risk patterns and help the allocation of proper resources to hotspot areas. Keywords: Cancer prevalence, Census tracts, geographically weighted machine learning models, Deep neural network, XGBoost, Random Forest, Ordinary Least Squares, risk factor, determinant
Sharifi Nowghabi, A.; Sharghilavan, S.; Bagheri, A.; Izadifar, M.
Show abstract
Wayfinding in hospitals is often hindered by ineffective signage; however, the cognitive mechanisms of healthcare wayfinding symbols comprehension remain under-researched. This study utilized eye-tracking and spatial gaze mapping to examine how visual complexity, abstraction, and human figuration modulate perception in 40 healthy adults viewing 24 hospital-related healthcare wayfinding symbols. Results indicate that pupil size is a sensitive physiological marker of cognitive load, significantly influenced by visual complexity ({chi}2 = 11.32, p = .022) and abstraction ({chi}2 = 7.49, p = .027). Human figuration reduced fixation duration and increased saccade amplitude, facilitating efficient semantic integration. Furthermore, human-centric healthcare wayfinding symbols elicited streamlined gaze trajectories, whereas abstract/complex designs induced chaotic scanpaths. These findings suggest that human figuration acts as a cognitive scaffold, reducing mental effort. We provide evidence-based guidelines for optimizing healthcare wayfinding symbols by prioritizing human body representations and balancing abstraction levels. HighlightO_LIPupil size indexes cognitive load during symbol comprehension. C_LIO_LIHuman figuration cuts fixation duration, boosting wayfinding efficiency. C_LIO_LIAbstract symbols increase pupil dilation, raising cognitive load. C_LI
Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.
Show abstract
Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.
Jo, A. A.
Show abstract
Maternal healthcare prediction systems often suffer from algorithmic biases due to socio-economic disparities and imbalanced datasets, limiting their effectiveness for equitable healthcare policymaking. This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India. The framework focuses on three critical health indicators:(1) Tetanus Toxoid (TT) booster uptake,(2) immunization coverage rates, and (3) the percentage of pregnant women completing four or more Antenatal Care (ANC) visits. To address fairness, we propose Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training. AESO is model-agnostic and adapts group equity weights in response to real-time disparities. We integrate SHAP, LIME, and feature permutation techniques for explainability, enabling transparent global and local interpretation. Empirical results demonstrate that MaternaAI significantly improves fairness metrics and model accuracy across diverse machine learning and deep learning models, offering interpretable and equitable decision support for public health stakeholders.
Delporte, M.; Tamimi, R.; Mehta, S.; Choi, E.; Zhang, Y.; Shi, Y.
Show abstract
Objective To develop and evaluate an automated large language model (LLM)-based framework for conducting meta-analyses of nutrition-related exposures and the risk of breast, ovarian, and uterine cancers. Design We developed MetaFemina, an automated evidence-synthesis pipeline for women's cancers that integrates keyword-based literature retrieval, LLM-assisted evidence extraction, and random-effects meta-analysis. We evaluated its performance against two recently published peer-reviewed meta-analyses and compared exposure-outcome associations across the three cancer types. Data sources PubMed articles identified through keyword-based searches of titles and abstracts. Methods MetaFemina was developed as a web platform that identifies relevant scientific articles, automatically extracts relevant information using LLMs, and synthesizes extracted evidence using random-effects meta-analysis. Additional analyses included assessment of heterogeneity, publication bias, and leave-one-out sensitivity analyses. The platform also provides sample size calculations based on synthesized effect sizes and generates visual summaries and plain-language interpretations. Results Compared with two recent peer-reviewed meta-analyses of folate and vitamin E intake in relation to breast cancer risk, MetaFemina demonstrated high sensitivity (81.82% and 80%, respectively) in identifying eligible studies and additionally retrieved relevant articles that had been missed by manual screening (27 and 13, respectively). Among 226 exposures considered, lutein and beta-carotene were significantly associated with lower risks of breast, ovarian, and uterine cancers. Vitamin D, antioxidants, and soy were significantly associated with lower risks of both breast and ovarian cancers, whereas calcium and folic acid were significantly associated with lower risks of both breast and uterine cancers. In contrast, iron, red meat, and copper were significantly associated with higher risks of both breast and uterine cancers. omega-6 fatty acids showed contrasting associations, being significantly associated with higher breast cancer risk but lower ovarian cancer risk. After restriction to dietary-intake studies, these cross-cancer significant associations remained statistically significant except for copper, which no longer met the two-study threshold for either breast or uterine cancer. Additionally, calcium became significantly associated with lower ovarian cancer risk, resulting in significant negative associations across all three cancer types, while vitamin E became significantly associated with lower breast cancer risk and remained significantly associated with lower ovarian cancer risk. Conclusions MetaFemina demonstrated high sensitivity for identifying relevant scientific literature, extracts key evidence, and performs statistically rigorous automated meta-analyses. The framework may facilitate more rapid evidence synthesis in nutritional epidemiology and may support researchers in study design, hypothesis generation, and interpretation of emerging evidence.
Roberts, L.
Show abstract
Objective. Triage of rheumatology outpatient referrals is a high-volume administrative task that consumes senior specialist time without advancing patient care. The human triage system is only moderately accurate and reproducible. We assessed whether contemporary large language models (LLMs) are able to perform well enough to support automating this task in practice. In addition, the effects of different prompting techniques on triage accuracy and cost was assessed to help identify to optimal approach. Methods. Twenty referral scenarios spanning the urgency spectrum, based on real referrals were created by a certified Australian rheumatologist. Four rheumatologists triaged all cases independently and blinded, to produce a consensus reference standard. Twenty-three LLMs each triaged every referral into one of five urgency categories, three times (1380 outputs per condition). The experiment was run with a simple prompt and repeated with a advanced prompt supplying explicit triage expectations and worked examples. Results. All 2760 attempts returned valid categories. Under the simple prompt, performance separated into distinct tiers, larger models were more accurate (Spearman rho=0.42; P=.047) and accuracy tracked cost. Advanced prompting minimised between-model variance in accuracy 5.3-fold (0.014 to 0.003; Levene P=.01), abolished the size-accuracy association (rho=-0.05; P=.83) and removed the accuracy-cost relationship. Leading models matched expert consensus on most cases, within or above the range reported for human triage. Under-triage errors persisted with some LLMs. Conclusion. Contemporary LLMs categorise rheumatology referral urgency as well or better than published human triage systems. Advanced LLM prompting methods substitute for the reasoning capability of larger models, suggesting that LLM performance on this task may not require the most expensive models. The tools to automate this administrative task appear to already exist. Strong candidate LLMs that might serve a production ready solution have been identified.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Nayak, K. S.; Nirgude, A. S.; Das, R.
Show abstract
Background Stroke remains one of the leading causes of mortality and long-term disability worldwide, with low- and middle-income countries bearing a disproportionate share of the global disease burden. In India, delays in risk identification, fragmented referral pathways, and limited continuity of preventive care present significant challenges, particularly in rural communities. As a frontline health worker Accredited Social Health Activists (ASHAs) are strategically positioned to support community-based stroke prevention; however, existing workflows are frequently constrained by multi-tasking, predominantly paper-based documentation and fragmented digital systems. Advances in mobile health, artificial intelligence along with digital health ecosystem provided by Ayushman Bharat Digital Mission (ABDM) provide an opportunity to strengthen community healthcare through integrated digital platforms. Objective This protocol describes the design, system architecture, and prospective evaluation framework of ASHA Assist India, an integrated AI-assisted mobile health platform intended to support community-based stroke prevention by connecting citizens, ASHA workers, Primary Health Centres (PHCs), and higher levels of healthcare facilities within a unified digital ecosystem. Methods ASHA Assist India has been designed as a modular, cloud-based digital health platform supporting standardized data collection, longitudinal health monitoring, referral management, and AI-assisted clinical decision support. The proposed system comprises four user-facing applications corresponding to citizens, ASHA workers, PHCs, and referral hospitals, integrated through a centralized backend providing authentication, secure data management, interoperability, analytics, and notification services. The AI framework includes three planned analytical modules: (i) population-level stroke risk stratification, (ii) longitudinal stroke risk prediction, and (iii) acute stroke symptom recognition. A prospective implementation study is planned to evaluate platform usability, feasibility, workflow integration, implementation outcomes, and operational performance within routine community healthcare settings. Future validation of the AI modules will be conducted using prospectively collected longitudinal datasets. Expected Impact The proposed platform aims to strengthen community-based stroke prevention by improving digital workflow integration, facilitating coordinated referral pathways, and supporting longitudinal monitoring through the existing healthcare providers at health and wellness centres like ASHA, Community Health Officers (CHOs), ANM, etc. Beyond stroke prevention, the modular architecture is intended to provide a scalable framework for future digital health programmes addressing multiple non-communicable diseases within primary healthcare systems. Publication of this protocol establishes a transparent implementation and evaluation framework that may guide future research, digital health innovation, and implementation science in resource-constrained settings.
Barzideh, A.; Devasahayam, A. J.; Marzolini, S.; Munce, S.; Sibley, K. M.; Inness, E. L.; Mansfield, A.
Show abstract
Background: Aerobic exercise is recommended during stroke rehabilitation to improve cardiorespiratory fitness and support recovery; however, participation rates remain low. While institutional and system-level barriers have been widely examined, less is known about how individual patient factors influence engagement in aerobic exercise during rehabilitation. Objectives: We aimed to determine whether depressive symptoms, apathy, self-efficacy and outcome expectations for exercise, perceived barriers, or past exercise history were associated with aerobic exercise participation in stroke rehabilitation. Methods: In this prospective cohort sub-study, adults admitted to in- or out-patient stroke rehabilitation at three urban hospitals completed validated questionnaires assessing depressive symptoms, apathy, exercise self-efficacy, outcome expectations for exercise, perceived barriers to being active, and premorbid exercise history. Participants were separated into two groups for analysis: those who completed aerobic exercise during rehabilitation and those who did not. Equivalence testing and between-group comparisons were performed. Results: Sixty-two participants were enrolled; 16 participated in aerobic exercise and 46 did not. Groups were not equivalent on any individual-level factors. Compared to non-participants, those who performed aerobic exercise had significantly higher depressive symptom scores (p=0.0025) and lower self-efficacy for exercise (p=0.0087). Non-participants demonstrated significantly higher apathy (p=0.0007). No significant differences were found for outcome expectations, perceived barriers, or exercise history. Conclusion: Depressive symptoms and lower self-efficacy did not impede aerobic exercise participation during rehabilitation. Increased apathy, however, was associated with non-participation. Findings highlight the need for individually tailored aerobic exercise prescriptions that consider motivational and affective factors to optimize engagement during stroke rehabilitation.
Roach, A.; Amow, A.; Haraksingh, R.; Archer, N.; Cyrus, E.; Evans, A. N.; Calleja, N.; Croes, R.; Forghani, I.; Bajnath, A.; Hadley, D.
Show abstract
Importance. Cancer is the second leading cause of death among patients in the Caribbean, where outcomes are associated with delayed clinical navigation to screening, diagnosis, and treatment. Artificial intelligence is increasingly used to guide patients with cancer to care, but whether these systems provide clinically actionable, facility-verified guidance for individuals in this population, and whether governance of the system is associated with the quality of that guidance, has not been evaluated. Objective. We tested whether a governed community learning platform navigates Caribbean cancer patients better than four ungoverned AI systems, and we tracked how community intelligence accumulates over time. Design, Setting, and Participants. We deployed a community learning ledger (CaribChat.ai) across ten Caribbean jurisdictions beginning March 2, 2026, and report all sessions through June 1, 2026 (N=207). An initial actively-promoted accrual period (March 2 - April 6, 2026; 168 sessions) was followed by continued organic use after active clinical promotion ceased. We then submitted the same 28 patient screening queries to ChatGPT (GPT-4o), Claude Haiku 4.5, DeepSeek-Chat, and OpenEvidence on April 5-6, 2026. Claude Haiku 4.5 powers CaribChat; testing it without governance isolates the governance effect. The platform requires no registration. Exempt under 45 CFR 46.104(d)(4)(ii). Main Outcomes and Measures. We classified 207 community sessions by thematic domain and temporal phase. We scored each of five systems on Caribbean facility citation, actionable navigation, and US-resource leakage across 28 screening queries. Results. The ledger accumulated 207 sessions - 168 during an actively-promoted accrual period (March 2 - April 6) and 39 after active clinical promotion ceased. Community engagement evolved from screening questions to active treatment navigation and diaspora engagement. CaribChat cited verified Caribbean facilities in 28/28 (100%) responses versus 10/28 (35.7%) for ChatGPT and 9/28 (32.1%) for OpenEvidence. CaribChat provided actionable navigation in 28/28 (100%) versus 2/28 (7.1%) for OpenEvidence (P<=.001). The same model scored 100% with governance and 54% without (P<=.001). DeepSeek cited US resources in 57.1% of Caribbean responses. After active clinical promotion ceased, off-codebook queries rose from 2.4% to 26.7% across phases while the governance contract continued to reject every adversarial probe - the community persisted but drifted from the cancer codebook absent clinician curation. The deployment operated within the OECS Health Strategy 2030 and CARICOM regional health frameworks, with queries originating across Caribbean jurisdictions led by Trinidad and Tobago. Conclusions and Relevance. Every ungoverned AI system we tested failed Caribbean cancer navigation. The best scored 68%. The most widely adopted physician platform scored 7%. The same foundation model scored 100% with governance and 54% without. Community intelligence accumulated from the population it serves, not published literature, is what makes health AI work in SIDS. The post-promotion decay shows the requirement is bidirectional: sustained, on-codebook engagement depends on patients and clinicians working together - community participation and active clinical curation are jointly necessary for maximum AI leverage.
Aguilar Ticona, J. P.; Ferreira-Stagliorio, A. F.; de Oliveira Costa, G. N.; Moreira, L.; Dias, A. S. B.; Costa, C.; Santos, A. O.; de Queiroz, A. A.; de Oliveira Pacheco, R. R.; Montano-Castellon, I.; Arriaga, M. B.; Netto, E. M.
Show abstract
Background The global increase in autism spectrum disorder (ASD) diagnoses is expected to substantially increase demand for long-term rehabilitation services. However, little is known about how this increase affects rehabilitation service utilization and capacity in low- and middle-income countries. Methods We conducted a retrospective longitudinal study of children receiving developmental care at a tertiary rehabilitation center in Salvador, Brazil (2017-2026). Patients were classified into Childhood Autism, Other ASD, and non-ASD diagnostic groups according to ICD-10 diagnoses. Temporal trends in admissions and patients under follow-up were analyzed using generalized additive models and segmented Poisson regression. Factors associated with follow-up duration were evaluated using multivariable Cox proportional hazards models. Results Among 2,123 eligible children, 833 (39.2%) had Childhood Autism, 462 (21.8%) had Other ASD, and 828 (39.0%) had non-ASD diagnoses. Compared with children with non-ASD diagnoses, those with Childhood Autism entered care at younger ages, were predominantly male (77.9% vs. 57.2%), attended more visits, and remained under follow-up longer (all P<0.001). Admissions of children with Childhood Autism increased by 30.2% annually before 2023 but declined thereafter (-19.6% annually; P<0.001). Despite this decline, the number of children with Childhood Autism receiving ongoing rehabilitation continued to increase, reflecting prolonged follow-up. In adjusted analyses, Childhood Autism was associated with a substantially lower hazard of reaching the last recorded follow-up visit than non-ASD diagnoses (adjusted hazard ratio, 0.35; 95% CI, 0.30-0.40; P<0.001). Conclusions The rapid increase in ASD admissions fundamentally reshaped rehabilitation service utilization. Because children with ASD remained under follow-up substantially longer than those with other developmental conditions, they accounted for an increasing share of the rehabilitation caseload, even after new admissions began to decline. These findings highlight the importance of planning rehabilitation services according to both new admissions and the cumulative demand generated by long-term follow-up.
Nahas, C.; Monfort, E.; Gandit, M.
Show abstract
Introduction: Computerized cognitive training (CCT) is a promising and innovative solution to improve the quality of life for those experiencing age-related cognitive decline. The comprehension of instructions for CCT plays a crucial role in determining technology engagement. This study delves into the relationship between the presentation modes of CCT serious games instructions, their comprehension, and the resulting acceptability among older adults (aged over 65) without any known cognitive impairments. Methodology: In a within-subjects experimental design, two types of CCT instructions were submitted to 128 older participants (mean age 71.5, 70% female): without visual cues and with visual cues. This approach was complemented by a study of the influence of self-efficacy and technology-related anxiety on the acceptability of the games. Results: Instructions without salient visual cues were more acceptable for a complex functional game. Additionally, individuals with lower confidence in their cognitive abilities were less receptive to cognitive training, except for a highly familiar game. Conclusion: The study highlights that older individuals may prefer simpler instructions for complex functional games, suggesting a preference for reduced cognitive load. It also shows the subtle role of self-efficacy in technology acceptance, except for the most familiar games, with higher cognitive self-confidence linked to greater acceptability. It emphasizes the importance of metacognition and self-efficacy in engagement when CCT involves mobilizing cognitive resources. It points the need for simple and personalized instructions to improve acceptance of CCT, and to contribute to the development of tailor-made interventions for older people.
Dang, Z.; Ren, G.; Wang, Z.; Su, W.; Ma, Y.; Li, P.; Ji, D.; Li, L.; Gao, J.
Show abstract
Background: Under the DRG/DIP payment reform, the cost structure and its driving factors for laparoscopic cholecystectomy (LC) in resource-limited plateau regions remain unclear. Methods: Based on a single-center cohort of 605 plateau LC patients from May 2020 to October 2025, natural log transformation was applied to total hospitalization costs. Pearson/Spearman correlation, multivariate linear regression (traditional clinical model vs system-driven model with year dummies), and quantile regression were used. Results: Mean hospitalization cost 8097.49+/-936.85 CNY, CV=11.6%, Gini=0.062, demonstrating high homogenization. Traditional six-variable clinical model yielded R^2=0.008 (F=0.78, P=0.587), no significant predictors. The system-driven model achieved R^2=0.143 (F=3.42, P=0.001), with year dummies as dominant predictors. The study proposes the SAO (System-Allocation-Outcome) paradigm to replace the traditional SPO framework.
Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.
Show abstract
Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.
Mandich, A.; Koirala, S.; Westen, S.; Adhikari, S.; Acharya, A.; Shrestha, A.
Show abstract
Language discordance can impede community-based research and health communication where trained interpreters are limited. Although multimodal artificial intelligence systems can provide real-time spoken translation, performance with under-resourced languages during spontaneous field interactions remains poorly characterized. We evaluated ChatGPT-4o during bidirectional English-Nepali voice translation in a community setting near Dhulikhel Hospital, Nepal. In this cross-sectional field study, 30 primarily Nepali-speaking adults were recruited by convenience sampling. ChatGPT-4o mediated conversations using standardized English questions and spontaneous Nepali responses. A bilingual Nepali-English reviewer assessed 485 translated utterances using a 3-point accuracy scale and an inductively developed framework for translation and conversational deviations. Of 485 translations, 282 (58.1%) received the highest accuracy rating, 134 (27.6%) a moderate rating, and 69 (14.2%) the lowest. Mean accuracy was higher for English-to-Nepali than Nepali-to-English translation (2.63 {+/-} 0.53 vs 2.23 {+/-} 0.86); 63 of 69 low-accuracy translations (91.3%) occurred in the Nepali-to-English direction. Among 329 deviation tags, the most frequent were distortion of intended meaning (17.1%), overly formal or unnatural phrasing (14.7%), omission (14.2%), and addition of content (11.5%). Some fluent outputs substantially altered meaning or introduced information not expressed by the speaker. ChatGPT-4o demonstrated potential for real-time English-Nepali communication but also produced errors that could alter interpretation of participant responses. Accuracy was lower and more variable for Nepali-to-English translation; however, translation direction was confounded with input type because Nepali inputs were spontaneous and English inputs standardized, limiting conclusions about directional performance. These findings support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or participant understanding. As multimodal AI evolves, performance should be reevaluated across languages, real-world conditions, and model versions, with bilingual oversight and community partnership remaining central to responsible use.