A Simulation Study to Advance Human-Centred Artificial Intelligence via Digital Citizen Science: Can Large Language Models Transform Current Approaches to Missing Data Imputation?
Patel, J.; Banik, K.; Tolulope Ibrahim, S.; Katapally, T. R.
Show abstract
BackgroundMissing data is a persistent challenge in digital health research, and traditional approaches like Multiple Imputation by Chained Equations (MICE) may not capture complex patterns. While large language models (LLMs) could offer a viable alternative, their use in this context remains understudied. Moreover, a critical gap remains in embedding human-centred artificial intelligence (AI) approaches that integrate equity, transparency, and stakeholder participation. Digital citizen science, which leverages citizen-owned devices for ethical, participatory big data collection, offers a foundation to advance such approaches in digital health. ObjectiveTo evaluate and compare the imputation accuracy of MICE with the OpenAI o3 model for categorical variables in a simulated digital health dataset under different missingness mechanisms and levels, while situating this evaluation within the broader vision of human-centred AI enabled by digital citizen science. MethodsA complete digital health dataset collected through a digital citizen science platform was used to simulate missingness under Missing at Random (MAR) and Missing Completely at Random (MCAR) at 10%, 25%, and 50%. MICE used logistic regression with five imputations and ten iterations per chain. For the o3 model, structured prompts were generated for each missing entry using all available non-missing variables from the same record. Both methods were evaluated on each simulated dataset using classification accuracy and a closeness metric representing similarity to the original data. Statistical differences were tested with a two-sample Z-test, and misclassification patterns were examined by variable type and category frequency. ResultsUnder MAR conditions, MICE and o3 performed similarly with an average accuracy of 0.60 and 0.59, and closeness metrics of 0.83 and 0.85, respectively. Under MCAR, both methods achieved 0.59 accuracy, with closeness metrics of 0.84 and 0.85. No statistically significant differences were found across conditions (all p > 0.05). ConclusionWhile MICE remains preferred for continuous data, the o3 model shows promise as a complementary tool for categorical imputation in smaller datasets. Beyond methodological comparability, this study demonstrates how digital citizen science can serve as an ethical foundation for embedding human-centred AI into digital health research, positioning large language models not only as technical tools but also as vehicles for advancing equity, transparency, and participatory innovation in healthcare.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Conversational, Longitudinal, Ecological Assessment (CLEA): Exploring a new AI-driven method for qualitative data collection in a behavioural health context 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 93%
- Evaluating and mitigating unfairness in multimodal remote mental health assessments 93%
Similar papers in this journal
- Simulated Misuse of Large Language Models and Clinical Credit Systems 94%
- Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets 93%
- Can co-designed educational interventions help consumers think critically about asking ChatGPT health questions? Results from a randomised-controlled trial 93%
Similar papers in this journal
- Obtaining prevalence estimates of COVID-19: A model to inform decision-making 92%
- Augmenting Fact and Date of Death in Electronic Health Records using Internet Media Sources: A Validation Study from Two Large Healthcare Systems 91%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 91%
Similar papers in this journal
Similar papers in this journal
- Scalable information extraction from free text electronic health records using large language models 94%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 92%
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.