Generation of Synthetic Data in Health Surveys Using Large Language Models
Villarreal-Zegarra, D.; Bellido-Boza, L.
Show abstract
BackgroundGenerating synthetic data using artificial intelligence, such as large language models (LLMs), is a useful strategy in public health because it can reduce time and costs, expand access to data, and facilitate information sharing without compromising confidentiality. ObjectiveTo evaluate the consistency and psychometric plausibility of synthetic data generated by an LLM to simulate the responses of survey participants (user personas) in a national health survey in Peru. MethodsWe conducted a cross-sectional study based on the National Health Satisfaction Survey (ENSUSALUD 2016) of ambulatory health service users. We used the GPT-OSS-20B model to generate synthetic responses in Spanish, conditioned on narrative profiles derived from sociodemographic and clinical variables. We evaluated consistency between responses and profile characteristics (sex, age, and comorbidities) using performance metrics (accuracy, precision, recall, F1 score, and AUC). We compared distributions between real and synthetic data using t-tests and chi-square tests. For latent variables, we conducted confirmatory factor analyses of the PHQ-9, PHQ-8, and GAD-7 (WLSMV; polychoric matrices) and estimated internal consistency ( and {omega}). We examined normality (Jarque-Bera test) and stability through correlations between real measures (PHQ-2 and EQ-5D) and synthetic measures (PHQ-2, PHQ-8, PHQ-9, GAD-2, and GAD-7). ResultsThe model showed strong concordance with the profile for sex, age, and chronic disease status, with metrics close to 1 for most variables; overall consistency was high in the vast majority of cases. The synthetic PHQ-9, PHQ-8, and GAD-7 instruments showed optimal factor fit and high internal consistency. Synthetic measures were positively and significantly correlated with the real PHQ-2 and negatively correlated with EQ-5D, with moderate to high correlations, particularly for PHQ-8/PHQ-9 and GAD-7. ConclusionsAn LLM can generate plausible synthetic data for health surveys when its output is conditioned on user personas, preserving high coherence with demographic and clinical characteristics and maintaining adequate psychometric properties in depression and anxiety scales. However, relevant deviations were identified (e.g., overestimation of obesity, unexpected distributions in some variables, and missing values in a sensitive item), which supports the need for rigorous validation and bias control before using these data for inferential purposes or public policy.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Suitability of just-in-time adaptive intervention in post-COVID-19-related symptoms: A systematic scoping review 93%
- Evaluating and mitigating unfairness in multimodal remote mental health assessments 92%
- Impact of a pilot mHealth intervention on treatment outcomes of TB patients seeking care in the private sector using Propensity Scores Matching – Evidence collated from New Delhi, India 92%
Similar papers in this journal
- Health Complexity Assessment in Primary Care: a validity and feasibility study of the INTERMED tool 94%
- Demographic and socioeconomic determinants of access to care: A subgroup disparity analysis using new equity-focused measurements 93%
- Psychosocial factors associated with mental health and quality of life during the COVID-19 pandemic among low-income urban dwellers in Peninsular Malaysia 93%
Similar papers in this journal
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 93%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 92%
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests 91%
Similar papers in this journal
- Impact of COVID-19 on Mental Health: A Longitudinal Study Using Penalized Logistic Regression 93%
- Development, validation, and usage of metrics to evaluate clinical research hypothesis quality 92%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.