SynthCraft: an AI partner for synthetic data generation to support data access and augmentation in healthcare
Callender, T.; Boyd, A.; Davis, R.; Ruhrberg Estevez, S.; Lavista Ferres, J. M.; van der Schaar, M.
Show abstract
BackgroundAccess to high-quality data provides the foundation for biomedical research. But data access is often limited or challenging due to privacy constraints, whilst the data themselves may be unrepresentative or sparse. Synthetic data can support privacy-preserving data access, data augmentation, as well as complex analytical workflows for the development of digital twins or to evaluate the impacts of data distribution shifts. However, the use of synthetic data remains limited due to the complexity of the methods themselves and their evaluation, as well as the need for advanced programming skills. MethodsWe developed SynthCraft, a tool for AI-human collaboration to support the principled, transparent, use of state-of-the-art synthetic data generation methods. SynthCraft uses Large Language Models (LLMs) combined with a reinforcement learning-based reasoning engine to orchestrate the necessary workflow to generate synthetic data based on dynamic interaction with the user using natural language. We demonstrate the capability of SynthCraft with both tabular and genomic datasets: National Health and Nutrition Examination Survey (NHANES) and the Cancer Genome Atlas (TCGA). ResultsUsing SynthCraft, we analysed the privacy, statistical fidelity, and downstream utility of four different synthetic data generators both with and without explicit privacy-preserving designs when applied to both the NHANES and TCGA datasets. We show that how different generators perform differently - and that no single method was optimal - across varying use-cases and datasets. Furthermore, we demonstrate how SynthCraft can be used for data augmentation as part of a workflow to attempt to mitigate imbalances in the proportion of individuals from different ethnic backgrounds. ConclusionsAn LLM-based, human-in-the-loop, AI partner can support the generation of synthetic datasets. Such tools could improve the quality, reproducibility, and transparency of research methods, whilst increasing their accessibility. Research into their use across different methodological areas is warranted.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 95%
- GRouNdGAN: GRN-guided simulation of single-cell RNA-seq data using causal generative adversarial networks 94%
- scPRINT: pre-training on 50 million cells allows robust gene network predictions 94%
Similar papers in this journal
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 94%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 94%
- Understanding the robustness of vision-language models to medical image artefacts 94%
Similar papers in this journal
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 92%
- Drug-combination wide association studies of cancer 92%
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.