DualAlign: Generating Clinically Grounded Synthetic Data
Li, R.; Wang, X.; Yu, H.
Show abstract
Synthetic clinical data are increasingly important for advancing AI in healthcare, given strict privacy constraints on real-world EHRs, limited availability of annotated rare-condition data, and systemic biases in observational datasets. While large language models (LLMs) can generate fluent clinical text, producing synthetic data that is both realistic and clinically meaningful remains challenging. We introduce DualAlign, a framework that enhances statistical fidelity and clinical plausibility through dual alignment: (1) statistical alignment, which conditions generation on patient demographics and risk factors; and (2) semantic alignment, which incorporates real-world symptom trajectories to guide content generation. Using Alzheimers disease (AD) as a case study, DualAlign produces context-grounded symptom-level sentences that better reflect real-world clinical documentation. Fine-tuning an LLaMA 3.1-8B model with a combination of DualAlign-generated and human-annotated data yields sub-stantial performance gains over models trained on gold data alone or unguided synthetic baselines. While DualAlign does not fully capture longitudinal complexity, it offers a practical approach for generating clinically grounded, privacy-preserving synthetic data to support low-resource clinical text analysis.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpretable deep learning approach for extracting cognitive features from hand-drawn images of intersecting pentagons in older adults 93%
- Zero Shot Health Trajectory Prediction Using Transformer 92%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 91%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- ARCH: Large-scale Knowledge Graph via Aggregated Narrative Codified Health Records Analysis 94%
- Medication information extraction using local large language models 92%
- Causal feature selection using a knowledge graph combining structured knowledge from the biomedical literature and ontologies: a use case studying depression as a risk factor for Alzheimer's disease 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.