Back

Trilingual (EN/ZH-CN/JP) synthetic dataset of cerebral infarction patient-nurse bedside dialogs with metadata

Taira, K.; Yada, S.; Itaya, T.; Ogata, S.; Kiyoshige, E.; Hanada, A.

2026-01-06 nursing
10.64898/2026.01.05.25342252 medRxiv
Show abstract

We propose a large-scale synthetic dataset that correlates structured background information aligned with the actual distribution of patients with cerebral infarction, nurse characteristics, and nurse-patient dialogues across diverse scenarios. Medical dialogue corpora are scarce due to privacy and access restrictions. Even when available, they primarily focus on physician-patient interactions and offer limited metadata (clinical covariates, staff characteristics, etc.). To address this gap, this resource conditions large language models on patient covariate tables and nurse characteristics, generating multilingual, daily-structured dialogues (7 scenarios per day) using a standardized JSONL schema. Potential applications include nursing education under conditions of limited clinical practice time (supporting scenario-based therapeutic communication and formative assessment), research involving controlled experiments on nursing record formats and quality indicators, and training or benchmarking AI models when actual texts cannot be shared (subject to the selected models terms of use). All content is synthetic data and does not contain protected health information. Reproducible scripts pair patients with nurses, assign care pathways, and generate conversations on a scale.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.