Back

Creating a General-Purpose Generative Model for Healthcare Data based on Multiple Clinical Studies

Maruyama, H.; Bito, K.; Saito, Y.; Hibi, M.; Katada, S.; Kawakami, A.; Oono, K.; Charoenphakdee, N.; Gao, Z.; Igata, H.; Yoshikawa, M.; Ota, Y.; Okui, H.; Akita, K.; Yamaguchi, S.; Sugawara, Y.; Maeda, S.-i.

2025-01-25 health informatics
10.1101/2025.01.23.25320504 medRxiv
Show abstract

Data for healthcare applications are typically customized for specific purposes but are often difficult to access due to high costs and privacy concerns. Rather than prepare separate datasets for individual applications, we propose a novel approach: building a general-purpose generative model applicable to virtually any type of healthcare application. This generative model encompasses a broad range of human attributes, including age, sex, anthropometric measurements, blood components, physical performance metrics, and numerous healthcare-related questionnaire responses. To achieve this goal, we integrated the results of multiple clinical studies into a unified training dataset and developed a generative model to replicate its characteristics. The model can estimate missing attribute values from known attribute values and generate synthetic datasets for various applications. Our analysis confirmed that the model captures key statistical properties of the training dataset, including univariate distributions and bivariate relationships. We demonstrate the models practical utility through multiple real-world applications, illustrating its potential impact on predictive, preventive, and personalized medicine.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.