Back

Knowledge-Driven Online Multimodal Automated Phenotyping System

Xiong, X.; Sweet, S. M.; Liu, M.; Hong, C.; Bonzel, C.-L.; Ayakulangara Panickan, V.; Zhou, D.; Wang, L.; Costa, L.; Ho, Y.-L.; Geva, A.; Mandl, K. D.; Cheng, S.-C.; Xia, Z.; Cho, K.; Gaziano, J. M.; Liao, K. P.; Cai, T.; Cai, T.

2023-10-02 health informatics
10.1101/2023.09.29.23296239 medRxiv
Show abstract

Though electronic health record (EHR) systems are a rich repository of clinical information with large potential, the use of EHR-based phenotyping algorithms is often hindered by inaccurate diagnostic records, the presence of many irrelevant features, and the requirement for a human-labeled training set. In this paper, we describe a knowledge-driven online multimodal automated phenotyping (KOMAP) system that i) generates a list of informative features by an online narrative and codified feature search engine (ONCE) and ii) enables the training of a multimodal phenotyping algorithm based on summary data. Powered by composite knowledge from multiple EHR sources, online article corpora, and a large language model, features selected by ONCE show high concordance with the state-of-the-art AI models (GPT4 and ChatGPT) and encourage large-scale phenotyping by providing a smaller but highly relevant feature set. Validation of the KOMAP system across four healthcare centers suggests that it can generate efficient phenotyping algorithms with robust performance. Compared to other methods requiring patient-level inputs and gold-standard labels, the fully online KOMAP provides a significant opportunity to enable multi-center collaboration.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.