Back

Identifying Different Phenotypes Of Symptomatic Gastroesophageal Reflux Disease Using Artificial Intelligence

Gardner, J. D.; Triadafilopoulos, G.

2023-04-18 gastroenterology
10.1101/2023.04.14.23288596 medRxiv
Show abstract

IntroductionThe Lyon Consensus Conference proposed criteria for the clinical diagnosis of three different phenotypes of gastroesophageal reflux disease: nonerosive gastroesophageal reflux disease, Reflux Hypersensitivity, and Functional Heartburn. MethodsIn the present study, we examined the potential of ChatGPT, an artificial intelligence-based conversational large language model to describe how one can identify the different phenotypes as identified by the Lyon Consensus Conference, and to provide a diagnosis when given important clinical findings in a given patient with a particular phenotype. ResultsAlthough in our analyses ChatGPT captured correct information regarding symptoms, upper gastrointestinal endoscopy findings and response to gastric antisecretory agents when asked to describe different phenotypes, it failed, however, to return correct information on esophageal acid exposure time and the association of symptoms with esophageal reflux episodes. ChatGPT was even less effective in returning the correct diagnosis after being given specific clinical features of a particular phenotype. ConclusionsAlthough it seems likely that the ability of ChatGPT to capture information from multiple sources will improve with future use and refinement, presently it is inadequate as a standalone tool for processing information for the description or diagnosis of different clinical disease states. On the other hand, artificial intelligence might prove useful to clinicians in performing tasks that involve obtaining data from a single source such as the electronic medical record and generating a document having a standardized format such as a discharge summary.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.