Back

Interactive Evaluation of an Adaptive-Questioning Symptom Checker Using Standardized Clinical Vignettes

Madda, P.; Kondru, J.

2025-08-24 health informatics Community evaluation
10.1101/2025.08.21.25333628 medRxiv
Show abstract

ObjectiveTo evaluate the triage performance and history-taking quality of an adaptive-questioning symptom checker (CareRoute) using an interactive protocol on standardized clinical vignettes (Semigran et al., BMJ 2015; 45 cases). MethodsEach session began with only the presenting complaint; CareRoute asked follow-up questions adaptively, and the evaluator answered concisely per the vignette. At the end of questioning, CareRoute issued a triage recommendation. We compared CareRoutes issued triage with the reference triage and computed history-taking quality from normalized features derived from each vignettes Condensed Format. History-taking quality comprised (i) elicitation coverage--the percentage of a vignettes normalized features obtained through questioning, and (ii) elicitation fraction--the proportion of surfaced normalized features (elicited or volunteered) that were obtained through questioning. Primary outcomes were triage concordance and history-taking quality; the secondary outcome was user burden (time spent answering questions). We did not evaluate possible diagnoses, though CareRoute issues them. ResultsExact 3-tier triage concordance was 88.9% (40/45; 95% CI 76.5-95.2%). Elicitation coverage had a median of 67% (IQR 60-71%), and elicitation fraction had a median of 70% (IQR 62-75%). CareRoute asked a median of 19 questions overall (IQR 16-20), with urgency-conditioned questioning: Emergency Care median 10 questions (IQR 4-14), Doctor Visit median 19 questions (IQR 18-20), Self Care median 19 questions (IQR 17-20). ConclusionsIn an interactive, vignette-constrained evaluation starting from only the presenting complaint, CareRoute achieved high 3-tier triage concordance (88.9%) with no under-triage on Emergency-reference vignettes, while eliciting most normalized features (median elicitation coverage 67%; median elicitation fraction 70%) with acceptable user burden via urgency-conditioned questioning.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.