Back

Synthetic Validation of Pediatric Trust Instruments using Persona-Driven Large Language Models

Boone, E.; Loban, K.; Guadagno, E.; Poenaru, D.

2025-11-27 health systems and quality improvement
10.1101/2025.11.25.25340922 medRxiv
Show abstract

ObjectivesTrust is foundational to patient-physician relationships and is associated with improved care-seeking and adherence in primary care. However, validated trust instruments for pediatric emergency and surgical contexts are lacking, and traditional instrument development is slow and resource-intensive. Large language models (LLMs) could streamline the validation process by serving as scalable, systematic expert panel surrogates. Materials and MethodsWe developed four new trust assessment instruments: two for patient-families and two for physicians. Two-phase content validation was conducted using two parallel synthetic and human expert panels. Synthetic panels consisted of three persona-prompted LLMs (Claude Sonnet 4, GPT-5, Grok4). Human panels served as traditional comparators. Scale-Content Validity Index (S-CVI) and Fleiss kappa (k) acceptance thresholds were set at [≥]0.80. ResultsCombined human-synthetic expert panels revealed substantial inter-rater reliability across all instruments. Fleiss kvalues for dimensional validation were: patient-family = 0.84 (95% CI [0.72, 0.96]), physician = 0.87 (95% CI [0.72, 1.00]);contextual validation: patient-family = 0.83 (95% CI [0.73, 0.93]), physician = 0.88 (95% CI [0.80, 0.96]). All instruments exceeded S-CVI [≥]0.80 thresholds across both validation phases. DiscussionPersona-prompted LLMs demonstrated comparable validity outcomes to human experts while accelerating validation timelines from months to weeks. Future research needs to evaluate this approach across psychometric testing phases. ConclusionThis synthetic instrument validation methodology offers a scalable blueprint for healthcare measurement development, enabling faster creation of validated tools to support evidence-based patient care.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.