Diagnostic analytics of routine Clinical Competency Committee data of six cohorts in family medicine program in the UAE, utilizing Milestones, EPA, and ITE
Baynouna Alketbi, L. M.; Nagelkerke, N.; Alzarouni, A.; AlKwuiti, M.
Show abstract
In Competency-based medical education (CBME), longitudinal data is generated continuously. The judgments a Clinical Competency Committee (CCC) makes about trainee learning and performance are a valuable resource, supporting both resident and program development. Such data as well can enables the evaluation of rating quality and of CBME instruments such as Milestones and Entrustable Professional Activities (EPAs) which can help address a gap in the CBME literature, where evidence on the performance of these instruments remains limited. Objective Routinely gathered CCC data of six cohorts in a four-training ACGME-I-accredited family medicine residency in Al Ain, United Arab Emirates, was studied to describe growth trajectories, rating-system behavior, and the concurrent agreement of CBME instruments. As well as investigating the prospective predictive validity of two CBME instruments, EPA and Milestones, and the In-Training Exam (ITE). Methods The longitudinal CCC data for 80 residents across six cohorts (2019-20 to 2024-25) were assessed at up to eight time points (mid- and end-year; R1-R4). The pooled dataset included 10,458 EPA item ratings across 334 resident-time points, 5,021 Competency Milestone item ratings across 285 resident-time points, and 185 ITE scores. Five research questions were examined: growth trajectories; within- and between-resident variation and straight-lining (identical scores on assessed items at a single time point); EPA-Milestones agreement; the validity of supervisor ratings against the ITE (anchoring diagnostic, same-year correlations, prospective regressions); and EPA blueprint fidelity (the mapping of EPAs against the ACGME-I subcompetency). Al Ain trajectories were benchmarked against an international family medicine reference. Results All three instruments rose steadily across the eight timepoints. By End R4, the Milestones mean (4.00, range 3.83-4.24) matched US end-of-training norms (3.84-4.02). With regards to rating quality, pooled R1-R3 Milestones straight-lining was 2.3% (EPA 0%), below US benchmarks; between-resident discrimination was preserved (SD 0.41-0.54); and longitudinal halo was ruled out (within-domain growth-slope r = 0.61 vs across-domain r = 0.37). End R1 Overall EPA was the strongest prospective predictor of Final Competency (B = 0.96, p < 0.001) and Final ITE (B = 96.88, p = .006). Medical Knowledge ratings were independent of prior ITE scores from Mid R2 onward, and End R2 MK ratings predicted ITE 17 months later at r=0.88, confirming supervisor judgment was not anchored to test results. With regards whether individual EPAs correlate with individual Milestone subcompetencies at each timepoint, a significant EPA and Milestones correlations were negligible at End R1 (1 of 222 item-level cells significant) and converged by End R3 (36 cells), while resident-mean stepwise regressions showed the two instruments (EPA and Milestones) behaved as overlapping predictors throughout, indicating that EPAs and Milestones are complementary at the level of specific content but convergent at the level of aggregate resident judgment. Blueprint fidelity rose from 30% of cells reaching r [≥] 0.40 at End R2 to 80% at End R3 in the same cohort, indicating that apparent fidelity is materially affected by measurement timing. Conclusion By graduation, residents demonstrated substantial and progressive competency achievement across both instruments, with the majority reaching the entrustable threshold on both EPA and Milestone ratings. The rating system demonstrated disciplined assessment behavior of supervisors and both concurrent and prospective validity relative to the ITE. Overall EPA at End R1 was the strongest prospective predictor of all three terminal outcomes, final ITE score, graduating Competency Milestones, and graduating overall EPA, outperforming Milestones and baseline knowledge. Routine CCC data support an evidence-based quality assurance framework spanning rater-process diagnostics, outcome-validity diagnostics, and the asymmetric-instrument diagnostic, requiring no additional data collection beyond existing program processes.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.