Back

Joining the Conversation with a Dedicated Medical Education Corpus

Ow, G. M.; Stetson, G. V.; Costello, J. A.; Artino, A. R.; Maggio, L. A.

2024-12-19 medical education
10.1101/2024.12.17.24319205 medRxiv
Show abstract

PROBLEMMedical education scholars struggle to join ongoing conversations in their field due to the lack of a dedicated medical education corpus. Without such a corpus, scholars must either search too widely across thousands of irrelevant journals, or too narrowly by relying on PubMeds Medical Subject Headings (MeSH). In our tests, MeSH missed 34% of medical education papers. APPROACHWe developed the MEC (Medical Education Corpus), the first dedicated collection of medical education papers, through a three-step process. First, using the Core-Periphery model, we created the MEJ (Medical Education Journals), a collection of three groups of journals based on participation and influence in medical education discourse: the MEJ-Core (formerly the MEJ-24, 24 journals), MEJ-Adjacent (127 journals), and MEJ-Peripheral (a theoretical group). Second, we developed and evaluated a machine learning model, the MEC Classifier, trained on 4,032 manually labeled papers to identify medical education content. Finally, we applied the MEC Classifier to extract medical education papers from the MEJ-Core and MEJ-Adjacent journals. OUTCOMESThe MEC currently contains 119,137 medical education papers from the MEJ-Core (54,927 papers) and MEJ-Adjacent journals (64,210 papers). In our evaluation using 1,358 test papers, the MEC Classifier demonstrated significantly improved sensitivity compared to MeSH (90% vs 66%, p = 0.001), even while maintaining a similar positive predictive value (82% vs 81%). NEXT STEPSThe MEC provides a focused, searchable corpus that enables medical education scholars to more easily join conversations in the field. Scholars can now rely on the MEC when reviewing literature to frame their work, and the MEC also creates opportunities for field-wide analyses and meta-research. The MEC is freely available in the Supplement and is regularly updated on MedEdMentor (mededmentor.org), where we will also incorporate community feedback to further improve and expand the corpus.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.