Back

Temporal Drift in the Semantic Meaning of Pediatric Anxiety Terms in Electronic Healthcare Records

Tschida, J.; Chandra Shekar, M.; Hanson, H. A.; Goethert, I.; Bhatnagar, S.; Santel, D.; Pestian, J.; Stawn, J. R.; Glauser, T.; Kapadia, A. J.; Agasthya, G. A.

2025-03-12 health informatics
10.1101/2025.03.09.25323626 medRxiv
Show abstract

AbstractO_ST_ABSObjectiveC_ST_ABSTo identify and measure semantic drift (i.e., the change in semantic meaning over time) in expert-provided anxiety-related (AR) terminology and compare it to other common electronic health record (EHR) vocabulary in longitudinal clinical notes. MethodsComputational methods were used to investigate semantic drift in a pediatric clinical note corpus from 2009 to 2022. First, we measured the semantic drift of a word using the similarity of temporal word embeddings. Second, we analyzed how a words contextual meaning evolved over successive years by examining its nearest neighbors. Third, we investigated the Laws of Semantic Change to measure frequency and polysemy. Words were categorized as AR or common EHR vocabulary. Results98% of the AR terminology maintained a cosine similarity score of 0.00 - 0.50; at least 90% of common EHR vocabulary maintained a cosine similarity score of 0.00 - 0.25. Laws of Semantic Change indicated that frequently occurring vocabulary words remained contextually stable (Frequency Coefficient = 0.04); however, words with multiple meanings, such as abbreviations, did not show the same stability (Polysemy Coefficient = 0.630). The semantic change over time within the AR terminology was slower on average than the semantic change within the common EHR vocabulary (Type Coefficient = -0.179); this was further validated by interacting the year and Type (Coef = -0.09 - -0.523). ConclusionsThe semantic meaning of anxiety terms remains stable within our dataset, indicating slower overall semantic drift compared to common EHR vocabulary. However, failure to capture nuanced changes may impact the accuracy and reliability of clinical decision support systems over time.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.