Temporal Drift in the Semantic Meaning of Pediatric Anxiety Terms in Electronic Healthcare Records
Tschida, J.; Chandra Shekar, M.; Hanson, H. A.; Goethert, I.; Bhatnagar, S.; Santel, D.; Pestian, J.; Stawn, J. R.; Glauser, T.; Kapadia, A. J.; Agasthya, G. A.
Show abstract
AbstractO_ST_ABSObjectiveC_ST_ABSTo identify and measure semantic drift (i.e., the change in semantic meaning over time) in expert-provided anxiety-related (AR) terminology and compare it to other common electronic health record (EHR) vocabulary in longitudinal clinical notes. MethodsComputational methods were used to investigate semantic drift in a pediatric clinical note corpus from 2009 to 2022. First, we measured the semantic drift of a word using the similarity of temporal word embeddings. Second, we analyzed how a words contextual meaning evolved over successive years by examining its nearest neighbors. Third, we investigated the Laws of Semantic Change to measure frequency and polysemy. Words were categorized as AR or common EHR vocabulary. Results98% of the AR terminology maintained a cosine similarity score of 0.00 - 0.50; at least 90% of common EHR vocabulary maintained a cosine similarity score of 0.00 - 0.25. Laws of Semantic Change indicated that frequently occurring vocabulary words remained contextually stable (Frequency Coefficient = 0.04); however, words with multiple meanings, such as abbreviations, did not show the same stability (Polysemy Coefficient = 0.630). The semantic change over time within the AR terminology was slower on average than the semantic change within the common EHR vocabulary (Type Coefficient = -0.179); this was further validated by interacting the year and Type (Coef = -0.09 - -0.523). ConclusionsThe semantic meaning of anxiety terms remains stable within our dataset, indicating slower overall semantic drift compared to common EHR vocabulary. However, failure to capture nuanced changes may impact the accuracy and reliability of clinical decision support systems over time.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 94%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 93%
Similar papers in this journal
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 94%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 93%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.