Phenotyping Adolescent Endometriosis: Characterizing Symptom Heterogeneity Through Note- and Patient-Level Clustering
Cohen, R. M.; Leventhal, E.; Nukavarapu, N.; Lazarov, V.; Hanif, S.; Elovitz, M. A.; Glazer, K. B.; Ensari, I.
Show abstract
IntroductionPelvic pain (dysmenorrhea and non-menstrual) is the most common presentation of adolescent endometriosis, but symptoms vary between and within patients. Other presentations, such as gastrointestinal (GI) symptoms, are often misattributed, leading to diagnostic delays. Patients incur frequent primary and specialty care visits, generating multiple and diverse clinical notes. These offer insights into disease trajectory and symptom heterogeneity, which can be rigorously investigated using clustering methods. This study aims to 1) evaluate phenotypes using electronic health records (EHRs) and 2) compare two clustering models (note-vs patient-level) for their ability to identify symptom patterns. MethodsWe queried the Mount Sinai Data Warehouse for clinical notes from patients aged 13-19 years with a SNOMED endometriosis diagnosis, yielding an initial sample of 7,221 notes. A randomly selected subsample was annotated with 12 disease-relevant labels, including symptoms, hormone use, and medications. The final analytic sample included 695 notes from 26 unique patients. Pelvic pain, dysmenorrhea, chronic pain, and GI symptoms were selected as model predictors based on principal component analysis. Two unsupervised machine learning (ML) methods were then applied for note-vs patient-level analyses: Partitioning Around Medoid (PAM) and Multivariate Mixture Models (MGM). ResultsThe PAM model identified K=3 clusters with average silhouette width of 0.76, indicating strong between-cluster separation. The "feature-absent" (abs) phenotype (76%) was distinct for absence of all 4 features. The "classic" phenotype (8%) exhibited pelvic pain, dysmenorrhea, and chronic pain. The "GI" phenotype (16%) was dominated by GI symptoms. The MGM identified K=2 stable patient-level clusters ({Delta} weighted model deviance = -224.93 from K=2 to 3) with a mean cluster membership probability of 0.97: A "classic" phenotype (50%), characterized by pelvic pain and chronic pain, and a "non-classic" phenotype (50%), defined by the absence of these features. PAM-based classic phenotype had significantly higher rates of hormonal intervention (78% vs 26% abs, 49% GI) and pain medication (68% vs 9% abs, 14% GI). For the patient-level, the classic phenotype also had higher average rates per person of hormonal therapy (26% vs 7%) and prescription pain medications (27% % vs 9%) (p<0.01 for all). ConclusionsBoth methods captured classic and non-classic phenotypes, with the note-level model uniquely identifying a feature-absent group. The classic phenotypes link to higher hormonal and pain intervention underscores the importance of recognizing non-classic symptoms. This study, the first to directly compare note-and patient-level clustering of EHR notes in endometriosis, demonstrates the ability to detect the less clinically recognizable phenotypes. This proof-of-concept can be applied to larger datasets to refine phenotype identification, aiding in earlier diagnosis.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sociodemographically Differential Patterns of Chronic Pain Progression Revealed by Analyzing the All of Us Research Program Data 93%
- Feasibility of Continuous Distal Body Temperature for Passive, Early Pregnancy Detection 91%
- Hospital-wide Natural Language Processing summarising the health data of 1 million patients 90%
Similar papers in this journal
- Identifying clusters of people with Multiple Long-Term Conditions using Large Language Models: a population-based study 91%
- Improving Pre-eclampsia Risk Prediction by Modeling Individualized Pregnancy Trajectories Derived from Routinely Collected Electronic Medical Record Data 91%
- Zero-shot Interpretable Phenotyping of Postpartum Hemorrhage Using Large Language Models 91%
Similar papers in this journal
- Who is pregnant? defining real-world data-based pregnancy episodes in the National COVID Cohort Collaborative (N3C) 90%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 90%
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 89%
Similar papers in this journal
- Subtyping of common complex diseases and disorders by integrating heterogeneous data. Identifying clusters among women with lower urinary tract symptoms in the LURN study 94%
- Measurement of changes to the menstrual cycle: A transdisciplinary systematic review evaluating measure quality and utility for clinical trials 93%
- Pre- and post-operative psychological interventions to prevent pain and fatigue after breast cancer surgery (PREVENT): a randomized controlled trial 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.