Early Detection of Erythropoietic Protoporphyria Using Sequential Machine Learning on Longitudinal Electronic Health Records
Ayati, A.; Onal, G.; Sur, A.; Azzam, S.; Wang, B.; Rudrapatna, V. A.
Show abstract
Objective: Erythropoietic protoporphyria (EPP) is a rare photodermatosis marked by multi-year diagnostic delays. We developed and externally validated machine learning models to identify patients with EPP earlier from longitudinal electronic health record (EHR) data and estimate undiagnosed disease burden. Materials and Methods: In a retrospective case-control study at two San Francisco health systems, an academic referral center (UCSF) and a safety-net hospital (ZSFG) we identified 74 confirmed EPP cases using combined diagnostic coding, biochemical criteria, and specialty chart review. Symptom-enriched controls were sampled at a 40:1 ratio. Longitudinal diagnoses, laboratory results, medications, procedures, and encounters preceding the outcome date were modeled with a gradient-boosting classifier (CatBoost) and a state-space sequence model (MAMBA). The best model was deployed across the UCSF population and externally validated at ZSFG without retraining. Results: On the UCSF held-out test set (n=1,865; 43 cases), MAMBA outperformed CatBoost (AUC ROC 0.91 vs 0.89; average precision 0.42 vs 0.27; precision 65% vs 20%), flagging cases a median of 229 days before documented diagnosis. Deployed across 297,967 symptom-compatible patients, it identified 310 high-risk individuals, implying a prevalence approaching genetic estimates. External validation at ZSFG showed attenuated performance (AUC ROC 0.72; average precision 0.10) while preserving early detection (median 264 days). Discussion: A sequence model integrating temporal EHR signals detected EPP months before clinical recognition, corroborating genetic evidence of substantial underdiagnosis. Cross-site attenuation reflects population and documentation differences and underscores the need for local recalibration. Conclusion: Longitudinal EHR-based machine learning can shorten EPP diagnostic delay and prioritize patients for confirmatory testing, supporting proactive rare-disease case finding.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning models predict long COVID outcomes based on baseline clinical and immunologic factors 90%
- Computational Assessment of Memory Function in Kidney Transplant Recipients and Donors 88%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 88%
Similar papers in this journal
- Diagnostic Yield of Exome Sequencing in a Diverse Pediatric and Prenatal Population is not Associated with Genetic Ancestry 89%
- Discordance between a deep learning model and clinical-grade variant pathogenicity classification in a rare disease cohort 87%
- Comprehensive reanalysis for CNVs in ES data from unsolved rare disease cases results in new diagnoses 86%
Similar papers in this journal
- Detecting Rare Diseases in Electronic Health Records Using Machine Learning and Knowledge Engineering: Case Study of Acute Hepatic Porphyria 92%
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 91%
- CohortDiagnostics: phenotype evaluation across a network of observational data sources using population-level characterization 91%
Similar papers in this journal
- Clinical Study Applying Machine Learning to Detect a Rare Disease: Results and Lessons Learned 93%
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 90%
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 89%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.