Latent structure in EHR data: reconstruction of diabetes markers with sparse NMF
Elhussein, A.; Hripcsak, G.
Show abstract
The dimensionality of electronic health record (EHR) data continues to grow as more clinical variables are recorded, often resulting in redundancy, sparsity, and analytical intractability. In this study, we apply non-negative matrix factorization (NMF) to a high-dimensional laboratory dataset of patients with type II diabetes to estimate the minimum latent dimensionality required to preserve clinically meaningful information. Using both within-patient imputation and across-patient generalization tasks, we evaluate the ability of the learned representations to reconstruct two key clinical lab values: blood glucose and HbA1c. Our findings show that clinically acceptable accuracy can be achieved with a dimensionality reduction of up to 80% and a dimensionality of 230 to 300, supporting the presence of a compact, low-dimensional latent structure underlying high-dimensional clinical data.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 95%
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 95%
- Graph-Based Clinical Recommender: Predicting Specialists Procedure Orders using Graph Representation Learning 93%
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 92%
- Zero Shot Health Trajectory Prediction Using Transformer 92%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 92%
Similar papers in this journal
- Leveraging Large Language Models to Analyze Continuous Glucose Monitoring Data: A Case Study 92%
- Emergency department admissions during COVID-19: explainable machine learning to characterise data drift and detect emergent health risks 92%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.