Evaluating and Improving the Performance and Racial Fairness of Algorithms for GFR Estimation
Zhang, L.; Richter, L. R.; Kim, T.; Hripcsak, G.
Show abstract
Data-driven clinical prediction algorithms are used widely by clinicians. Understanding what factors can impact the performance and fairness of data-driven algorithms is an important step towards achieving equitable healthcare. To investigate the impact of modeling choices on the algorithmic performance and fairness, we make use of a case study to build a prediction algorithm for estimating glomerular filtration rate (GFR) based on the patients electronic health record (EHR). We compare three distinct approaches for estimating GFR: CKD-EPI equations, epidemiological models, and EHR-based models. For epidemiological models and EHR-based models, four machine learning models of varying computational complexity (i.e., linear regression, support vector machine, random forest regression, and neural network) were compared. Performance metrics included root mean squared error (RMSE), median difference, and the proportion of GFR estimates within 30% of the measured GFR value (P30). Differential performance between non-African American and African American group was used to assess algorithmic fairness with respect to race. Our study showed that the variable race had a negligible effect on error, accuracy, and differential performance. Furthermore, including more relevant clinical features (e.g., common comorbidities of chronic kidney disease) and using more complex machine learning models, namely random forest regression, significantly lowered the estimation error of GFR. However, the difference in performance between African American and non-African American patients did not decrease, where the estimation error for African American patients remained consistently higher than non-African American patients, indicating that more objective patient characteristics should be discovered and included to improve algorithm performance.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating the kidney disease progression using a comprehensive patient profiling algorithm: A hybrid clustering approach 96%
- Structural modeling for Oxford histological classifications of immunoglobulin A nephropathy 95%
- Progression of chronic kidney disease among black patients attending a tertiary hospital in Johannesburg, South Africa 95%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 93%
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 93%
- Combining symbolic regression with the Cox proportional hazards model improves prediction of heart failure deaths 92%
Similar papers in this journal
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 92%
- Modeling physician variability to prioritize relevant medical record information 92%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 91%
Similar papers in this journal
- Causal modeling of chronic kidney disease in a participatory framework for informing the inclusion of social drivers in health algorithms 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 91%
- Learning Decision Thresholds for Risk-Stratification Models from Aggregate Clinician Behavior 90%
Similar papers in this journal
- PK-RNN-V E: A Deep Learning Model Approach to Vancomycin Therapeutic Drug Monitoring Using Electronic Health Record Data 91%
- Graph-Based Clinical Recommender: Predicting Specialists Procedure Orders using Graph Representation Learning 90%
- A scoping review of fair machine learning techniques when using real-world data 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.