BOLD: Blood-gas and Oximetry Linked Dataset - Open Source Research
Matos, J.; Struja, T.; Gallifant, J.; Nakayama, L. F.; Charpignon, M.-L.; Liu, X.; Economou-Zavlanos, N.; Cardoso, J. S.; Johnson, K. S.; Bhavsar, N.; Gichoya, J. W.; Celi, L. A.; Wong, A.-K. I.
Show abstract
Pulse oximeters measure peripheral arterial oxygen saturation (SpO2) noninvasively, while the gold standard (SaO2) involves arterial blood gas measurement. There are known racial and ethnic disparities in their performance. BOLD is a new comprehensive dataset that aims to underscore the importance of addressing biases in pulse oximetry accuracy, which disproportionately affect darker-skinned patients. The dataset was created by harmonizing three Electronic Health Record databases (MIMIC-III, MIMIC-IV, eICU-CRD) comprising Intensive Care Unit stays of US patients. Paired SpO2 and SaO2 measurements were time-aligned and combined with various other sociodemographic and parameters to provide a detailed representation of each patient. BOLD includes 49,099 paired measurements, within a 5-minute window and with oxygen saturation levels between 70-100%. Minority racial and ethnic groups account for [~]25% of the data - a proportion seldom achieved in previous studies. The codebase is publicly available. Given the prevalent use of pulse oximeters in the hospital and at home, we hope that BOLD will be leveraged to develop debiasing algorithms that can result in more equitable healthcare solutions.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical prediction rule for SARS-CoV-2 infection from 116 U.S. emergency departments 94%
- Hemodynamic profiles by non-invasive monitoring of cardiac index and vascular tone in acute heart failure patients in the emergency department: external validation and clinical outcomes 94%
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 94%
Similar papers in this journal
- Development and validation of automated computer aided-risk score for predicting the risk of in-hospital mortality using first electronically recorded blood test results and vital signs for COVID-19 hospital admissions: a retrospective development and validation study 95%
- Performance of digital Early Warning Score (NEWS2) in a cardiac specialist setting: retrospective cohort study 94%
- Protocol for the development and validation of a machine-learning based tool for predicting the risk of hypertriglyceridemia in critically-ill patients receiving propofol sedation 93%
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 95%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 94%
Similar papers in this journal
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 95%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 95%
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 94%
Similar papers in this journal
- Utility of skin tone on pulse oximetry in critically ill patients: a prospective cohort study 94%
- ABCDEF Bundle Implementation: The influence of access to bundle-enhancing supplies and equipment 92%
- A Multicenter Evaluation of Blood Purification with Seraph 100 Microbind Affinity Blood Filter for the Treatment of Severe COVID-19: A Preliminary Report 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.