Participant Flow Diagrams for Health Equity in AI Research
Ellen, J.; Matos, J.; Viola, M.; Gallifant, J.; Quion, J. M.; Celi, L. A.; Hussein, N. S. A.
Show abstract
Biases in sample creation can arise at any study phase, including initial patient recruitment, exclusion criteria, input-level exclusion and outcome-level exclusion, and often reflect the underrepresentation or exclusion of demographic groups historically disadvantaged in medical research. The use of non-representative samples to construct clinical algorithms in artificial intelligence (AI) and machine learning (ML) applications may further amplify this selection bias. Building on the "Data Cards" initiative for transparency in AI research, we advocate for the addition of a detailed participant flow diagram for AI studies, emphasizing the need to detail excluded participant demographic characteristics at every study phase. This tracking of excluded participants enhances understanding of potential algorithmic biases before their clinical implementation, and thus deserves to be detailed in any medical AI study. We include both a model for this flow diagram as well as a brief case study explaining how it could be implemented in practice. Through standardized reporting of participant flow diagrams, we can better gauge the potential inequity embedded in AI applications, facilitating more reliable and equitable clinical algorithms.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 95%
- Diversity and inclusion: A hidden additional benefit of Open Data 94%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 93%
Similar papers in this journal
Similar papers in this journal
- Clinical code sets and the problem of redundancy in code set repositories 94%
- Common misconceptions held by health researchers when interpreting linear regression assumptions, a cross-sectional study 94%
- A machine learning approach to identifying important features for achieving step thresholds in individuals with chronic stroke 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.