Back

Assessing DEI Bias in Gene Expression Omnibus (GEO) Datasets Based on Gender and Ethnicity

Gondal, M. N.

2024-12-05 bioinformatics
10.1101/2024.11.29.622327 bioRxiv
Show abstract

The Gene Expression Omnibus (GEO) is a widely used repository for gene expression data, but its datasets may be subject to biases in terms of diversity, equity, and inclusion. This study aims to assess DEI-related biases in GEO datasets specifically focusing on gender and ethnicity representation. We curated a subset of GEO datasets by applying filters for organism type, sample size, and DEI-related metadata (gender and ethnicity). Following rigorous data extraction and cleaning, we analyzed 211 datasets, evaluating gender balance and ethnicity/race distribution using quantitative metrics. For gender representation, we calculated a gender ratio for each dataset, with a ratio closer to 1 indicating balanced representation. Ethnicity/race representation was assessed using Chi-square goodness-of-fit tests to identify disparities in ethnic distribution. We ranked datasets based on these DEI criteria to identify those with the highest and lowest bias. Our findings revealed that while many datasets displayed balanced gender ratios, significant bias in ethnicity representation was observed, with a predominance of White and African American participants. Additionally, we observed some variation in DEI metrics for datasets published before and after 2015, suggesting a shift in gender recruitment practices. These results highlight the need for more inclusive data collection and reporting in genomic research. We emphasize the importance of incorporating DEI criteria into data curation and encourage future efforts to standardize DEI data reporting across repositories to mitigate bias and improve the generalizability of genomic research.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.