Assessing DEI Bias in Gene Expression Omnibus (GEO) Datasets Based on Gender and Ethnicity
Gondal, M. N.
Show abstract
The Gene Expression Omnibus (GEO) is a widely used repository for gene expression data, but its datasets may be subject to biases in terms of diversity, equity, and inclusion. This study aims to assess DEI-related biases in GEO datasets specifically focusing on gender and ethnicity representation. We curated a subset of GEO datasets by applying filters for organism type, sample size, and DEI-related metadata (gender and ethnicity). Following rigorous data extraction and cleaning, we analyzed 211 datasets, evaluating gender balance and ethnicity/race distribution using quantitative metrics. For gender representation, we calculated a gender ratio for each dataset, with a ratio closer to 1 indicating balanced representation. Ethnicity/race representation was assessed using Chi-square goodness-of-fit tests to identify disparities in ethnic distribution. We ranked datasets based on these DEI criteria to identify those with the highest and lowest bias. Our findings revealed that while many datasets displayed balanced gender ratios, significant bias in ethnicity representation was observed, with a predominance of White and African American participants. Additionally, we observed some variation in DEI metrics for datasets published before and after 2015, suggesting a shift in gender recruitment practices. These results highlight the need for more inclusive data collection and reporting in genomic research. We emphasize the importance of incorporating DEI criteria into data curation and encourage future efforts to standardize DEI data reporting across repositories to mitigate bias and improve the generalizability of genomic research.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- All of gene expression (AOE): an integrated index for public gene expression databases 93%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 93%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 93%
Similar papers in this journal
- Quantitative monitoring of nucleotide sequence data from genetic resources in context of their citation in the scientific literature 93%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 93%
- FAIR Data Station for Lightweight Metadata Management & Validation of Omics Studies 93%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 93%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 92%
- Discovering Key Transcriptomic Regulators in Pancreatic Ductal Adenocarcinoma using Dirichlet Process Gaussian Mixture Model 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.