Stage-wise algorithmic bias, its reporting, and relation to classical systematic review biases in AI-based automated screening in health sciences: A structured literature review
Pardal-Refoyo, J. L.; Pardal-Pelaez, B.
Show abstract
IntroductionAlgorithmic bias in systematic reviews that use automatic screening is a major challenge in the application of AI in health sciences. This article presents preliminary findings from the project titled "Identification, Reporting, and Mitigation of Algorithmic Bias in Systematic Reviews with AI-Assisted Screening: Systematic Review and Development of a Checklist for its Evaluation" registered in PROSPERO with the registration number CRD420251036600 (https://www.crd.york.ac.uk/PROSPERO/view/CRD420251036600). The results presented here are preliminary and part of ongoing work. ObjectiveTo synthesize knowledge about the taxonomies of algorithmic bias, reporting, relationships with classical biases, and use of visualizations in AI-supported systematic reviews in health sciences. MethodsA specific literature review was conducted, focusing on systematic reviews, conceptual frameworks, and reporting standards for bias in AI in healthcare, as well as studies cataloguing detection and mitigation strategies, with an emphasis on taxonomies, transparency practices, and visual/illustrative tools. ResultsA mature body of work describes stage-based taxonomies and mitigation methods for algorithmic bias in general clinical AI. Common improvements in reporting and transparency (e.g. CONSORT-AI, SPIRIT-AI) are described. However, there is a notable absence of direct application to AI-automated screening of systematic reviews or empirical analyses of the interactions of biases with classical biases at the review level. Visualization techniques, such as bias heatmaps and pipe diagrams, are available, but have not been adapted to review workflows. ConclusionsThere are fundamental methodologies to identify and mitigate algorithmic bias in AI in health, but significant gaps remain in the understanding and operationalization of these frameworks within AI-assisted systematic reviews. Future research should address this translational gap to ensure transparency, fairness, and methodological rigor in the synthesis of evidence.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Fast and frugal decision tree for the rapid critical appraisal of systematic reviews 96%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 96%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 96%
Similar papers in this journal
- Completeness of reporting of clinical prediction models developed using supervised machine learning: A systematic review 94%
- Creating an Indexing Scheme for Case Series Articles 94%
- Does advance contact with research participants increase response to questionnaires: A Systematic Review and meta-Analysis 94%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 96%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 95%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 94%
Similar papers in this journal
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 96%
- Agreeability testing of AMSTAR-PF, a tool for quality appraisal of systematic reviews of prognostic factor studies 96%
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 95%
Similar papers in this journal
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
- Diversity and inclusion: A hidden additional benefit of Open Data 94%
- The NASSS (Non-Adoption, Abandonment, Scale-Up, Spread and Sustainability) framework use over time: A scoping review 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.