The Power of Open Health Data: Impact, Representation, and Knowledge Diffusion
Gorijavolu, R.; Armengol de la Hoz, M. A.; Bielick, C.; Cajas, S.; Charpignon, M.-L.; El Mir, A.; Gichoya, J. W.; Kwak, H. G.; Madapati, K.; Mattie, H.; McCullum, L.; Mwavu, R.; Nair, V.; Nakayama, L. F.; Nanyonjo, J.; Nazer, L.; Patel, M. S.; Sauer, C. M.; Celi, L. A.
Show abstract
Background Open health data repositories receive billions in public funding, yet no systematic framework exists to evaluate their downstream scholarly impact, the composition of the research communities they cultivate, or the breadth of disciplines they reach. We introduce a two-degree citation methodology to quantify knowledge diffusion from open data, normalized by funding, and apply it to four major health data repositories. Methods Using the OpenAlex bibliometric database (January-February 2026), we identified all first-degree citing publications (n = 30,049) and their second-degree citing publications (n = 485,396), defined as papers citing those first-degree publications, for MIMIC (versions I-IV; retrospective EHR data; $14.4M), UK Biobank (prospective cohort with genomics; $525.5M), OpenSAFELY (federated EHR platform; $53.7M), and All of Us (prospective national cohort with biobanking and community engagement; $2,160M). We extracted author demographics (gender via Genderize.io, institutional country income via World Bank 2024 classifications) and research topics. Chi-square tests with odds ratios assessed demographic differences across repositories. Results Funding-normalized first-degree papers per $1M ranged from 689 (MIMIC) to 1 (All of Us), though these figures reflect total program investment, which included community engagement and biobanking for prospective cohorts in addition to data-curation costs. The citation amplification ratio was consistent across these four repositories (9.3-11.5x). Author demographics differed significantly (p < 0.001): LMIC authorship ranged from 41.8% (MIMIC) to 4.3% (All of Us), while female authorship showed the opposite pattern, lowest for MIMIC (31.8%) and highest for All of Us (43.2%). Female authors were consistently underrepresented in senior (last-author) compared with first-author positions across all repositories. Differences in scope, design, and what funding covers limit direct comparisons. Conclusions Open health data generates a consistent ~10x indirect citation amplification beyond its direct users, a ratio that held across repositories spanning over two orders of magnitude in funding. The large differences in funding-normalized output partly reflect structural differences between retrospective databases and prospective cohorts. Low-cost access combined with intentional community building attracted globally diverse research communities with LMIC investigators in intellectual leadership positions, while a persistent gender gap in senior authorship across all repositories reflects disciplinary and structural inequities that data access policies alone cannot address. Future evaluations of open data investments should examine who is producing research, from where, in what positions, and whether their participation translates into locally relevant knowledge production.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Diversity and inclusion: A hidden additional benefit of Open Data 97%
- Artificial Intelligence's Contribution to Biomedical Literature Search: Revolutionizing or Complicating? 91%
- Inferring Gender from First Names: Comparing the Accuracy of Genderize, Gender API, and the gender R Package on Authors of Diverse Nationality 91%
Similar papers in this journal
- The Rise of Open Data Practices Among Bioscientists at the University of Edinburgh 95%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 94%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 94%
Similar papers in this journal
- Quantifying new threats to health and biomedical literature integrity from rapidly scaled publications and problematic research 94%
- COVID-19 L·OVE repository is highly comprehensive and can be used as a single source for COVID-19 studies 90%
- Re-use of trial data in the first 10 years of the data-sharing policy of the Annals of Internal Medicine: a survey of published studies 90%
Similar papers in this journal
- A Web-based Tool for Automatically linking Clinical Trials to their Publications 91%
- Automating literature screening and curation with applications to computational neuroscience 90%
- Beyond Metrics to Methods: A Scoping Review of Large Language Models for Detection of Social Drivers of Health in Clinical Notes 90%
Similar papers in this journal
- Publishing at any cost: a cross-sectional study of the amount that medical researchers spend on open-access publishing each year 95%
- A systematic examination of preprint platforms for use in the medical and biomedical sciences setting 94%
- NIH Funding of COVID-19 Research in 2020: a Cross Sectional Study 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.