Back

Who Funds Open Data Sharing? Analysis of data availability statements in biomedical publications

Lawrimore, J.; Li, C.; Moraczewski, D.; Poline, J.-B.; Thomas, A.

2026-07-20 scientific communication and education
10.64898/2026.07.17.739022 bioRxiv
Show abstract

Open data sharing is increasingly mandated by research funders, journals, and institutions, yet large-scale compliance measurement remains challenging. We analyzed 951,949 open access biomedical research articles published between January 2024 and June 2025 using a dual-source pipeline: PDF-based text extraction (MinerU) where PDFs were available and PMC XML otherwise, followed by algorithmic detection of data sharing statements (oddpub v7.2.3), enriched with funder, journal, and institutional metadata from OpenAlex. The corpus included 294,172 PDF-covered articles (30.9%) and 657,777 XML-only articles (69.1%). We found an over-all open data rate of 8.7%, rising to 11.7% among funder-linked articles (those with at least one funder identified in the metadata). Rates varied more than tenfold across the research ecosystem: leading major funders reached observed open data rates of 20-24%, while top journals reached observed rates of 70-86%, with corrected estimates as high as 92.9% (Nature Genetics) after adjusting for XML-only coverage limitations. PDF-based detection identified approximately 52% more data sharing statements than XML-based methods on the same articles. These observed rates differ markedly across funders and journals, and current overall sharing remains far below universal compliance. These patterns provide an empirical baseline against which future policy changes can be measured. An interactive dashboard at https://www.opensciencemetrics.org enables stakeholders to explore and benchmark these results.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.