Back

A versioned, analysis-ready archive of United States State Cancer Profiles county- and state-level estimates

Davis, S.

2026-08-27 health informatics
10.64898/2026.08.24.26361254 medRxiv
Show abstract

State Cancer Profiles (statecancerprofiles.cancer.gov), maintained by the National Cancer Institute with the Centers for Disease Control and Prevention, is a widely used source of county- and state-level cancer statistics in the United States, used for cancer-center catchment-area surveillance and for geographic studies of cancer burden, screening, and access to care. The site offers no API, no bulk download, and no archive of prior estimates: its sole machine-readable export returns one statistical stratum per HTTP request, and when the underlying data are updated the previous estimates are overwritten and become unrecoverable. This resource provides the complete national county- and state-level extract of all four State Cancer Profiles data topics (incidence, mortality, screening and risk factors, and demographics) as typed, analysis-ready files with the stratifying dimensions as columns, published under pinned, citable version DOIs on Zenodo (concept DOI 10.5281/zenodo.11098814). One version DOI is minted per distinct upstream data vintage, the set of values the site served between successive replacements. Three vintages have been captured to date; at each observed vintage boundary roughly 97% of estimate values changed, so which vintage an analysis draws on affects its results. From the 2026-08-24 release forward, cells that the upstream site suppresses are retained as typed nulls with an explicit suppression-reason column. Capture has been automated on an approximately monthly cadence since February 2025, and each future upstream revision will be preserved as a new vintage.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.