Back

A Machine-Learning-Imputed Global Atlas of Ultra-Processed Food Supply Shares and PIF-Based Burden Estimates for Non-Communicable Diseases

Huang, S.; Wang, X.

2026-08-03 endocrinology
10.64898/2026.08.01.26359471 medRxiv
Show abstract

No open atlas of ultra-processed food supply exists with documented predictive validity across countries, and the global non-communicable disease burden linked to such foods, estimated under counterfactual exposure scenarios with transparently decomposed uncertainty, has not been quantified within a single framework. We built a machine-learning-imputed atlas of the share of dietary energy from ultra-processed food for 174 countries over 2010 to 2023, mapping 44 published national estimates onto food supply structure via Random Forest regression with leave-one-country-out validation. Five linked evaluations follow: an ecological phenome-wide association study across 27 non-communicable disease outcomes; a potential impact fraction estimation with time-varying exposure and a five-layer sensitivity decomposition; a 60-year generational evaluation of processed macro-ingredient supply; a synthetic-control assessment of sugar-sweetened beverage taxation across 16 countries; and a cross-level comparison of ecological and individual-level effect magnitudes using nationally representative survey data. The supply-side estimate yields a range of 1.6 to 9.5 million disability-adjusted life years in 2021. A credibility-discounted figure places the burden at roughly 3.6 million. Burden growth from 2010 to 2021 was entirely denominator-driven: population ageing and disease prevalence expansion supplied 103% of the increase, while changing supply contributed minus 3 percent. This holds consistently with a 20-year generation-lag between dietary-structure change and population obesity. Leave-region-out cross-validation returns zero generalisability for Latin America and the Caribbean. The denominator-driven pattern does not imply ultra-processed food is harmless; it indicates the processed-food environment was structurally established in most countries by 2010. The atlas, sensitivity framework, and all code are released as public-health infrastructure.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.