Back

The Open Constellation of Nutritional Values: Mapping the Health Star Rating System

Lafargue, V.; Suguem, F. N.

2025-11-05 bioinformatics
10.1101/2025.11.05.686692 bioRxiv
Show abstract

BackgroundNutritional awareness is increasing worldwide due to rising obesity rates and the widespread use of nutrition-tracking technologies. However, nutritional labels on food packaging remain difficult for many consumers to interpret. To improve clarity, public authorities have introduced simplified front-of-pack labeling systems, such as the Health Star Rating (HSR) used in Australia and New Zealand, which summarizes a products overall nutritional quality. Despite its adoption, the lack of open data on Health Star Rating labeled products limits independent research, reproducibility, and data-driven policy evaluation. MethodsWe introduce OpenHSR, the first open, FAIR-compliant dataset (Findable, Accessible, Interoperable, and Reusable) dedicated to the Health Star Rating system. The dataset includes 246 unique food products collected from Australian supermarket websites and brand databases. Using these data, we trained and compared several regression models linear, tree-based, and neural approaches to predict Health Star Rating values from nutritional attributes, thereby developing a transparent, interpretable ("white-box") Health Star Rating calculator. ResultsMachine learning models achieved strong predictive performance, with the Support Vector Regression model yielding the best overall accuracy (Mean Squared Error = 0.28). The trained models were then applied to predict Health Star Rating values for 6,554 Open Food Facts products, demonstrating scalability and generalizability of the approach. ConclusionsOpenHSR provides the first openly available, standardized dataset linking nutritional composition and Health Star Rating values. Its transparent design supports reproducible research, automated Health Star Rating estimation, and cross-dataset integration. This resource enables future studies in nutrition science, food labeling policy, and machine learning applications for public health.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.