In-Context Learning with Large Language Models for Scalable Glycemic Index Assignment to Food Composition Databases: Development, Validation, and Reproducibility
Della Corte, K. A.; Ebbert, J. L.; Brand-Miller, J.; Atkinson, F.; Della Corte, D.
Show abstract
Assigning glycemic index (GI) values to food composition databases is a critical bottleneck in nutritional epidemiology. We developed an in-context learning approach using large language models (LLMs), in which a structured knowledge system (termed a skill) loads GI reference databases ([~]11,000 entries), expert decision rules, and error-correction heuristics into the models context window ([~]300,000 tokens). The LLM performs GI assignments without scripted logic, functioning simultaneously as a semantic matching engine, numerical reasoning system, and expert curator. We validated this approach in two experiments. In Validation Study 1, the skill predicted the expert-curated US National GI Database (9,428 foods) using only European reference data, achieving within {+/-}10 agreement of 73.7% without manual review - compared with 31.3% retention of previously published cosine-similarity approach. In Validation Study 2, the skill was augmented with US GIDB and applied to 1,157 European food descriptions classified using the EFSA FoodEx2 system, achieving ICC = 0.79 with the expert (weighted {kappa} = 0.65; triplicate ICC = 0.88). We then applied the skill prospectively to extend US dietary GI and GL surveillance to two additional NHANES cycles (2019-2023), identifying a continued decline in energy-adjusted glycemic load. Reproducibility was assessed through triplicate runs (temperature = 0, pinned model version). The skill architecture is described in sufficient detail to inform future applications of in-context learning for nutritional database construction. STATEMENT OF SIGNIFICANCEThis paper introduces a fundamentally new approach to glycemic index (GI) database construction. Rather than using programmatic text-matching algorithms followed by extensive manual curation, we demonstrate that a large language model (LLM), when loaded with the complete GI reference literature and formalized expert decision rules, can perform one-shot GI assignments at accuracy levels comparable to human expert ratings (ICC = 0.79 with expert, weighted {kappa} = 0.65 for GI category agreement). The approach is validated across two independent food databases spanning US and European food supplies. The method reduces the time required to assign GI values to a new national food database from months of expert labor to hours of computation, while maintaining reproducibility through a structured, versionable skill architecture. This has immediate practical implications for enabling GI-based dietary surveillance and epidemiologic research in countries that currently lack GI databases or need to update existing databases.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Dietaryindex: A User-Friendly and Versatile R Package for Standardizing Dietary Pattern Analysis in Epidemiological and Clinical Studies 95%
- Racial/Ethnic Heterogeneity in Diet of Low-income Adult Women in the United States: Results from National Health and Nutrition Examination Surveys 2011-2018 92%
- Adjustment for energy intake in nutritional research: a causal inference perspective 92%
Similar papers in this journal
- Fecal metagenomics to identify biomarkers of food intake in healthy adults: Findings from randomized, controlled, nutrition trials 95%
- Assessing sustainable and healthy diets in large-scale surveys: validity and applicability of a dietary index based on a brief food group propensity questionnaire representing the EAT-Lancet planetary health diet 94%
- Development and Validation of a Nutrient Profiling Model for Shopping Baskets: The Grocery Basket Score (GBS) 92%
Similar papers in this journal
- A novel web-based 24-hour dietary recall tool in line with the Nova food processing classification: description and evaluation 94%
- Avoidable burden of stomach cancer and potential gains in healthy life years from gradual reductions in salt consumption in Vietnam, 2019 to 2030: a modelling study 91%
- Demographic, spatial, and temporal dietary intake patterns among 526,774 23andMe research participants 91%
Similar papers in this journal
- A methodological framework for deriving the German food-based dietary guidelines 2024: food groups, nutrient goals, and objective functions 93%
- Estimating the dietary and health impact of implementing mandatory front-of-package nutrient disclosures in the US: a policy scenario modeling analysis 92%
- Modeling the microbial contribution to human Energy Balance using the Digestion, Absorption, and Microbial Metabolism (DAMM) model 92%
Similar papers in this journal
- Application of n-of-1 clinical trials in personalized nutrition research: a trial protocol for Westlake N-of-1 Trials for Macronutrient Intake (WE-MACNUTR) 92%
- How do the indices based on the EAT-Lancet recommendations measure adherence to healthy and sustainable diets? A comparison of measurement performance in adults from a French national survey 91%
- Interpretable machine learning framework reveals novel gut microbiome features in predicting type 2 diabetes 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.