Back

Prediction of Nutritional Content in Peruvian Lunch Meals by Large Language Models: A One-Shot Evaluation

Carrillo-Larco, R. M.; Gallo Ruelas, M.; Matsuzaki, M.; Tarazona-Meza, C.

2025-10-22 nutrition
10.1101/2025.10.21.25338310 medRxiv
Show abstract

BACKGROUNDVarious artificial intelligence applications have been developed to predict the nutrient content of meals. However, none have been evaluated in the context of Peruvian cuisine, characterized by diverse ingredients and recipes across geographical regions. We assessed whether large language models (LLMs) could predict the nutritional content of Peruvian lunch meals. METHODSUsing a dataset of 510 unique lunch images extracted from a nationally representative Peruvian cookbook, we compared nutrient values from recipe data (ground truth) against predictions generated by three LLMs (Gemma-3 4B, 12B, and 27B). The LLMs were given the meal name and a photograph and prompted to produce narrative descriptions of the meal. Using only the descriptions, the same LLMs were prompted to estimate six nutrients: energy (kcal/serving), protein (g/serving), carbohydrates (g/serving), iron (mg/serving), vitamin A (g/serving), and zinc (mg/serving). Agreement proportions and errors metrics were calculated against ground truth. RESULTSThe 27B LLM achieved the highest agreement proportions across most nutrients--calories (45%), carbohydrates (31%), iron (15%), vitamin A (19%), and zinc (31%)--while the 12B model performed best for protein (70% agreement). The 27B model yielded the lowest mean absolute error (MAE) for calories (108 kcal), carbohydrates (26 g), iron (4 mg), and zinc (1 mg). The 12B LLM had the lowest MAE for protein (6 g) and vitamin A (667 g). The 4B LLM showed the poorest performance across metrics. CONCLUSIONSLLMs can generate estimates of nutrient content from narrative descriptions of Peruvian lunch meals, but current performance levels fall short of the accuracy needed for clinical or consumer-facing applications.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.