Back

NutriRAG: Unleashing the Power of Large Language Models for Food Identification and Classification through Retrieval Methods

Zhou, H.; Chow, L.; Harnack, L.; Panda, S.; Manoogian, E.; Li, M.; Xiao, Y.; Zhang, R.

2025-03-20 health informatics
10.1101/2025.03.19.25324268 medRxiv
Show abstract

ObjectiveThis study explores the use of advanced Natural Language Processing (NLP) techniques to enhance food classification and dietary analysis using raw text input from a diet tracking app. Materials and MethodsThe study was conducted in three stages: data collection, framework development, and application. Data were collected via the myCircadianClock app, where participants logged their meals in free-text format. Only de-identified food-related entries were used. We developed the NutriRAG framework, an NLP framework utilizing a Retrieval-Augmented Generation (RAG) approach to retrieve examples and incorporating large language models such as GPT-4 and Llama-2-70b. NutriRAG was designed to identify and classify user-recorded food items into predefined categories and analyzed dietary patterns from free-text entries in a 12-week randomized clinical trial (RCT: NCT04259632). The RCT compared three groups of obese participants: those following time-restricted eating (TRE, 8-hour eating window), caloric restriction (CR, 15% reduction), and unrestricted eating (UR). ResultsNutriRAG significantly enhanced classification accuracy and effectively identified nutritional content and analyzed dietary patterns, as noted by the retrieval-augmented GPT-4 model achieving a Micro F1 score of 82.24. Both interventions showed dietary alterations: CR participants ate fewer snacks and sugary foods, while TRE participants reduced nighttime eating. ConclusionBy using AI, NutriRAG marks a substantial advancement in food classification and dietary analysis of nutritional assessments. The findings highlight NLPs potential to personalize nutrition and manage diet-related health issues, suggesting further research to expand these models for wider use.

Published in Journal of the American Medical Informatics Association (predicted rank #6) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.