Vision-Language Models for Image-Based Dietary Assessment: A Benchmark of Accuracy, Cost, and Prompt Strategies Across Ten Models
Sterling, S.; Berube, L. T.; Glenn, A. J.; Shaukat, A.; Barua, S.; Grams, M. E.; Tsirigos, A.
Show abstract
BackgroundDietary assessment is the cornerstone of clinical management and research studies evaluating diet and health. Traditional methods such as food diaries and 24-hour recalls can be burdensome, prone to recall bias, and difficult to adhere to. Image-based dietary assessment using vision-language models (VLMs) offers a potential solution. ObjectiveOur goal was to benchmark state-of-the-art VLMs for automated food recognition, weight estimation, and calorie estimation using Googles Nutrition5k dataset. MethodsWe evaluated 3,229 food images using ten approaches: proprietary VLMs (Gemini 2.0 Flash, 2.5 Flash, 3.0 Flash, and 3.1 Flash Lite; GPT- 4o, GPT-4o-mini, and GPT-5 Mini; and Claude Haiku 4.5), an open-source VLM (Qwen2-VL-7B), and a commercial food recognition API (FatSecret). We assessed calorie and weight estimation using Lins Concordance Correlation Coefficient (CCC) and component detection using Jaccard similarity. ResultsGemini 3.0 Flash achieved the best calorie estimation (CCC 0.767, MAE 80.7 kcal), while Gemini 3.1 Flash Lite offered very comparable accuracy (CCC 0.754) with the highest ingredient recognition (Jaccard 0.655) at the lowest cost among top-performing models ($0.59/1K images). Among earlier-generation models, Gemini 2.0 Flash remained competitive (CCC 0.742, Jaccard 0.621) at a fraction of the cost ($0.10/1K images). A human validation study in which four annotators reviewed 440 images revealed systematic omissions in the original Nutrition5k labels. After correction, the extrapolated ingredient-overlap score for Gemini 2.0 Flash increased from 0.62 to an estimated 0.82, suggesting that raw Jaccard scores substantially underestimate true model performance. ConclusionsCurrent VLMs can perform automated dietary assessment with reasonable accuracy from single overhead photographs. Our results inform model selection for dietary assessment applications and highlight remaining challenges in calorie estimation and component detection for complex, multi-item meals.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A methodological framework for deriving the German food-based dietary guidelines 2024: food groups, nutrient goals, and objective functions 93%
- Optimization of nutritional strategies using a mechanistic computational model in prediabetes: Application to the J-DOIT1 study data 91%
- Deep Learning Classification of Lipid Droplets in Quantitative Phase Images 90%
Similar papers in this journal
- Multi-Dimensional Machine Learning Approaches for Fruit Shape Recognition and Phenotyping in Strawberry 91%
- ChronoRoot: High-throughput phenotyping by deep segmentation networks reveals novel temporal parameters of plant root system architecture 90%
- ShinyLearner: A containerized benchmarking tool for machine-learning classification of tabular data 90%
Similar papers in this journal
- Healthier decisions in the cued-attribute food choice paradigm have high test-retest reliability across five task repetitions over 14 days 91%
- Fitness tracking reveals task-specific associations between memory, mental health, and physical activity 91%
- On evaluation metrics for medical applications of artificial intelligence 90%
Similar papers in this journal
- Assessing generalizability of an AI-based visual test for cervical cancer screening 91%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 89%
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 89%
Similar papers in this journal
- Bayesian Structural Time Series for Biomedical Sensor Data: A Flexible Modeling Framework for Evaluating Interventions 90%
- Repeated measures ASCA+ for analysis of longitudinal intervention studies with multivariate outcome data 88%
- How well do rudimentary plasticity rules predict adult visual object learning? 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.