Back

Dietary pattern and diversity analysis using 'DietR' package in R

Sadohara, R.; Jacobs, D.; Pereira, M. A.; Johnson, A. J.

2023-07-12 nutrition
10.1101/2023.07.07.23292390 medRxiv
Show abstract

There are scarce resources available for analyzing 24-hour dietary records. Here we introduce DietR, a set of functions written in R for the analysis of 24-hour dietary recall or records data, collected with either the Automated Self-Administered 24-hour (ASA24) dietary assessment tool or two-day data from the National Health and Nutrition Examination Survey (NHANES). The R functions are intended for food and nutrition researchers who are not computational experts. DietR provides users with functions to (1) clean dietary data; (2) analyze 24-hour dietary intakes in relation to other study-specific metadata variables; (3) visualize percentages of calorie intake from macronutrients; (4) perform principal component analysis (PCA) or k-means to group participants by similar dietary patterns; (5) generate foodtrees based on the hierarchical information of food items consumed; (6) perform principal coordinate analysis (PCoA) taking food classification information into account; (7) and calculate diversity metrics for overall diet and specific food groups. DietR includes a set of tutorials available on a website (https://computational-nutrition-lab.github.io/DietR/), which are designed to be self-paced study materials. DietR enables users to visualize dietary data and conduct data-driven dietary pattern analyses using R to answer research questions regarding diet. As a demonstration of DietR, we applied DietR to a set of created 24-hour dietary records data to demonstrate the basic functions of the package. We also applied DietR to a subset of 24-hour recall data from NHANES to demonstrate analyses using dietary diversity metrics. We present the results of this example NHANES analysis comparing legume diversity with waist circumference.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.