Back

Analyzing open-ended questions in research: A commonly used category selection methodology

Agosto Arroyo, L. D.; Fitzmaurice, A.; Feric, Z.; Alshawabkeh, A.; Meeker, J. D.; Kaeli, D.; Velez-Vega, C. M.; Cordero, J. F.; Cardona Cordero, N. R.

2022-05-30 epidemiology
10.1101/2022.05.27.22275646 medRxiv
Show abstract

A closer examination of consumer product brands and how they are associated with levels of potential endocrine disrupting chemicals should be explored. The large number of brands available and changes in consumer preferences for certain brands makes it difficult to develop questionnaires that include all brands. Open-ended brand reporting questions are an option, but they bring challenges in identifying each brand given the multiple possibilities of variations in brand name reporting. We report a method for transforming product brand data reported as text to brand codes that allows quantitative analysis of brand use and its association with endocrine disrupting chemicals. We selected 14 consumer products to be included in our analyses. To evaluate commonly used brand selection, we used Cohens power calculations for two-sample t-tests in R (version 1.3.0). Considering a moderate effect size (Cohens d) of 0.5, each test will include the most used brand and the least used brand among the commonly used brands per product and visit. We compared how the commonly used brand selection differ per product in terms of the number of brands it selected, the total sample size and the power calculated by creating a correlation matrix and analyzing the relationship between power, commonly used brands, and brand usage. The correlation coefficient between the commonly used brand frequency of each visit approximated 0.99. From all products, fabric softener, conditioner, and lotion where the products that attained the highest power. The differences in brand use distributions per product provided an optimal environment for evaluating the performance of the commonly used brand selection methodology. It provides enough flexibility when selecting exposure groups that it could be applied to any open-ended questions, and it proves significantly useful when accounting for repeated measures.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.