Related Experiment Video
Updated: Aug 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Exploring the potential of large language models in nutrition behavior prediction: evidence from college students
Misha M Madhavan1, Satyapriya1, Subhashree Sahu1
1ICAR - Indian Agricultural Research Institute, New Delhi, India.
Introduction:
The integration of Artificial Intelligence (AI) and Large Language Models (LLMs) in behavioral nutrition is expanding rapidly to predict the behavioral data. Recent developments in LLM-assisted analytical tools have opened new possibilities for analyzing survey datasets and exploring behavioral patterns related to nutrition.
Objectives:
The study aims to evaluate the ability of a large language model-assisted analytical workflow to predict healthy eating behavior scores using survey data collected from undergraduate students. The study also compares the predictive performance of the LLM-assisted workflow with a baseline OLS regression model and examines how prompt-based conditioning using different dataset sizes influences prediction accuracy.
Methods:
This study employed a cross-sectional design using primary survey data collected from 914 undergraduate students from agricultural universities in India between December 2024 and February 2025. The Healthy Eating Behavior (HEB) scale was used to assess behavioral outcomes. The dataset was analyzed using an LLM-assisted analytical environment (ChatGPT-4o), where the model was prompted to generate predicted HEB scores across different training-test splits. The predicted scores were compared with observed survey scores using statistical tests. Additionally, qualitative analysis examined if the model's predicted determinants of healthy eating behavior were aligned with the previous literature.
Results:
The results indicate that predictions generated using the LLM-assisted analytical workflow gradually converged toward the observed survey scores as the size of the training dataset increased. When a larger proportion of the data was used for training, the predicted scores did not differ significantly from the observed mean scores. However, predictive accuracy could be further strengthened using larger and more diverse datasets. The qualitative analysis also revealed similarity between the determinants of healthy eating behavior identified by the model and those reported in prior studies.
Conclusion:
The findings suggest that LLM-assisted analytical tools can support exploratory prediction tasks in behavioral nutrition research when sufficient training data are available.
