Related Experiment Video
Updated: May 23, 2026

'Boden Food Plate': Novel Interactive Web-based Method for the Assessment of Dietary Intake
Published on: September 18, 2018
Evaluation of AI-generated versus registered dietitian-authored nutrition responses: a cross-sectional study
Kathryn Ayres1, Maryam Nadery2, Pia Henfridsson2
1Department of Public and Allied Health, College of Health and Human Services, Bowling Green State University, Bowling Green, OH, USA.
Background:
Artificial intelligence (AI) has emerged as a potential tool in nutrition counseling, but its performance compared with registered dietitians (RDs) remains unclear. This study aimed to compare the clinical quality, empathy, and readability of nutrition responses by a large language model (LLM, ChatGPT-4o) with those provided by licensed RDs, as assessed by RD evaluators.
Methods:
In this cross-sectional study, 100 nutrition-related questions were selected from public online forums where RDs had provided answers. Each question was paired with an AI-generated response. Licensed RDs (n=8), blinded to the source, rated responses for quality and empathy (5-point Likert scales) and overall performance (0-100). Readability was assessed using the Flesch Reading Ease Score (FRES), Flesch-Kincaid Grade Level (FKGL), syllables per word, and words per sentence. Statistical analyses included independent two-tailed t-tests, z-tests for proportions meeting a threshold for "acceptable" (≥4), Pearson correlations, and sensitivity analyses by response length.
Results:
AI-generated responses scored higher than RD-authored responses for quality (4.48±0.31 vs. 2.56±0.76; P<0.001), empathy (4.62±0.37 vs. 3.21±0.62; P<0.001), and overall performance (91.10±5.38 vs. 66.83±14.71; P<0.001). AI scores clustered at the upper end, while RD scores were more variable. Quality and empathy were not correlated for AI (r=-0.10, P=0.32) but showed a moderate positive correlation for RDs (r=0.37, P<0.001). Nearly all AI responses met the ≥4 threshold for quality (96%) and empathy (97%), compared with few RD responses (3% and 14%; P<0.001). Word count did not differ, and longer RD responses were not associated with higher ratings. RDs' responses were more readable, with higher FRES (53.5±13.5 vs. 46.2±12.5; P<0.001), and simpler vocabulary (1.60±0.1 vs. 1.73±0.1 syllables/word, P<0.001), though both groups averaged a 10th-grade level on the FKGL, exceeding Centers for Disease Control and Prevention (CDC) and National Institutes of Health (NIH) recommendations.
Conclusions:
AI-generated nutrition responses demonstrated consistently high perceived quality and empathy, independent of length, while RD-authored responses display greater variability but higher readability. These findings highlight both the promise and the limitations of LLMs in nutrition counseling, suggesting that AI may complement, but not replace, human expertise, provided that accuracy, transparency, and professional oversight are maintained.
