Related Experiment Video
Updated: Oct 7, 2026

Iterative Development of an Innovative Smartphone-Based Dietary Assessment Tool: Traqq
Published on: March 19, 2021
Evaluating sports nutrition advice provided by general-purpose large language models for recreational female marathon
Hao Sun1, Xinshuo Chu2, Ziyang Li1
1School of Sports Training, Wuhan Sports University, Wuhan, China.
Background:
General-purpose large language models (LLMs) are increasingly used for health and nutrition information. We evaluated the quality and practical limitations of LLM-generated sports nutrition advice for a standardized recreational female marathon runner, focusing on numerical recommendations, safety boundaries, and female-specific considerations.
Methods:
We conducted a time-stamped cross-sectional descriptive audit of five LLM configurations using 15 consumer-style prompts covering pre-race nutrition, in-race fueling and hydration, unexpected race situations, supplements, and female-specific nutritional risks. Each prompt was repeated 10 times per configuration, yielding 750 responses. The standardized profile was a 39-year-old, 58-kg recreational female marathon runner targeting a 4-h finish during the luteal phase. Responses were evaluated for scientific accuracy, numerical appropriateness and stability, safety-boundary handling, female-specific applicability, and readability. The coding framework was calibrated on 50 responses and finalized before blinded full-corpus evaluation by two author-raters and two independent external raters.
Results:
All configurations produced structured advice, but performance and female-specific integration varied by configuration, prompt, and domain. In the directly prompted Q15 scenario, ChatGPT covered all seven indicators in all 10 runs. Gemini's Q5 recommendations ranged from 5-6 gels and 50-60 g/h carbohydrate. DeepSeek and Doubao Q5 carbohydrate recommendations ranged from 15-120 g/h and 5-145 g/h, respectively, crossing both limits of the 30-90 g/h comparison range. One Doubao response stated a 115-145 g/h target while its gel schedule provided 37.5-45 g/h. Configurations differed in response stance toward strong pre-race coffee and in maintaining the training-first recommendation for new supplements (5/10-10/10). Model-level mean Flesch Reading Ease scores ranged from 45.5 to 57.8. Raw agreement between the independent experts was 91.24% for binary judgments (κ = 0.780); agreement with final author ratings was 87.76%-89.98% (κ = 0.703-0.763); exact numeric agreement was 75.0%; indicator-specific ICCs were 0.264-0.928.
Conclusion:
General-purpose LLMs can provide accessible and structured sports nutrition information for recreational female marathon runners, but their advice remains uneven in female-specific integration, numerical consistency, and safety-boundary communication. Apparent strengths were configuration-, prompt-, and context-specific and do not support a stable overall ranking of model families. LLM outputs may therefore be useful as supplementary information but should not replace individualized guidance from a sports dietitian or medical professional.

