Related Experiment Video
Updated: Jan 15, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation Strategies for Large Language Model-Based Models in Exercise and Health Coaching: Scoping Review
Xiangxun Lai1,2, Yue Lai3, JiaCheng Chen1
1Research and Communication Center for Exercise and Health, Xiamen University of Technology, 600 Ligong Road, Jimei District, Xiamen, Fujian Province, 361024, China, 86 15606951380.
Current evaluations of large language model (LLM)-based AI health coaches are fragmented and lack rigor. Developing standardized validation frameworks is crucial for safe and effective AI coaching in exercise and health.
Area of Science:
- Artificial Intelligence
- Health Informatics
- Exercise Science
Background:
- Large language model (LLM)-based AI coaches offer potential for personalized health and exercise interventions.
- Current evaluation methods for these AI coaches are fragmented, lacking standardized frameworks for safety and real-time feedback.
Purpose of the Study:
- To systematically map and analyze current evaluation strategies for LLM-based AI coaches in exercise and health.
- To identify strengths, limitations, and propose directions for robust, standardized validation of these AI coaching tools.
Main Methods:
- A scoping review following PRISMA-ScR guidelines was conducted, searching 6 major databases.
- Included studies explicitly reported evaluation methods for LLM-based exercise and health coaching.
- A 5-point Evaluation Rigor Score (ERS) was developed to assess methodological depth.
Main Results:
- Twenty studies were included, predominantly using proprietary LLMs like ChatGPT (75%).
- Evaluation strategies were heterogeneous, combining human ratings (80%) and automated metrics (40%).
- Methodological rigor was low, with a median ERS of 2.5/5 and 55% of studies rated as low rigor; key gaps included limited real-world data use and inconsistent reliability reporting.
Conclusions:
- The evaluation of LLM-based AI health coaches is currently fragmented and methodologically weak.
- Multidimensional validation frameworks integrating technical benchmarks and human-centered methods are needed.
- Ensuring safe, effective, and equitable deployment of AI coaches requires robust validation.
More Related Videos
Related Concept Videos
Models of Health Promotion and Illness Prevention II
The agent-host-environment model states that disease results...
Models of Health Promotion and Illness Prevention I
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Exercise and Muscle Performance
Endurance exercises
Endurance exercises involve running, swimming, or cycling, which require repetitive movements with low force output. When a person engages in endurance exercise, a few noticeable changes occur in their skeletal muscles. For instance, the number of capillaries...

