Related Experiment Video
Updated: Aug 14, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation of Large Language Models in Generating Physical Exercise Rehabilitation Programs for Musculoskeletal
Yu Fu1,2, Hairui Li3, Mingke You1,2
1Sports Medicine Center, West China Hospital, Sichuan University, Chengdu 610041, China.
Healthcare (Basel, Switzerland)
|August 13, 2026
Summary
Large language models (LLMs) can create moderate-to-high-quality physical exercise rehabilitation plans for musculoskeletal disorders. However, readability and supporting materials need improvement for optimal patient use.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Musculoskeletal Rehabilitation
Background:
- Artificial intelligence (AI) and large language models (LLMs) are rapidly advancing in medicine.
- Their application in developing physical exercise rehabilitation programs and understanding musculoskeletal (MSK) disorders is not well-established.
- This study investigates the quality and readability of LLM-generated content for MSK disorder patient consultations.
Purpose of the Study:
- To evaluate the quality of LLM-generated responses for physical exercise rehabilitation programs in musculoskeletal disorders.
- To assess the readability of these LLM-generated responses across different clinical scenarios.
- To explore the potential of LLMs as supportive tools in orthopedic patient care.
Main Methods:
- Fifty patients with musculoskeletal disorders participated.
- Frequently asked questions were extracted from Google searches related to MSK disorders.
- Four LLMs generated responses to simulated clinical consultation questions.
- Response quality was assessed using the DISCERN instrument by orthopedic specialists and therapists.
- Readability was evaluated using six validated indices.
Main Results:
- LLM-generated rehabilitation programs showed consistent adherence to query requirements.
- Response quality, assessed by DISCERN scores, ranged from poor (26) to excellent (68), with a mean of 55.60.
- Readability scores indicated that most LLM responses exceeded recommended reading levels.
- Interrater agreement for quality assessment was moderate (ICC=0.68).
Conclusions:
- LLMs can produce moderate-to-high-quality physical exercise rehabilitation recommendations for MSK disorders.
- Current LLM outputs may have limitations due to suboptimal readability and insufficient supporting materials.
- With physician oversight and improvements in readability and resources, LLMs show significant promise as assistive tools in orthopedics.
