Related Experiment Video
Updated: Jul 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Cross-model consistency of AI-generated exercise prescriptions: A repeated generation study across three large
1Data Convergence Team, Office of Hospital Information, Seoul National University Bundang Hospital, Seongnam, South Korea.
Background:
Large language models (LLMs) are increasingly applied to exercise prescription, yet cross-model differences in output consistency under repeated generation conditions remain unexamined.
Objectives:
This study aimed to systematically compare repeated generation consistency of exercise prescription outputs across three widely used LLMs under identical conditions.
Methods:
GPT-4.1, Claude Sonnet 4.6, and Gemini 2.5 Flash each generated prescriptions for six clinical scenarios 20 times (360 total outputs) under temperature = 0 conditions. Outputs were analyzed across four dimensions: semantic similarity (SBERT cosine similarity), output reproducibility, FITT component classification, and safety expression.
Results:
Mean semantic similarity was highest for GPT-4.1 (0.955), followed by Gemini 2.5 Flash (0.950) and Claude Sonnet 4.6 (0.903), with significant inter-model differences confirmed (H = 458.41, p < 0.001, ε2 = 0.134). These scores reflected fundamentally different generative behaviors: GPT-4.1 produced entirely unique outputs (100%) with stable semantic content, while Gemini 2.5 Flash showed pronounced output repetition (27.5% unique outputs), indicating that its high similarity score derived from text duplication rather than consistent reasoning. Safety expression reached ceiling levels across all models (mean Safety Total: 3.93-3.99 out of 4.00), confirming its limited utility as a differentiating metric.
Conclusions:
These findings suggest that model-specific output behavior should be considered when evaluating LLMs for exercise prescription support, particularly in applications requiring reproducible and guideline-compatible outputs.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Typical Model Studies