Related Experiment Video
Updated: Jul 13, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Cross-model consistency of AI-generated exercise prescriptions: A repeated generation study across three large
1Data Convergence Team, Office of Hospital Information, Seoul National University Bundang Hospital, Seongnam, South Korea.
International Journal of Medical Informatics
|July 11, 2026
Summary
Large language models (LLMs) show varied consistency in exercise prescription. GPT-4.1 offers stable semantic content, while Gemini 2.5 Flash repeats outputs, impacting reliability for clinical applications.
Area of Science:
- Artificial Intelligence
- Exercise Science
- Clinical Applications
Background:
- Large language models (LLMs) are increasingly used for exercise prescription.
- Cross-model consistency in LLM outputs for exercise prescription is unexamined.
Purpose of the Study:
- Systematically compare repeated generation consistency of exercise prescription outputs across three leading LLMs.
- Evaluate differences in semantic similarity, reproducibility, FITT component classification, and safety expression.
Main Methods:
- GPT-4.1, Claude Sonnet 4.6, and Gemini 2.5 Flash generated exercise prescriptions for six clinical scenarios, each 20 times (360 total outputs).
- Analysis included semantic similarity (SBERT cosine similarity), output reproducibility, FITT component classification, and safety expression under controlled conditions (temperature=0).
Main Results:
- GPT-4.1 exhibited the highest semantic similarity (0.955) with 100% unique outputs.
- Gemini 2.5 Flash showed high similarity (0.950) but low uniqueness (27.5%), indicating repetition over consistent reasoning.
- Claude Sonnet 4.6 had lower similarity (0.903); safety expression was high across all models, limiting its differentiation value.
Conclusions:
- Model-specific output behavior is crucial for evaluating LLMs in exercise prescription.
- Applications requiring reproducible and guideline-compatible exercise outputs necessitate careful consideration of LLM generation consistency.
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Per-Unit Sequence Models
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Typical Model Studies
Fluid mechanics model studies often utilize scaled-down systems to predict fluid behavior in full-scale environments, such as river flows, dam spillways, and structures interacting with open surfaces. Maintaining Froude number similarity in river models is crucial, as it replicates surface flow features like wave patterns and velocities.