Related Experiment Video
Updated: Aug 6, 2026

Collecting Sleep, Circadian, Fatigue, and Performance Data in Complex Operational Environments
Published on: August 8, 2019
From GPT-4 to Expert-Endorsed Athlete Guidance: A Delphi Consensus on Sleep and Jet Lag
Jacopo Vitale1,2,3, Alan McCall4,5,6,
1Department of Theoretical and Applied Sciences, eCampus University, Novedrate, Como, Italy. jacopo.vitale@uniecampus.it.
Background:
Large language models (LLMs) including GPT-4 are increasingly used to generate health information, but concerns persist about their accuracy and relevance, particularly for elite athletes.
Objective:
This study used GPT-4-generated frequently asked question (FAQ) responses on sleep and jet lag as the starting material for expert evaluation and consensus development, with the goal of producing consensus-based, athlete-specific guidance while identifying the limitations of AI-generated content.
Methods:
Between November 2024 and March 2025, n = 17 international sleep and circadian experts from the Athlete Travel & Sleep Interest Group (ATSIG) participated in a two-round Delphi process. Experts rated 20 GPT-4-generated FAQ responses (10 on sleep, 10 on jet lag) for appropriateness using a 6-point Likert scale and provided qualitative feedback. Items were revised after round 1 using inductive thematic coding. Consensus was defined as ≥ 70% of participants rating an item as appropriate (scores 5-6) and ≤ 15% as non-appropriate (scores 1-2). Statistical analyses included Wilcoxon signed-rank tests, convergence metrics and dissent detection (outlier and bipolarity analysis).
Results:
In round 1, 15 of 20 items (75%) met consensus; by round 2, 18 of 20 (90%) achieved consensus. For sleep items, 7 of 10 reached consensus in round 1 and 9 in round 2; for jet lag, 8 items reached consensus in round 1 and 9 in round 2. Sleep Q6 (sleep and injury risk) narrowly missed the consensus threshold with 64.7% agreement, while Jet Lag Q9 (melatonin and sleep aids) remained below the 70% threshold. No item showed bimodal score distributions, suggesting no polarized disagreement. Descriptive rating patterns, increased consensus and qualitative expert feedback indicated improved clarity, accuracy and athlete-specific relevance after the Delphi process, although item-level statistical comparisons did not remain significant after Bonferroni correction for multiple testing. Qualitative analysis identified common concerns: for sleep items-imprecise or misleading content (55%), lack of athlete-specific relevance (30%) and outdated evidence (11%); for jet lag-outdated evidence (36%), imprecise or misleading content (34%) and formatting issues (17%).
Conclusions:
GPT-4-derived content may serve as useful preliminary material for expert discussion but should not be used as standalone guidance. Expert evaluation improved the clarity, safety and athlete-specific relevance of most sleep and jet-lag responses, while the final outputs should be interpreted as consensus-based guidance rather than definitive proof of correctness.
Related Concept Videos
Sleep Apnea
The condition is more prevalent among...
Sleep-Wake Cycles
NREM Sleep
NREM sleep comprises four progressive stages that seamlessly merge:
Blind Procedures