Related Experiment Video
Updated: Feb 15, 2026

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Artificial intelligence in pediatric otorhinolaryngology: Assessing readability, understandability, and actionability
Alya AlZabin1, Hisham AlMutawa2, Lulwah AlTurki3
1College of Medicine, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia.
Background:
Despite the increasing use of artificial intelligence (AI) platforms in healthcare communication, their efficacy in producing patient-facing materials for pediatric surgeries is still unexplored. The readability, understandability, and actionability of postoperative instructions generated by three AI platforms (ChatGPT, Gemini, and DeepSeek) for common pediatric otorhinolaryngology (ORL) surgeries were compared in this study.
Methods:
Postoperative instructions were generated from each AI platform using a standardized prompt. Readability was assessed using the Flesch-Kincaid Reading Ease (FKRE) and Grade Level (FKGL). Understandability and actionability were evaluated using the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P). Comparative statistical analyses and Pearson correlation coefficients were calculated.
Results:
ChatGPT showed the highest readability with FKRE 64.57 (±7.20) and the lowest FKGL 7.40 (±1.54), significantly outperforming others in FKRE (p = 0.030). The Gemini had the highest understandability (83.20% ± 1.56; p = 0.027), while the DeepSeek led in actionability (71.10% ± 3.81; p = 0.140). No AI platform met all the adequacy thresholds (FKGL <6, FKRE >60, PEMAT-P > 80%) at one time. Across the procedures, the differences in the readability, understandability, and the actionability were not statistically significant (all p > 0.5). The correlation analysis showed that there is significant inverse relationships of FKRE with FKGL (r = -0.711, p = 0.032) and FKRE with understandability (r = -0.669, p = 0.049). The actionability was not significantly correlated with any metric.
Conclusion:
Although each AI platform shown strengths in some domains, none consistently fulfilled the criteria for optimal patient education. These findings underscore the necessity for enhanced AI training utilizing pediatric and caregiver-specific data to improve the quality of postoperative guidance in pediatric ORL care.
Related Concept Videos
Understanding the Self
Intelligence
Fixed Action Patterns
Understanding Self-Concept
Understanding Deception
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...

