Related Experiment Video
Updated: Aug 16, 2026

Using an Automated Hirschberg Test App to Evaluate Ocular Alignment
Published on: March 24, 2020
Large Language Model Simplification of Open Access Pediatric Strabismus Literature: Cross-Sectional Validation of
Mingming Jiang1,2,3, Mingming Zhou1,2,3, Xiaomei Wan1,2,3
1Eye Institute of Shandong First Medical University, Qingdao Eye Hospital of Shandong First Medical University, No.5 Yan'erdao Road, Shinan District, Qingdao, Shandong, 266000, China, 86 15898825081.
Background:
Peer-reviewed medical literature consistently violates established health literacy readability targets, creating a gap that effectively excludes patients and caregivers from accessing evidence-based information.
Objective:
This study aimed to evaluate whether a large language model (LLM) can generate plain-language summaries of pediatric strabismus literature while preserving clinical fidelity and meeting established health literacy readability targets.
Methods:
This cross-sectional study analyzed 85 open access, peer-reviewed pediatric strabismus articles published between 2022 and 2025, stratified by strabismus subtype, surgical relevance, and publication type. Full-text articles were processed using DeepSeek-V3 (DeepSeek) via a structured prompt, which instructed the model to provide a simplified summary meeting the following requirements for each article: a seventh-grade or lower reading level, a maximum length of 800 words, and strict preservation of medically significant data. Primary outcomes were readability scores measured by the Flesch-Kincaid Grade Level (FKGL) and Simple Measure of Gobbledygook (SMOG) indices. Secondary outcomes included clinical fidelity, which was independently assessed by 2 fellowship-trained pediatric strabismus specialists.
Results:
Baseline articles demonstrated a mean FKGL score of 15.79 (SD 1.53) and a mean SMOG score of 14.41 (SD 1.09). Following LLM simplification, the mean FKGL score significantly decreased from 15.79 (SD 1.53) to 7.84 (SD 1.30), representing a mean difference of 7.95 (95% CI 7.52-8.38; P<.001). Similarly, the mean SMOG score decreased from 14.41 (SD 1.09) to 7.68 (SD 0.94), representing a mean difference of 6.73 (95% CI 6.42-7.04; P<.001). Postsimplification readability did not differ significantly by strabismus subtype or surgical relevance (all adjusted P>.05). However, case reports retained slightly higher FKGL scores (mean 8.35, SD 0.89) compared to reviews (mean 7.89, SD 1.37) and original research (mean 7.48, SD 1.40) (adjusted P=.003). Out of the 85 summaries, clinical fidelity was rated good in 81 (95.29%), moderate in 4 (4.71%; these were exclusively summaries of review articles), and poor in 0 (0%).
Conclusions:
DeepSeek-V3 effectively reduced the reading level of complex pediatric strabismus literature by approximately 8 grade levels, achieving National Institutes of Health-recommended eighth-grade or lower targets without compromising clinical accuracy. When integrated with clinician oversight, LLM-generated summaries offer a scalable, equitable tool to enhance health literacy and support shared decision-making for patients and caregivers.

