Related Experiment Videos
A Comparative Study of ad^Writer and Human Evaluations of Guideline Reporting Quality: Performance of ChatGPT and
Jie Zhang1, Qianling Shi2, Hui Liu1
1Research Unit of Evidence-Based Evaluation and Guidelines, Chinese Academy of Medical Sciences (2021RU017), School of Basic Medical Sciences, Lanzhou University, Lanzhou, China.
Aim:
The proliferation of clinical practice guidelines presents significant challenges for ensuring evidence-based care, since evaluating their trustworthiness with the Reporting Items for practice Guidelines in HealThcare (RIGHT) criteria is both time-consuming and resource-intensive. To address this, the Artificial intelligence-empowered DeVelopment and AssessmeNt of healthCare guidElines and stanDards (ADVANCED) working group developed ad^Writer, an artificial intelligence system, designed to automate guideline reporting quality assessment. This study aimed to determine the agreement and reliability of ad^Writer compared with human evaluations.
Methods:
Guidelines were identified from four published systematic reviews, and through direct communication with the original authors, 60 were ultimately included. These guidelines were evaluated using ad^Writer with two different large language models, each performing three independent assessments. A majority rule was applied to synthesize results. Results from systematic reviews, which served as the reference standard. Agreement, consistency, and stability of the artificial intelligence-based assessments were examined using Bland-Altman plots, accuracy measures, and intraclass correlation coefficient (ICC), with a focus on identifying patterns of performance across individual RIGHT items.
Results:
Both models showed good agreement with human but had wide limits of agreement. DeepSeek performed better (mean difference = -1.40, standard deviation = 5.00, 95% limits of agreement ([-11.22, 8.42]) and had higher accuracy scores. Both achieved ≥ 75% accuracy on 23 items (65.71%), however, the results varied. ChatGPT showed higher consistency (ICC = 0.93) compared to DeepSeek (ICC = 0.88).
Conclusions:
The ad^Writer system leverages large language models to demonstrate the potential for rapid, accurate, and reproducible guideline quality assessments based on RIGHT. By serving as a supportive tool for manual evaluation, this approach enhances the efficiency of the guideline appraisal process, saving time and human resources for stakeholders.