Related Experiment Video
Updated: Mar 4, 2026

11:29
Measuring the Functional Abilities of Children Aged 3-6 Years Old with Observational Methods and Computer Tools
Published on: June 20, 2020
9.9K
Development and Performance Validation of an Automated Generative Pretrained Transformer-Based Evaluation Tool for
Yu-Jeng Ju1, Yi-Ching Wang1, Fan Chou2
1School of Occupational Therapy, College of Medicine, National Taiwan University, Taipei, Taiwan.
Archives of Physical Medicine and Rehabilitation
|March 2, 2026
Summary
An automated GPT-PEDro evaluation tool (GPT-PEDro ET) shows high reliability and validity for assessing randomized controlled trial (RCT) quality. This tool can streamline methodological quality evaluations in research.
Area of Science:
- Medical Informatics
- Clinical Research Methodology
- Artificial Intelligence in Healthcare
Background:
- The PEDro scale is a widely used tool for assessing the methodological quality of randomized controlled trials (RCTs).
- Manual evaluation of RCTs using the PEDro scale can be time-consuming and resource-intensive.
- Automating this process could enhance efficiency and consistency in research quality assessment.
Purpose of the Study:
- To develop an automated GPT-PEDro evaluation tool (GPT-PEDro ET).
- To validate the performance of the GPT-PEDro ET, including its intra-rater reliability and concurrent validity.
- To assess the tool's effectiveness in evaluating the methodological quality of RCTs using the PEDro scale.
Main Methods:
- A psychometric validation study employing a repeated measurements design was conducted.
- One hundred and twenty-five RCTs on neurofacilitation interventions in stroke rehabilitation were analyzed.
- Performance was evaluated using intraclass correlation coefficients (ICC) and prevalence-adjusted and bias-adjusted kappa (PABAK) statistics, alongside Bland-Altman analysis.
Main Results:
- The GPT-PEDro ET demonstrated almost perfect intra-rater reliability for both total scores (ICC=1.00) and individual items (PABAK=0.94-1.00).
- Concurrent validity was found to be moderate to high, with ICCs of 0.83 for total scores and PABAKs ranging from 0.68 to 0.97 for individual items.
- The tool showed strong agreement with established reliability and validity metrics.
Conclusions:
- The GPT-PEDro ET is a potentially valuable automated tool for evaluating the methodological quality of RCTs.
- The tool exhibits strong psychometric properties, suggesting its utility in research settings.
- It has the potential to reduce the workload associated with manual quality assessments and support clinical and research applications.
Keywords:
ChatGPTConcurrent validityIntrarater reliabilityPEDro scaleRandomized controlled trialsRehabilitation
