Related Experiment Videos
Assessing the Accuracy, Completeness, and Reference Quality of GPT-4 and Google for Gynecologic Cancer Information:
Alexandra H Smick1, Martins Ayoola1, Rajeshree Rajpara2
1Gynecologic Oncology, University of California, Los Angeles, Los Angeles, CA, United States.
Background:
Patients with newly diagnosed gynecologic cancers often seek information online, but the quality of available resources may be inconsistent. Although GPT-4 may offer an alternative to traditional internet search engines, its performance remains largely under-studied in gynecologic oncology.
Objective:
This study aimed to compare the completeness, accuracy, and reference quality of responses generated by GPT-4 with those generated by Google in clinical scenarios involving a new diagnosis of a gynecologic cancer.
Methods:
Clinical scenarios representing early- and advanced-stage endometrial, ovarian, and cervical cancers were developed by gynecologic oncologists using publicly available patient education materials. Each scenario included 4 standardized questions addressing etiology, prognosis, treatment, and treatment efficacy. GPT-4 and Google were queried for each question, with new sessions for GPT-4 and private browsing for Google to minimize bias. Responses were independently evaluated by 4 gynecologic oncology experts who were blinded to each other's ratings. Accuracy was scored on a 6-point Likert scale; completeness and reference quality were scored on 3-point scales. Reference quality was categorized as low (commercial), medium (institutional or government), or high (peer reviewed). Optional free-text reviewer comments were collected and summarized descriptively to provide context for the quantitative findings. Descriptive statistics and Wilcoxon signed-rank tests were used for analysis.
Results:
Across 6 clinical scenarios and 21 standardized questions (N=84 total responses), GPT-4 outperformed Google across all evaluated domains. The median accuracy score was 6.00 (IQR 5.00-6.00) for GPT-4 and 5.00 (IQR 4.00-6.00) for Google (P=.04). The median completeness score was 3.00 (IQR 3.00-3.00) for GPT-4 and 2.00 (IQR 1.00-3.00) for Google (P=.009). Reference quality was also higher for GPT-4, with a median score of 3.00 (IQR 3.00-3.00) compared to 2.00 (IQR 2.00-2.00) for Google (P=.009). Reviewer comments noted that GPT-4 provided more accurate, comprehensive, and personalized responses, while Google returned less detailed content from general consumer health websites rather than peer-reviewed sources.
Conclusions:
GPT-4 may serve as a reliable and high-quality resource for information in gynecologic oncology. Compared to Google, it delivered more accurate and complete content, with higher-quality references. Further research is needed to assess the readability and accessibility of GPT-4-generated content across diverse patient populations.