Related Experiment Videos
Reliability and readability of Artificial Intelligence generated information on acute cholecystitis: A comparative
Mabel Lucero Olarte Jurado1, Natalia Pimiento Blanco1, María José Prieto Otero1
1Facultad de Medicina, Universidad Autónoma de Bucaramanga, Bucaramanga, Colombia.
Objective:
Acute cholecystitis (AC) is among the most frequently encountered conditions in emergency care settings. Recently, artificial intelligence (AI)-driven language models have emerged as innovative resources for accessing and synthesising medical information; however, their reliability and readability across different languages remain insufficiently established. Therefore, this study aimed to assess the reliability and readability of AC information provided by AI-based language models.
Methods:
Seven standardised questions were formulated in Spanish and English. Two independent reviewers subsequently evaluated the reliability of the responses via a validated assessment tool. Spanish readability was measured via the Flesch-Szigrist formula. In English, the Flesch Reading Easy score and the Flesch‒Kincaid readability score were employed. The results were then compared by tool and language.
Results:
Regarding reliability in Spanish, Perplexity® generated the highest percentage of complete responses (57.14%), followed by both ChatGPT® and Gemini® (42.86%). In English, Gemini® demonstrated the strongest performance with a score of 85.71%, followed by both Perplexity® and ChatGPT® with 71.43% each. The readability analysis revealed significant differences among the AI models in both Spanish and English (p < 0.05). In Spanish, Gemini® generated the most readable responses, whereas Perplexity® produced the least readable text. In English, ChatGPT® achieved the lowest Flesch-Kincaid grade level (highest readability), while Perplexity® generated the most complex responses.
Conclusion:
Despite the absence of statistically significant differences in reliability among the AI models, Gemini® generated the highest proportion of complete responses in English, whereas Perplexity® showed the highest proportion in Spanish. Moreover, significant differences in readability were observed among the three AI language models, with Perplexity® consistently generating the least readable content across both languages. There is a clear need for improvements to optimise the accuracy and accessibility of AI-generated medical information about AC.