Related Experiment Video
Updated: May 28, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The Quality of AI-Generated CABG Counseling: A Blinded Comparison of Two Language Models
Alper Özbakkaloğlu1, Ömer Faruk Rahman1, Ercan Keleş1
1Department of Cardiovascular Surgery, İzmir Bakırçay University, Menemen 35665, Turkey.
Abstract:
Objectives: Coronary artery bypass grafting (CABG) remains a fundamental surgical treatment for advanced coronary artery disease. With the increasing use of large language models to obtain health information, patients are increasingly turning to these systems to understand surgical options. However, their performance in generating patient-oriented CABG information has not been sufficiently evaluated. Therefore, this study aimed to compare the responses generated by ChatGPT and DeepSeek-R1 to patient questions about CABG in terms of scientific accuracy, comprehensibility, and level of unnecessary detail. Methods: Forty patient-oriented questions were developed based on online sources and clinical experience. Responses were obtained from ChatGPT and DeepSeek under standardized conditions. A blinded panel of four cardiovascular surgeons evaluated the responses using a five-point Likert scale across three domains. Statistical analyses were performed using paired tests. Results: DeepSeek generated significantly longer responses than ChatGPT (212.88 ± 48.13 vs. 188.7 ± 50.34 words; p < 0.001). Accuracy scores were higher for DeepSeek (median 4.5 vs. 4.25; p = 0.004), whereas comprehensibility and unnecessary detail scores were similar between the models. Overall scores were high for both models (4.32 ± 0.28 vs. 4.27 ± 0.30; p = 0.34). Conclusions: The responses generated by both models were generally evaluated favorably by the expert panel, with only limited differences observed between them. DeepSeek demonstrated higher accuracy, whereas ChatGPT tended to produce shorter and more concise responses. However, given the variability observed at the individual-question level, these findings should be interpreted with caution. Large language models may support patient information delivery but should not be considered reliable stand-alone sources for clinical decision-making or patient counseling.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
