Related Experiment Video
Updated: Sep 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing the accuracy and readability of large language models in answering Turkish-language patient-generated
1Department of Radiology, Faculty of Medicine, Necmettin Erbakan University, Konya, Türkiye.
Objective:
To assess the clinical accuracy, structural readability, and Ensuring Quality Information for Patients (EQIP-36) quality of Large Language Model (LLM) responses to real patient queries regarding radiological contrast agents and to define their safe boundaries in practice.
Methods:
The 30 most popular Turkish patient questions reflecting fears and risk perceptions about contrast agents were identified via Google Trends and AlsoAsked.com. These queries were presented to ChatGPT-5.4, Gemini 3, and Claude 4.6 Sonnet using a zero-shot learning approach. Responses were evaluated by experts for clinical accuracy against the 2024/2025 American College of Radiology (ACR) Manual on Contrast Media. Information quality and readability were measured using the expanded EQIP-36 scale and the Ateşman and Bezirci-Yılmaz indices.
Results:
Claude 4.6 Sonnet demonstrated the highest clinical accuracy and high compliance with guidelines (>90%). Gemini 3 exhibited a more conservative stance, occasionally resulting in over-triage by exaggerating risks. While all models scored exceptionally well in the Content and Structure dimensions of the EQIP-36, generating readable texts, they universally failed in the Identification dimension by omitting author names, update dates, or bibliographic sources. Readability indices indicated that comprehending the texts required an average of 10.5 to 11.6 years of formal education.
Conclusion:
LLMs generate highly readable and subjectively reassuring Turkish-language texts in response to patient queries regarding radiological contrast agents. However, due to domain-specific structural gaps regarding quantitative thresholds and lack of citations, they should be positioned as hybrid communication tools subject to mandatory review by specialist physicians rather than standalone medical advisors.
