Related Experiment Video
Updated: May 8, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Large Language Models for Turkish Emergency CT Impression Drafting: Quality, Critical Omissions, and
Halil Tekdemir1, Esra Çıvgın2, Şebnem Akpınar2
1Department of Radiology, Ankara Etlik City Hospital, Ankara, Turkey. haltek04726@gmail.com.
Journal of Imaging Informatics in Medicine
|May 6, 2026
Summary
Large language models (LLMs) show promise for drafting Turkish emergency CT reports, achieving acceptable quality. However, occasional critical omissions necessitate clinician oversight for safe use in medical imaging.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Natural Language Processing
Background:
- Emergency computed tomography (CT) reports are crucial for patient diagnosis and management.
- Drafting accurate and concise impression text requires significant radiologist expertise.
- Large language models (LLMs) offer potential for automating aspects of medical report generation.
Purpose of the Study:
- To evaluate and compare the performance of different LLMs in generating Turkish emergency CT impression text.
- To assess the quality, risk of critical omissions, and readability of LLM-generated impressions across various anatomical regions.
- To quantify the performance differences between LLMs and identify areas for improvement.
Main Methods:
- Retrospective analysis of 802 emergency CT reports (abdomen, chest, cranial, head and neck).
- Four LLMs (Grok-2, ChatGPT-4o-Latest, Gemini-2.0-Flash, DeepSeek-V3-FW) generated impression drafts from report sections.
- Radiologist evaluation using a 4-point Likert scale, recording critical finding omissions, and calculating the Ateşman readability index.
Main Results:
- LLM impression quality varied by model and anatomical region, with higher scores for head/neck and cranial CTs.
- Critical omissions were infrequent but exhibited model- and region-specific patterns, notably in abdominal CT for one model.
- Readability was generally high and comparable between radiologist-generated text and top-performing LLM outputs.
Conclusions:
- LLM-generated Turkish CT impressions can achieve acceptable quality for specific applications.
- The risk of critical omissions, though low, persists and requires careful monitoring.
- LLMs should function as clinical decision-support tools, mandating radiologist oversight rather than autonomous deployment.