Related Experiment Video
Updated: Jun 4, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative Evaluation of Large Language Models for Translating Radiology Reports into Hindi.
Amit Gupta1, Ashish Rastogi1, Hema Malhotra1
1Department of Radiology, Dr. Bhim Rao Ambedkar Institute-Rotary Cancer Hospital, All India Institute of Medical Sciences, New Delhi.
Four large language models (LLMs) were evaluated for translating radiology reports into Hindi. GPT-4o and Gemini showed strong performance, with results varying by prompt, indicating LLM potential for simplifying medical information.
Area of Science:
- Medical imaging and artificial intelligence
- Natural language processing in healthcare
- Radiology report translation
Background:
- Radiology reports contain complex medical information.
- Accurate translation is crucial for patient understanding and care.
- Large Language Models (LLMs) offer potential for simplifying medical texts.
Purpose of the Study:
- To compare the performance of four leading LLMs (GPT-4o, GPT-4, Gemini, Claude Opus) in translating radiology report impressions into simple Hindi.
- To evaluate the accuracy and quality of LLM-generated Hindi translations.
- To assess the impact of different prompts on translation performance.
Main Methods:
- Retrospective analysis of 100 CT scan report impressions from a cancer center.
- Reference translations created by bilingual radiology staff and a radiologist.
- Two distinct prompts used to test LLM translation capabilities.
- Radiologist review for misinterpretations, omissions, and additions.
- Quantitative evaluation using BLEU, METEOR, TER, and CHRF scores.
Main Results:
- Overall, few errors (9 misinterpretations, 2 omissions) were noted across 800 LLM translations.
- Gemini excelled with Prompt 1 across multiple metrics (BLEU, METEOR, TER, CHRF).
- GPT-4o outperformed all models with Prompt 2 across all evaluated metrics.
- Prompt 2 generally yielded superior translation scores compared to Prompt 1.
Conclusions:
- All evaluated LLMs demonstrate significant potential for translating and simplifying radiology reports into Hindi.
- Translation quality is influenced by the specific LLM and the wording of the prompt.
- Further research can optimize LLM use for accessible medical communication.
More Related Videos
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Improving Translational Accuracy
Leaky Scanning