Related Experiment Video
Updated: May 20, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language model comparisons between English and Chinese query performance for cardiovascular prevention
Hongwei Ji1,2,3, Xiaofei Wang4, Ching-Hui Sia5,6
1Beijing Visual Science and Translational Eye Research Institute (BERI), Eye Center of Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University, Beijing, China.
ChatGPT-4.0 excels in providing accurate cardiovascular disease (CVD) prevention information in English, outperforming other large language models (LLMs). Performance slightly decreases for Chinese queries, highlighting potential language bias in LLMs.
Area of Science:
- Artificial Intelligence in Healthcare
- Natural Language Processing for Medical Information
- Cardiovascular Disease Prevention
Background:
- Large language models (LLMs) show potential for answering public health queries on cardiovascular disease (CVD) prevention.
- The accuracy and reliability of information from general LLMs for CVD prevention remain under investigation.
Purpose of the Study:
- To evaluate the accuracy and consistency of leading large language models (LLMs) in responding to cardiovascular disease (CVD) prevention queries.
- To compare the performance of BARD, ChatGPT-3.5, ChatGPT-4.0, and ERNIE in English and Chinese CVD prevention information retrieval.
Main Methods:
- Assessed capabilities of BARD, ChatGPT-3.5, ChatGPT-4.0, and ERNIE using 75 cardiovascular disease (CVD) prevention questions in English and Chinese.
- Responses were rated for accuracy as appropriate, borderline, or inappropriate.
- Evaluated temporal improvement and self-awareness of response correctness for each model.
Main Results:
- In English, ChatGPT-4.0 achieved the highest accuracy (97.3%), followed by ChatGPT-3.5 (92.0%) and BARD (88.0%).
- All models showed improvement over time, with ChatGPT-4.0 demonstrating superior temporal gains and self-correction.
- For Chinese queries, ERNIE (84.0%) and ChatGPT-3.5 (88.0%) showed higher accuracy than ChatGPT-4.0 (85.3%), with ERNIE excelling in improvement and self-awareness.
Conclusions:
- ChatGPT-4.0 is the leading LLM for English CVD prevention queries, demonstrating superior accuracy, temporal improvement, and self-awareness.
- A slight performance decrease in Chinese suggests potential language bias in current LLMs.
- Regular, rigorous evaluations are crucial to assess the quality and limitations of LLM-generated medical information across languages.
Related Concept Videos
Lung Capacity
Language and Cognition
Targeted Cancer Therapies
There are several types of targeted therapies against...

