Related Experiment Video
Updated: Jan 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Large Language Models in Nursing Examinations: Comparative Analysis of ChatGPT-3.5, ChatGPT-4 and
Peifang Li1, Menglin Jiang2, Jiali Chen1
1Department of Orthopedic Surgery, West China Hospital, Sichuan University/West China School of Nursing, Sichuan University, Chengdu, China.
Background:
While large language models (LLMs) have been widely utilised in nursing education, their performance in Chinese nursing examinations remains unexplored, particularly in the context of ChatGPT-3.5, ChatGPT-4 and iFLYTEK Spark.
Purpose:
This study assessed the performance of ChatGPT-3.5, ChatGPT-4 and iFLYTEK Spark on the 2022 China National Nursing Professional Qualification Exam (CNNPQE) at both the Junior and Intermediate levels. It also investigated whether the accuracy of these language models' responses correlated with the exam's difficulty or subject matter.
Methods:
We inputted 800 questions from the 2022 CNNPQE-Junior and CNNPQE-Intermediate exams into ChatGPT-3.5, ChatGPT-4 and iFLYTEK Spark to determine their accuracy rates in correctly answering the questions. We then analysed the correlation between these accuracy rates and the exams' difficulty levels or subjects.
Results:
The accuracy of ChatGPT-3.5, ChatGPT-4 and iFLYTEK Spark in the CNNPQE-Junior was 49.3% (197/400), 68.5% (274/400), and 61% (244/400), respectively, whereas it was 56.4% (225/399), 70.7% (282/399) and 57.6% (230/399) in the CNNPQE-Intermediate. When considering different grades, the differences in accuracy rates among the three models were statistically significant (M2 = 95.531, degrees of freedom (df) = 4, p < 0.001). These accuracy rates of ChatGPT-4 in the elementary knowledge, relevant professional knowledge, professional knowledge, and professional practice ability were 74.5%, 63.5%, 79% and 62.3%, respectively, leading in accuracy in other subjects in the CNNPQE. The results of the Cochran-Mantel-Haenszel (CMH) test showed that when considering different subjects, there was a statistically significant difference in accuracy rates of three LLMs (M2 = 97.435, df = 4, p < 0.001).
Conclusions:
ChatGPT-4 and iFLYTEK Spark performed well on Chinese nursing examinations and demonstrated potential as valuable tools in nursing education.
