Related Experiment Video
Updated: Jul 4, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Investigating the Impact of Prompt Engineering on the Performance of Large Language Models for Standardizing
Lei Wang1, Wenshuai Bi1, Suling Zhao1
1BGI Research, Shenzhen, China.
Large language models (LLMs) and BERT show comparable performance in standardizing obstetric diagnostic terms from electronic medical records (EMRs). LLMs offer superior efficiency in unsupervised settings, with specific prompt engineering techniques significantly boosting performance.
Area of Science:
- Medical Informatics
- Natural Language Processing in Healthcare
- Artificial Intelligence in Medicine
Background:
- Electronic medical records (EMRs) offer significant research value, especially in obstetrics.
- Standardizing diagnostic terminology across institutions is crucial for accurate medical data analysis.
- Large language models (LLMs) are increasingly utilized for diverse medical applications, with prompt engineering being critical for their effectiveness.
Purpose of the Study:
- To evaluate and compare the performance of LLMs using various prompt engineering techniques.
- To standardize obstetric diagnostic terminology using real-world obstetric data.
- To assess the efficiency and effectiveness of LLMs against traditional models in a clinical context.
Main Methods:
- A four-step approach was employed, beginning with similarity measures for mapping diagnoses to ICD-10.
- Candidate terms were collected based on similarity scores for training data.
- Two LLMs (ChatGLM2, Qwen-14B-Chat [QWEN]) were used for zero-shot learning, and three BERT variants (BERT, whole word masking BERT, MC-BERT) were used for unsupervised optimal mapping term generation.
Main Results:
- LLMs and BERT demonstrated comparable performance, with LLMs showing advantages in efficiency in unsupervised scenarios.
- Prompt engineering significantly impacted LLM performance; the self-consistency approach in QWEN improved F1-score by 5% and precision by 7.9%.
- Momentum contrastive learning with BERT (MC-BERT) achieved the highest performance among BERT variants, though differences were minor.
Conclusions:
- LLMs, particularly with optimized prompts like those used with QWEN, can effectively standardize obstetric diagnostic terms.
- The precision achieved by QWEN prompts was comparable to that of BERT models.
- Unsupervised LLM approaches show potential for enhancing diagnostic term alignment in research and unlocking insights from patient data.
More Related Videos
Related Concept Videos
Formulating and Validating Nursing Diagnosis I
There are thirteen domains...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Formulating and Validating Nursing Diagnosis II
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...

