Related Experiment Video
Updated: May 5, 2026

07:50
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
16.4K
Leveraging LLMs for Efficient Data Structure Standardization in Chinese Medical Examination Reports
Summary
Large language models (LLMs) can standardize Chinese medical reports by accurately predicting department and item names. This advancement improves health data consistency and automated analysis across hospitals.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Natural Language Processing
Background:
- Chinese medical examination reports lack standardized terminology, hindering health assessment and data analysis.
- Variations in reporting across hospitals create significant challenges for consistent data interpretation.
Purpose of the Study:
- To evaluate the efficacy of large language models (LLMs) in standardizing Chinese medical examination reports.
- To assess prompt engineering and LoRA fine-tuning for improving LLM performance on medical text data.
Main Methods:
- Utilized a large dataset of over 900,000 Chinese medical reports.
- Employed prompt engineering and LoRA fine-tuning techniques on various LLMs, including Qwen2.5-14B.
- Conducted cross-validation experiments to evaluate model generalization on unseen hospital data.
Main Results:
- Fine-tuned LLMs, especially Qwen2.5-14B, achieved high accuracy in standardizing department and item detail names.
- LLMs significantly outperformed traditional methods like BERT in medical report standardization.
- Achieved an average accuracy of 98.34% on unseen hospital data, demonstrating strong generalization capabilities.
Conclusions:
- LLMs show significant potential for standardizing Chinese medical examination reports.
- This research provides a foundation for automated processing and analysis of large-scale medical data.
- Fine-tuned LLMs offer a robust solution to the challenge of non-standardized medical terminologies.

