Related Experiment Video
Updated: Oct 28, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
516
Enhancing Pulmonary Disease Prediction Using Large Language Models With Feature Summarization and Hybrid
Ronghao Li1, Shuai Mao2, Congmin Zhu1
1School of Biomedical Engineering, Capital Medical University, No. 10, Xitoutiao, You An Men, Fengtai District, Beijing, 100069, China, 86 010-83911542.
Journal of Medical Internet Research
|June 11, 2025
Summary
Large language models (LLMs) with novel prompt engineering significantly improve pulmonary disease prediction accuracy. This advanced approach outperforms traditional models, offering a promising tool for clinical decision-making.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing for Healthcare
- Pulmonary Disease Diagnostics
Background:
- Rapid advancements in Large Language Models (LLMs) offer potential for managing complex clinical text.
- Challenges exist in applying prompt engineering to medical texts for diagnostic tasks due to complexity and specificity.
Purpose of the Study:
- To explore LLMs with novel prompt engineering for enhanced interpretability and pulmonary disease prediction.
- To improve upon traditional deep learning models for disease prediction using advanced AI techniques.
Main Methods:
- A retrospective dataset of 2965 chest CT radiology reports from healthy individuals and patients with tuberculosis, lung cancer, and pneumonia was used.
- A novel prompt engineering strategy integrating feature summarization (F-Sum), chain of thought (CoT) reasoning, and a hybrid retrieval-augmented generation (RAG) framework was developed.
- Three LLMs (GLM-4-Plus, GLM-4-air, GPT-4o) and a BERT-based model were evaluated on internal and external validation datasets.
Main Results:
- The proposed method with GLM-4-Plus achieved the highest F1-score (0.89) and accuracy (0.89) on the test dataset.
- GPT-4o demonstrated superior performance on the external validation dataset with an F1-score of 0.86 and accuracy of 0.92.
- The LLM-based F-Sum and auto-generated CoT strategy outperformed manually selected samples and doctor-designed CoT.
Conclusions:
- LLMs, enhanced by the proposed prompt engineering strategy, significantly outperform traditional models in pulmonary disease prediction.
- The approach demonstrates improved accuracy, flexibility, and adaptability in complex medical contexts, aiding clinical decision-making.
- This technology holds promise for advancing disease diagnosis, especially in resource-constrained settings.

