Related Experiment Video
Updated: Mar 13, 2026

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.8K
Advanced Prompting Techniques Informed by Clinical Expertise Improve the Accuracy of LLM Data Extraction but Increase
Yifei Wang1, Alejandro Alejandrez Cisneros2, Carly Stewart3
1Department of Epidemiology and Biostatistics, University of California San Francisco, San Francisco, CA, USA. yifei.wang@ucsf.edu.
Journal of Imaging Informatics in Medicine
|March 12, 2026
Summary
Prompt engineering for generative AI in medical text classification shows promise. External knowledge sources improve accuracy, while chain-of-thought and recursive criticism boost sensitivity but increase non-determinism.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Natural Language Processing
Background:
- Generative artificial intelligence (AI) and prompt engineering techniques have advanced significantly for classification tasks.
- The effectiveness of these methods in extracting structured data from unstructured medical text remains unclear.
- Radiology reports are a key source of unstructured medical data.
Purpose of the Study:
- To evaluate five large language model (LLM) prompting strategies for extracting structured categorical data from unstructured radiology reports.
- To assess the impact of different prompting techniques (external knowledge, recursive criticism, chain-of-thought) on accuracy and non-determinism.
- To determine the efficacy of prompt engineering in handling both explicit and implicit information within medical texts.
Main Methods:
- Five distinct LLM prompting strategies were applied to a structured categorical question posed against unstructured radiology reports.
- Strategies incorporated external knowledge sources, recursive criticism and improvement, and chain-of-thought techniques.
- Efficacy was measured by overall accuracy, sensitivity to a non-explicit category, and the rate of non-determinism (output variability across multiple runs).
Main Results:
- An external knowledge source significantly improved sensitivity for a category requiring medical knowledge extrapolation from 10% to 66%, with minimal impact on overall accuracy.
- Combining external knowledge with recursive criticism and chain-of-thought further increased sensitivity to 78%.
- However, these advanced techniques increased the proportion of exams with non-deterministic outputs from 7% to 14%.
Conclusions:
- Prompt engineering strategies, particularly those incorporating external knowledge, enhance the extraction of specific categories from radiology reports.
- While advanced techniques like chain-of-thought and recursive criticism improve sensitivity, they introduce greater variability in AI outputs.
- Further research is needed to balance accuracy gains with the reliability of AI-driven medical data extraction.
Keywords:
Artificial intelligenceChain-of-thought promptingLarge language modelNon-determinismRadiology reportRecursive criticism and improvement
