Related Experiment Video
Updated: Jan 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Leveraging large language models for structured information extraction from pathology reports
Jeya Balaji Balasubramanian1, Daniel Adams2,3, Ioannis Roxanis4
1Division of Cancer Epidemiology and Genetics, National Cancer Institute, 9609 Medical Center Dr, NCI Shady Grove, Room 7E554, Rockville, MD 20850, USA.
Large language models (LLMs) achieve human-level accuracy in extracting structured data from breast cancer histopathology reports. This automated approach enhances data accessibility for clinical research, offering a scalable alternative to manual extraction.
Area of Science:
- Computational pathology
- Natural Language Processing
- Medical Informatics
Background:
- Structured information extraction from unstructured histopathology reports is crucial for clinical research data accessibility.
- Manual extraction is time-consuming and limits scalability.
- Large language models (LLMs) offer automated extraction via zero-shot prompting, eliminating the need for labeled data or training.
Purpose of the Study:
- To evaluate the accuracy of LLMs in extracting structured information from breast cancer histopathology reports.
- To compare LLM performance against manual extraction by a trained human annotator.
Main Methods:
- Developed the Medical Report Information Extractor web application utilizing LLMs.
- Created a gold-standard dataset for evaluation.
- Assessed five LLMs, including GPT-4o and Llama 3 models, on 111 breast cancer histopathology reports, extracting 51 pathology features.
Main Results:
- Llama 3.1 405B (94.7% accuracy) and GPT-4o (96.1%) demonstrated comparable accuracy to the human annotator (95.4%).
- Llama 3.1 70B (91.6%) performed below human accuracy but offers a viable self-hosting option due to lower computational needs.
Conclusions:
- An open-source tool for structured information extraction achieved expert human-level accuracy using state-of-the-art LLMs.
- The tool is customizable via natural language and promotes data standardization, accessibility, and interoperability for analytics.
More Related Videos
07:35A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy