Related Experiment Video
Updated: Mar 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Ensuring Reliability of Curated Electronic Health Record-Derived Data: The Validation of Accuracy for Large Language
Melissa Estevez1, Nisha Singh1, Lauren Dyson1
1Flatiron Health, New York, NY.
Large language models (LLMs) enhance oncology data curation but require robust quality checks. This study introduces a framework to ensure the reliability, accuracy, and fairness of LLM-extracted clinical data for trustworthy AI applications.
Area of Science:
- Clinical Informatics
- Artificial Intelligence in Medicine
- Oncology Data Science
Background:
- Large language models (LLMs) offer scalable and efficient curation of real-world data (RWD) from electronic health records in oncology.
- Challenges exist in ensuring the reliability, accuracy, and fairness of LLM-extracted clinical data for research and clinical use.
- Current RWD and AI quality assurance frameworks do not fully address LLM-specific complexities and error modes.
Purpose of the Study:
- To propose a comprehensive framework for evaluating the quality of clinical data extracted by LLMs.
- To address the unique challenges in RWD curation using AI in oncology.
- To ensure the trustworthiness of AI-generated evidence in cancer research and practice.
Main Methods:
- Integrating variable-level performance benchmarking against expert human abstraction.
- Implementing verification checks for internal consistency and plausibility of extracted data.
- Conducting replication analyses comparing LLM-extracted data with human-abstracted datasets or external standards.
- Incorporating bias assessment by stratifying analyses across demographic subgroups.
Main Results:
- The proposed framework enables identification of variables needing improvement and systematic detection of latent errors.
- It confirms the fitness-for-purpose of LLM-extracted datasets for real-world research.
- Bias assessment across demographic subgroups is supported, enhancing fairness evaluation.
Conclusions:
- A rigorous and transparent framework is essential for assessing LLM-extracted RWD quality in oncology.
- This framework advances industry standards for evaluating AI-driven evidence generation.
- It supports the trustworthy application of LLMs in oncology research and clinical practice.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Related Concept Videos
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Methods of Documentation VII: EMR
Purpose of Health Records I
Here's a breakdown of how health records serve these purposes:
Reliability and Validity