Related Experiment Video
Updated: Feb 24, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Ground Truth Creation for Complex Clinical NLP Tasks - an Iterative Vetting Approach and Lessons Learned
Jennifer J Liang1, Ching-Huei Tsou1, Murthy V Devarakonda1
1IBM Research, Yorktown Heights, NY, USA.
Abstract:
Natural language processing (NLP) holds the promise of effectively analyzing patient record data to reduce cognitive load on physicians and clinicians in patient care, clinical research, and hospital operations management. A critical need in developing such methods is the "ground truth" dataset needed for training and testing the algorithms. Beyond localizable, relatively simple tasks, ground truth creation is a significant challenge because medical experts, just as physicians in patient care, have to assimilate vast amounts of data in EHR systems. To mitigate potential inaccuracies of the cognitive challenges, we present an iterative vetting approach for creating the ground truth for complex NLP tasks. In this paper, we present the methodology, and report on its use for an automated problem list generation task, its effect on the ground truth quality and system accuracy, and lessons learned from the effort.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Impression Management Techniques IV: Altercasting
Extraction: Advanced Methods
