Related Experiment Video
Updated: Feb 24, 2026

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.7K
Ground Truth Creation for Complex Clinical NLP Tasks - an Iterative Vetting Approach and Lessons Learned
Jennifer J Liang1, Ching-Huei Tsou1, Murthy V Devarakonda1
1IBM Research, Yorktown Heights, NY, USA.
Summary
Creating accurate ground truth data for natural language processing (NLP) in healthcare is challenging. An iterative vetting approach improves data quality for training clinical NLP algorithms, enhancing system accuracy.
Area of Science:
- Clinical Informatics
- Natural Language Processing (NLP)
- Medical Data Analysis
Background:
- Natural Language Processing (NLP) offers potential for analyzing electronic health records (EHRs) to reduce physician cognitive load.
- High-quality "ground truth" datasets are essential for training and validating clinical NLP algorithms.
- Creating ground truth for complex medical data is challenging due to the vast amount of information physicians must process.
Purpose of the Study:
- To present an iterative vetting methodology for creating ground truth datasets for complex NLP tasks in healthcare.
- To evaluate the impact of this methodology on ground truth quality and system accuracy for automated problem list generation.
Main Methods:
- Developed and implemented an iterative vetting approach for ground truth data creation.
- Applied the methodology to an automated problem list generation task using EHR data.
- Assessed the quality of the generated ground truth and the accuracy of the NLP system.
Main Results:
- The iterative vetting approach was successfully applied to create ground truth for a complex NLP task.
- The methodology demonstrated a positive effect on the quality of the ground truth data.
- Improved ground truth quality led to enhanced accuracy in the automated problem list generation system.
Conclusions:
- An iterative vetting process is an effective strategy for mitigating inaccuracies in ground truth creation for clinical NLP.
- This approach enhances the reliability of datasets used for training and testing medical NLP algorithms.
- Lessons learned from this effort provide valuable insights for future clinical NLP development.
Related Concept Videos
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
Improving Translational Accuracy
3.7K
3.7K
Data Validation
7.1K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
7.1K
Impression Management Techniques IV: Altercasting
197
Altercasting is a strategic communication technique in which an individual imposes a specific identity or social role onto another person to influence their behavior and shape the interaction. By presuming a role—such as “responsible leader” or “patient person”—altercasting encourages the target to conform to that identity, often aligning their behavior with the expectations associated with the role. The power of this tactic lies in its subtlety; once a role...
197
Extraction: Advanced Methods
1.2K
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
1.2K
