Related Experiment Video
Updated: May 20, 2026

07:50
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Feasibility of pooling annotated corpora for clinical concept extraction
Kavishwar Wagholikar1, Manabu Torii, Siddhartha Jonnalagadda
1Mayo Clinic, Rochester, MN;
Summary
Pooling clinical corpora for machine learning concept extraction decreased performance. Developing standard annotation guidelines is crucial for improving medical problem detection accuracy and portability of NLP taggers.
Area of Science:
- Clinical Natural Language Processing (NLP)
- Machine Learning in Healthcare
- Medical Informatics
Background:
- Annotated corpora are essential for training machine learning models for clinical concept extraction.
- Creating these corpora is resource-intensive for individual institutions.
- Pooling corpora from multiple sources is a potential strategy to overcome resource limitations.
Purpose of the Study:
- To investigate if pooling corpora from different sources improves the performance and portability of machine learning taggers for medical problem detection.
- To evaluate the impact of combining data from the 2010 i2b2/VA NLP challenge and Mayo Clinic Rochester datasets.
Main Methods:
- Utilized machine learning taggers trained on pooled corpora from two distinct sources: the 2010 i2b2/VA NLP challenge and Mayo Clinic Rochester.
- Evaluated tagger performance using F1-score for medical problem recognition.
- Examined annotation guidelines to identify sources of incompatibility between corpora.
Main Results:
- Contrary to expectations, pooling corpora resulted in a decrease in the F1-score, indicating reduced performance.
- Differences in annotation guidelines were identified as a key factor contributing to corpus incompatibility.
- The study highlights challenges in achieving effective model portability when combining diverse datasets.
Conclusions:
- Pooling annotated clinical corpora from different institutions may not improve, and can even decrease, the performance of machine learning models for medical problem detection.
- Standardized annotation guidelines are necessary for the clinical NLP community to ensure corpus compatibility and enhance model portability.
- Future efforts should focus on developing consensus guidelines to facilitate effective data sharing and improve NLP tool development in healthcare.
More Related Videos
Related Concept Videos
Clinical Trials: Overview
Clinical development focuses on how the drug will interact with the human body and encompasses four key phases of clinical trials, each serving a specific purpose in assessing the safety and effectiveness of new drugs. These phases overlap and build upon one another. Phase I involves a small group of healthy volunteers (typically 20-80 individuals) or, in cases where significant toxicity is expected, patients with the targeted disease, such as cancer or AIDS. The volunteers are tested for...
Clinical Trials
Clinical trials are prospective experimental studies conducted on humans to determine the safety and efficacy of treatments, drugs, diet methods, and medical devices. Using statistics in clinical trials enables researchers to derive reasonable and accurate conclusions from the collected data, allowing them to make wise decisions in uncertain situations. In medical research, statistical methods are crucial for preventing errors and bias.
There are four phases in a clinical trial. A phase one...
There are four phases in a clinical trial. A phase one...

