Related Experiment Video
Updated: Jan 16, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
MedError: A Machine-Assisted Framework for Systematic Error Analysis in Clinical Concept Extraction
Hongfang Liu1, Sunyang Fu2, Qiuhao Lu2
1University of Texas Health Science Center at Houston.
Abstract:
Error analysis is a critical step in evaluating and improving clinical concept extraction models, the most common clinical natural language processing (NLP) task. Unlike corpus annotation, which follows standardized protocols for creating gold-standard datasets, error analysis requires nuanced judgment grounded in both clinical expertise and NLP knowledge. This is especially important given the heterogeneity of clinical text, where variations in documentation style, note structure, and terminology can substantially influence model behavior. Despite its importance, there is currently no standardized, user-level framework to support systematic error analysis in clinical concept extraction task. In this study, we developed and validated MedError, a machine-assisted, human-in-the-loop framework designed to standardize and enhance error analysis for clinical concept extraction tasks. We collected and manually curated a corpus of 1,187 unique errors from a total of 4,237 notes across three different distinct hospitals. The error categories were defined using our previously validated error taxonomy and included 480 false negatives and 707 false positives across 25 error types and 48 clinical concept categories. We evaluated the performance of three proprietary and three open-source large language models (LLMs) in automatically classifying these errors into 26 and 15 predefined categories. We further developed a machine-assisted framework, MedError, which integrates best practices in error analysis, LLM-assisted classification and reasoning, and a user-friendly interface to enable more efficient, reproducible, and context-aware error analysis. The framework supports both single-site and federated multisite error analysis, facilitating the effective translation of clinical NLP systems into real-world settings.
More Related Videos
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Natural and Artificial Concepts
Random and Systematic Errors
Types of Errors: Detection and Minimization
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Concepts and Prototypes
The brain organizes this information using concepts, which are mental categories grouping linguistic data,...