Related Experiment Video
Updated: Jan 22, 2026

Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
Clinical trial cohort selection based on multi-level rule-based natural language processing system.
This study introduces a clinical natural language processing (NLP) system for identifying patients eligible for clinical trials. The system uses rule-based processing and integrates clinical knowledge resources like UMLS and UIMA. It was tested in the 2018 n2c2-1 challenge and achieved an F-measure of 0.9028, ranking fourth among participants. The system's performance was close to the top systems, even with limited training data. A separate general cNLP system was also developed, which showed promise in extracting clinical concepts from unstructured data. The authors suggest that combining both systems could lead to better performance. This work highlights the potential of rule-based systems in automating patient eligibility assessments for clinical trials.
Area of Science:
- Clinical informatics
- Natural language processing in healthcare
- Medical data analysis
Background:
Selecting eligible patients for clinical trials remains a complex and resource-intensive task. Traditional methods rely heavily on manual chart reviews, which are slow and prone to human error. While prior research has demonstrated the potential of automated systems, few have focused on integrating rule-based approaches with clinical knowledge resources. This gap motivated the development of a clinical NLP system that leverages structured and unstructured data from medical records. Existing tools often fail to capture nuanced eligibility criteria, especially when dealing with longitudinal patient data. The challenge of accurately parsing clinical text has driven the need for systems that can interpret both structured and unstructured information. Prior studies have shown that rule-based systems can be effective when combined with medical terminologies. However, no prior work had resolved how to optimize these systems for small training datasets. This paper addresses that limitation by proposing a novel framework for cohort selection.
Purpose Of The Study:
The aim of this work was to develop and evaluate a clinical NLP system for patient cohort selection in clinical trials. The system integrates rule-based processing with clinical knowledge resources to improve eligibility assessment accuracy. The authors sought to address the challenge of parsing longitudinal medical records, which often contain unstructured clinical notes. They aimed to create a system that could handle both structured and unstructured data efficiently. The study also explored the potential of combining rule-based and general cNLP approaches for improved performance. The authors focused on the 2018 n2c2-1 challenge dataset to test their system's effectiveness. Their goal was to demonstrate that a rule-based system could achieve high performance even with limited training data. This work contributes to the broader effort of automating clinical trial recruitment processes.
Main Methods:
The authors constructed a rule-based clinical NLP system using a generic framework enhanced with lexical, syntactic, and meta-level knowledge inputs. They integrated the Unified Medical Language System (UMLS) and Unstructured Information Management Architecture (UIMA) into their general cNLP system. The rule-based system was designed to parse clinical text and extract eligibility-relevant information. The general cNLP system focused on extracting clinical concepts from unstructured data. Both systems were trained and evaluated using the 2018 n2c2-1 challenge dataset. The authors implemented a multi-level processing approach to capture both explicit and implicit patient eligibility criteria. They used task-specific rules to guide the extraction of relevant clinical entities. The evaluation metrics included F-measure to assess system performance against other participants.
Main Results:
The rule-based clinical NLP system achieved an F-measure of 0.9028 in the 2018 n2c2-1 challenge. This performance ranked fourth among all participants and was within 1% of the top-performing system. The general cNLP system, while less accurate, demonstrated strengths in clinical concept extraction. The rule-based system's performance suggests that structured rules can effectively handle cohort selection tasks. The general cNLP system showed promise in capturing complex clinical concepts from unstructured text. The authors observed that combining both systems could enhance overall performance. The rule-based approach proved robust even with limited training data. These findings support the use of rule-based systems for clinical trial eligibility assessments.
Conclusions:
The authors concluded that a well-designed rule-based clinical NLP system can achieve strong performance on cohort selection tasks. Their system's F-measure of 0.9028 indicates its effectiveness in parsing clinical records for eligibility criteria. The rule-based system's performance was close to the top-performing systems in the challenge. The authors noted that the general cNLP system had complementary strengths in clinical concept extraction. They proposed that a hybrid system combining both approaches could improve performance further. The study demonstrated that rule-based systems are viable even with small training datasets. The results suggest that integrating structured rules with clinical knowledge resources is beneficial. The authors emphasized the importance of leveraging both structured and unstructured data for accurate cohort selection.
Frequently Asked Questions
The rule-based clinical NLP system achieved an F-measure of 0.9028 in the 2018 n2c2-1 challenge, ranking fourth among participants.
The rule-based system uses lexical, syntactic, and meta-level rules, while the general cNLP system relies on UMLS and UIMA for concept extraction.
UMLS helps standardize clinical terminology, improving the accuracy of concept extraction in unstructured clinical text.
The rule-based system achieved higher F-measure and performed well with limited training data.
This score indicates strong performance in identifying eligible patients for clinical trials.
The authors propose that combining rule-based and general cNLP systems could surpass current state-of-the-art performance.
Related Concept Videos
What is Natural Selection?
Clinical Trials
There are four phases in a clinical trial. A phase one...
Clinical Trials: Overview
Limits to Natural Selection
Natural Selection and Adaptation
Beyond physical adaptations,...
Leveling Effect and Non-Aqueous Acid-Base Solutions
The Leveling Effect of a Solvent
A generic acid (HA) reacts with the generic base (B-) to yield the corresponding conjugate base (A-) and conjugate acid (HB):

