Related Experiment Video
Updated: Dec 10, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Chia, a large annotated corpus of clinical trial eligibility criteria
Fabrício Kury1, Alex Butler1, Chi Yuan1
1Columbia University in the City of New York, New York, NY, United States.
Insights
We introduce Chia, a large annotated corpus for clinical trial eligibility criteria. This dataset aids in developing advanced methods for extracting crucial patient information from trial data.
Area of Science:
- Natural Language Processing
- Clinical Informatics
- Biomedical Data Science
Background:
- Clinical trial eligibility criteria are complex and often unstructured.
- Efficiently extracting this information is crucial for trial recruitment and analysis.
- Existing methods for processing eligibility criteria are limited.
Purpose of the Study:
- To present Chia, a novel, large annotated corpus of patient eligibility criteria.
- To provide a benchmark dataset for developing and testing information extraction methods.
- To facilitate the transformation of eligibility criteria into computable formats.
Main Methods:
- Extracted eligibility criteria from 1,000 interventional, Phase IV clinical trials registered in ClinicalTrials.gov.
- Annotated 12,409 eligibility criteria, identifying 41,487 entities and 25,017 relationships.
- Represented each criterion as a directed acyclic graph for potential Boolean logic transformation.
Main Results:
- Developed Chia, a comprehensive corpus of annotated clinical trial eligibility criteria.
- The corpus contains 15 entity types and 12 relationship types.
- Eligibility criteria are structured as directed acyclic graphs.
Conclusions:
- Chia serves as a valuable resource for advancing information extraction from clinical trial data.
- Enables the development of machine learning, rule-based, and hybrid approaches.
- Facilitates the creation of database queries from free-text eligibility criteria.
Abstract:
We present Chia, a novel, large annotated corpus of patient eligibility criteria extracted from 1,000 interventional, Phase IV clinical trials registered in ClinicalTrials.gov. This dataset includes 12,409 annotated eligibility criteria, represented by 41,487 distinctive entities of 15 entity types and 25,017 relationships of 12 relationship types. Each criterion is represented as a directed acyclic graph, which can be easily transformed into Boolean logic to form a database query. Chia can serve as a shared benchmark to develop and test future machine learning, rule-based, or hybrid methods for information extraction from free-text clinical trial eligibility criteria.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:00TBase - an Integrated Electronic Health Record and Research Database for Kidney Transplant Recipients
Published on: April 13, 2021
Related Concept Videos
Clinical Trials: Overview
Clinical Trials
There are four phases in a clinical trial. A phase one...
Hazard Ratio
For example, in a clinical trial...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Cardiovascular Drugs: Classification based on Therapeutic Indications
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...