Related Experiment Video
Updated: Jan 30, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Transformer-based relation extraction and concept normalization using an annotated clinical trials corpus
Leonardo Campillos-Llanos1, Ana Valverde-Mateos2, Adrián Capllonch-Carrión3
1ILLA, CCHS CSIC, Madrid, 28037, Spain. leonardo.campillos@csic.es.
None:
Healthcare professionals manually review electronic health records to select patients who meet eligibility criteria of clinical trials. Natural language processing offers a complement for this task, although few initiatives exist in Spanish. We present version 3 of the CT-EBM-SP corpus of 1200 clinical trials (292173 tokens), annotated with 23 entity types and 18 relation types, covering Unified Medical Language System (UMLS) semantic groups, drug-related information, temporal data, and negation/speculation. We encoded 11 attributes (e.g., event temporality and experiencer status) and normalized entities to UMLS Concept Unique Identifiers. The corpus contains 87037 entities, including nested and discontinuous entities, 16597 attributes and 68206 relationships. Inter-annotator agreement (IAA) achieved average F1 values of 0.861 (entities), 0.810 (attributes), and 0.791 (relations). 81.75% of entities were normalized (IAA: F1 = 0.966). We benchmarked this dataset by fine-tuning Transformer models for relation extraction (RE) and medical concept normalization (MCN). In RE, the average F1 ranged from 0.858 to 0.879, and for MCN, the accuracy at rank 1 was 0.896. The corpus and models are publicly available.
Related Concept Videos
Clinical Trials
There are four phases in a clinical trial. A phase one...
Clinical Trials: Overview
Relation of DFT to z-Transform
To understand how the DFT works, it's helpful to consider the z-transform, which is a method for representing discrete sequences in the complex frequency domain. The z-transform involves summing the...
Statistical Software for Data Analysis and Clinical Trials
Acid and Bases: Ka, pKa, and Relative Strengths
Genome Annotation and Assembly

