Related Experiment Video
Updated: Jul 20, 2025

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
An open source corpus and automatic tool for section identification in Spanish health records
Iker de la Iglesia1, María Vivó2, Paula Chocrón2
1HiTZ Basque Center for Language Technology Faculty of Engineering Bilbao University of the Basque Country (UPV/EHU), Spain(1).
This study introduces a Spanish open-source dataset for building and evaluating automatic section identification systems for Electronic Clinical Narratives (ECNs). The developed B2 metric and fine-tuned language model improve system performance, even in data-scarce situations.
Area of Science:
- Natural Language Processing
- Medical Informatics
- Computational Linguistics
Background:
- Electronic Clinical Narratives (ECNs) contain vital health information but lack sufficient open-source data.
- The structural heterogeneity of ECNs, from structured headings to unstructured notes, hinders automatic system development and evaluation.
- Developing automated systems for ECNs is crucial for extracting and utilizing patient data effectively.
Purpose of the Study:
- To provide a Spanish open-source dataset for developing and evaluating automatic section identification systems for ECNs.
- To design and implement a novel evaluation metric (B2) tailored for clinical narrative section identification.
- To create a fine-tuned language model optimized for this specific task.
Main Methods:
- Annotation of a corpus of Spanish clinical progress notes into seven major section types.
- Assessment of existing metrics and definition of a new B2 metric for improved task-specific evaluation.
- Development and release of a baseline language model fine-tuned for section identification.
Main Results:
- The open-source annotated corpus, evaluation script, and baseline model are publicly available.
- The baseline model achieved an average B2 score of 71.3 on the open-source dataset.
- The model demonstrated competence in data scarcity scenarios, achieving an average B2 of 67.0.
Conclusions:
- Automatic section identification in unstructured clinical narratives is feasible with adequate data and evaluation metrics.
- This contribution aims to accelerate the development of more robust and accurate systems for processing ECNs.
- The open-sourced resources will foster further research and innovation in clinical text analysis.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
Related Concept Videos
Purpose of Health Records II
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Methods of Documentation VII: EMR
Nursing Clinical Information System
A Nursing Clinical Information System (NCIS) is a specialized type of healthcare information system tailored to meet the unique needs of nursing practice. It incorporates the principles of nursing informatics to streamline information management and improve the quality of care delivery.
Critical attributes of NCIS include:
Statistical Software for Data Analysis and Clinical Trials
Methods of Documentation I: Source-Oriented Records
In an SOR, each discipline involved in patient care maintains a separate medical record section. This record-keeping method enables easy tracking of patient progress and ensures healthcare staff have access to up-to-date information.
Key Attributes include the following: