Related Experiment Video
Updated: Apr 4, 2026

07:11
Examination of Anatomical Features of Retinal Ganglion Cells Under N-methyl-D-aspartic Acid (NMDA)-induced Excitotoxicity
Published on: September 19, 2025
1.2K
Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus
1School of Library and Information Science, Simmons College, Boston, MA, USA.
Journal of Biomedical Informatics
|August 31, 2015
Summary
Researchers created a de-identified medical record corpus for research. This dataset, derived from longitudinal records, aids natural language processing in protecting patient privacy while enabling data analysis.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Health Data Privacy
Background:
- De-identification of medical records is crucial for privacy and research.
- Existing datasets for de-identification research are limited.
- Longitudinal medical records present unique de-identification challenges.
Purpose of the Study:
- To create a de-identified corpus of longitudinal medical records for research.
- To establish a gold standard for medical record de-identification.
- To support the 2014 i2b2/UTHealth natural language processing shared task.
Main Methods:
- De-identification of 1304 longitudinal medical records from 296 patients.
- Application of a broad interpretation of Health Insurance Portability and Accountability Act (HIPAA) guidelines.
- Utilized double-annotation, arbitration, sanity checking, and manual proofreading.
- Automated replacement of private health information with realistic surrogates.
Main Results:
- Achieved an average token-based F1 score of 0.927 for annotator agreement with the gold standard.
- Developed the first corpus of its kind for de-identification research.
- Systems in the 2014 i2b2/UTHealth shared task achieved a mean F-measure of 0.872 and a maximum of 0.964.
Conclusions:
- The developed corpus is a valuable resource for advancing de-identification research.
- The rigorous annotation process ensured high-quality de-identification.
- Facilitated advancements in natural language processing for clinical data privacy.
