Related Experiment Video
Updated: Jun 6, 2025

A High Throughput MHC II Binding Assay for Quantitative Analysis of Peptide Epitopes
Published on: March 25, 2014
Standardizing Free-Text Data Exemplified by Age and Data-Location Fields in the Immune Epitope Database
Sebastian Duesing1, Jason Bennett1, James A Overton2
1La Jolla Institute For Allergy & Immunology.
Background:
While unstructured data, such as free text, constitutes a large amount of publicly available biomedical data, it is underutilized in automated analyses due to the difficulty of extracting meaning from it. Normalizing free-text data, i.e., removing inessential variance, enables the use of structured vocabularies like ontologies to represent the data and allow for harmonized queries over it. This paper presents an adaptable tool for free-text normalization and an evaluation of the application of this tool to two different sets of unstructured biomedical data curated from the literature in the Immune Epitope Database (IEDB): age and data-location.
Results:
Free text entries for the database fields for subject age (4095 distinct values) and publication data-location (251,810 distinct values) in the IEDB were analyzed. Normalization was performed in three steps, namely character normalization, word normalization, and phrase normalization, using generalizable rules developed and applied with the tool presented in this manuscript. For the age dataset, in the character stage, the application of 21 rules resulted in 99.97% output validity; in the word stage, the application of 94 rules resulted in 98.06% output validity; and in the phrase stage, the application of 16 rules resulted in 83.81% output validity. For the data-location dataset, in the character stage, the application of 39 rules resulted in 99.99% output validity; in the word stage, the application of 187 rules resulted in 98.46% output validity; and in the phrase stage, the application of 12 rules resulted in 97.95% output validity.
Conclusions:
We developed a generalizable approach for normalization of free text as found in database fields with content on a specific topic. Creating and testing the rules took a one-time effort for a given field that can now be applied to data as it is being curated. The standardization achieved in two datasets tested produces significantly reduced variance in the content which enhances the findability and usability of that data, chiefly by improving search functionality and enabling linkages with formal ontologies.
More Related Videos
08:51Identification of Mouse and Human Antibody Repertoires by Next-Generation Sequencing
Published on: March 15, 2019
08:26The Isolation, Differentiation, and Quantification of Human Antibody-secreting B Cells from Blood: ELISpot as a Functional Readout of Humoral Immunity
Published on: December 14, 2016
Related Concept Videos
Cross-reactivity
Immunoprecipitation
Chromatin Immunoprecipitation
Chromatin immunoprecipitation, also known as ChIP, is used to study protein-DNA or...