Related Experiment Video
Updated: Apr 3, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
SORTA: a system for ontology-based re-coding and technical annotation of biomedical phenotype data.
Chao Pang1, Annet Sollie2, Anna Sijtsma3
1University of Groningen, University Medical Centre Groningen, Genomics Coordination Centre, Department of Genetics, Groningen, The Netherlands, University of Groningen, University Medical Center Groningen, Department of Epidemiology, Groningen, The Netherlands and.
Standardizing biomedical data is crucial for research. The SORTA system automates encoding free text or local codes to formal ontologies like SNOMED CT, saving time and improving data quality.
Area of Science:
- Biomedical Informatics
- Data Science
- Biobanking
Background:
- Standardizing biomedical data semantics, such as phenotypes, is essential for comparative and integrative analyses.
- Retrospective data standardization is often necessary due to varied data collection protocols.
- Manual data curation for matching to ontologies like SNOMED CT, ICD-10, and HPO is time-consuming.
Purpose of the Study:
- To develop SORTA, a computer-aided system to mechanize the encoding of free text or locally coded biomedical data to formal coding systems or ontologies.
- To evaluate SORTA's applicability and efficiency in real-world use cases.
Main Methods:
- SORTA matches original data values to target coding systems (Excel, OWL, OBO formats).
- It employs Lucene and n-gram based matching algorithms for semi-automatic candidate code shortlisting.
- The system can learn from human expert matches and handles semicolon-delimited data input.
Main Results:
- In the LifeLines biobank, SORTA recoded 90,000 free text values to Metabolic Equivalent of Task (MET) codes with 0.97 precision and 0.98 recall.
- In the CINEAS system, SORTA mapped to Human Phenotype Ontology (HPO), achieving 0.58 precision and 0.45 recall.
- Users reported SORTA as a significant time saver and quality improvement, reducing human error.
Conclusions:
- SORTA effectively automates and streamlines the data (re)coding process.
- The system demonstrates significant potential to aid numerous projects requiring biomedical data standardization.
- SORTA enhances data quality and accelerates research by reducing manual curation efforts.
More Related Videos
09:10A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
Published on: May 22, 2018
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Genome Annotation and Assembly
Genomics
Pedigree Analysis