Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genome Annotation and Assembly03:36

Genome Annotation and Assembly

22.0K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
22.0K
Genomics02:02

Genomics

41.8K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
41.8K
Pedigree Analysis01:35

Pedigree Analysis

91.0K
Overview
91.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Microencapsulated OPO Enhances Intestinal SCFA Production by Optimizing Lipid Digestion and Regulating Bile Acid Metabolism in Mice.

Foods (Basel, Switzerland)·2026
Same author

A multidisciplinary RNA-guided approach to complement genomic analysis of unsolved patients with an inborn error of immunity.

Frontiers in immunology·2026
Same author

A Foundation Model for Capturing Complexity of Menstrual Health Data.

npj women's health·2026
Same author

Toward scalable early cancer detection: evaluating EHR-based predictive models against traditional screening criteria.

NPJ precision oncology·2026
Same author

The association between chronotype and incident dementia: exploring age, educational-attainment and sex differences.

Epidemiology and psychiatric sciences·2026
Same author

Body Mass Index, Clinical Outcomes, and Mortality in Heart Failure: A Mendelian Randomization Study.

Journal of the American College of Cardiology·2026

Related Experiment Video

Updated: Apr 3, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

8.1K

SORTA: a system for ontology-based re-coding and technical annotation of biomedical phenotype data.

Chao Pang1, Annet Sollie2, Anna Sijtsma3

  • 1University of Groningen, University Medical Centre Groningen, Genomics Coordination Centre, Department of Genetics, Groningen, The Netherlands, University of Groningen, University Medical Center Groningen, Department of Epidemiology, Groningen, The Netherlands and.

Database : the Journal of Biological Databases and Curation
|September 20, 2015
PubMed
Summary

Standardizing biomedical data is crucial for research. The SORTA system automates encoding free text or local codes to formal ontologies like SNOMED CT, saving time and improving data quality.

More Related Videos

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
09:10

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes

Published on: May 22, 2018

10.2K
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.6K

Related Experiment Videos

Last Updated: Apr 3, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

8.1K
A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
09:10

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes

Published on: May 22, 2018

10.2K
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.6K

Area of Science:

  • Biomedical Informatics
  • Data Science
  • Biobanking

Background:

  • Standardizing biomedical data semantics, such as phenotypes, is essential for comparative and integrative analyses.
  • Retrospective data standardization is often necessary due to varied data collection protocols.
  • Manual data curation for matching to ontologies like SNOMED CT, ICD-10, and HPO is time-consuming.

Purpose of the Study:

  • To develop SORTA, a computer-aided system to mechanize the encoding of free text or locally coded biomedical data to formal coding systems or ontologies.
  • To evaluate SORTA's applicability and efficiency in real-world use cases.

Main Methods:

  • SORTA matches original data values to target coding systems (Excel, OWL, OBO formats).
  • It employs Lucene and n-gram based matching algorithms for semi-automatic candidate code shortlisting.
  • The system can learn from human expert matches and handles semicolon-delimited data input.

Main Results:

  • In the LifeLines biobank, SORTA recoded 90,000 free text values to Metabolic Equivalent of Task (MET) codes with 0.97 precision and 0.98 recall.
  • In the CINEAS system, SORTA mapped to Human Phenotype Ontology (HPO), achieving 0.58 precision and 0.45 recall.
  • Users reported SORTA as a significant time saver and quality improvement, reducing human error.

Conclusions:

  • SORTA effectively automates and streamlines the data (re)coding process.
  • The system demonstrates significant potential to aid numerous projects requiring biomedical data standardization.
  • SORTA enhances data quality and accelerates research by reducing manual curation efforts.