Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Genetic Lingo01:11

Genetic Lingo

Overview
Gene Evolution - Fast or Slow?02:05

Gene Evolution - Fast or Slow?

The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
Gene Evolution - Fast or Slow?02:05

Gene Evolution - Fast or Slow?

The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Enhancing the quality and trustworthiness of large language model-generated summaries of clinical oncology literature.

JAMIA open·2026
Same author

Impact of Varicocele Embolization on DNA Fragmentation Index (DFI) in Male Infertility.

Journal of vascular and interventional radiology : JVIR·2024
Same author

Supporting the working life exposome: Annotating occupational exposure for enhanced literature search.

PloS one·2024
Same author

Predicting potential target genes in molecular biology experiments using machine learning and multifaceted data sources.

iScience·2024
Same author

Student-Led Individualized Education Programs: A Gateway to Self-Determination.

Language, speech, and hearing services in schools·2024
Same author

Knowledge-enhanced Graph Topic Transformer for Explainable Biomedical Text Summarization.

IEEE journal of biomedical and health informatics·2023

Related Experiment Video

Updated: Jul 13, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

Learning string similarity measures for gene/protein name dictionary look-up using logistic regression.

Yoshimasa Tsuruoka1, John McNaught, Jun'ichi Tsujii

  • 1School of Computer Science, The University of Manchester, Manchester, UK. yoshimasa.tsuruoka@manchester.ac.uk

Bioinformatics (Oxford, England)
|August 19, 2007
PubMed
Summary

Accurate biomedical data integration is challenging due to term variations. A new logistic regression-based string similarity measure improves dictionary look-up accuracy for gene and protein names.

More Related Videos

A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

Related Experiment Videos

Last Updated: Jul 13, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

Area of Science:

  • Biomedical Informatics
  • Bioinformatics
  • Computational Biology

Background:

  • Biomedical data integration faces challenges due to variations in terminology.
  • Exact string matching fails to link biological concepts (e.g., gene IDs) to names because of minor differences.
  • Soft string matching offers a solution by considering name similarity, but its accuracy depends on the chosen measure.

Purpose of the Study:

  • To develop and evaluate a novel, accurate string similarity measure for biomedical data integration.
  • To address the limitations of existing similarity measures in associating biological terms with their corresponding identifiers.

Main Methods:

  • Employed logistic regression to learn an effective string similarity measure from a comprehensive dictionary.
  • Utilized large-scale gene and protein name dictionaries for training and testing the model.

Main Results:

  • The developed logistic regression-based similarity measure significantly outperformed existing measures in dictionary look-up tasks.
  • Demonstrated improved accuracy in identifying correct biological concept IDs despite variations in input names.

Conclusions:

  • Logistic regression provides a robust method for learning accurate string similarity measures.
  • The proposed approach enhances the reliability of biomedical data integration by improving the association of names with identifiers.