Semantic similarity measure in biomedical domain leverage web search engine

Chi-Huang Chen1, Sheau-Ling Hsieh, Yung-Ching Weng

  • 1Department of Electrical Engineering, National Taiwan University, No. 1, Sec. 4, Roosevelt Road, Taipei, 10617 Taiwan. vinchen@ntu.edu.tw

Summary

This study introduces a novel page-count-based semantic similarity measure for information retrieval and natural language processing. The method leverages web search engine data and machine learning for improved term relatedness in biomedical domains.

Related Concept Videos

Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Bioequivalence: Overview01:16

Bioequivalence: Overview

Pharmaceutical equivalents, by definition, are drug products with the same active ingredient in the same quantities, encapsulated in identical dosage forms, and intended for the same administration routes. These pharmaceutical equivalents are deemed bioequivalent if the bioavailability of the active entity in the drug preparations is similar. Moreover, pharmaceutical equivalents demonstrating bioequivalence are also regarded as therapeutically equivalent. This means that when used as directed,...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...