Related Experiment Video
Updated: Jun 6, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Semantic similarity measure in biomedical domain leverage web search engine
Chi-Huang Chen1, Sheau-Ling Hsieh, Yung-Ching Weng
1Department of Electrical Engineering, National Taiwan University, No. 1, Sec. 4, Roosevelt Road, Taipei, 10617 Taiwan. vinchen@ntu.edu.tw
This study introduces a novel page-count-based semantic similarity measure for information retrieval and natural language processing. The method leverages web search engine data and machine learning for improved term relatedness in biomedical domains.
Area of Science:
- Natural Language Processing
- Information Retrieval
- Biomedical Informatics
Background:
- Semantic similarity measurement is crucial for Information Retrieval (IR) and Natural Language Processing (NLP).
- Existing semantic similarity measures face challenges in accuracy and robustness, particularly in specialized domains like biomedicine.
- Previous research has explored various methods, but a gap remains for effective, scalable solutions.
Purpose of the Study:
- To propose and evaluate a novel page-count-based semantic similarity measure.
- To apply this measure within the biomedical domain, addressing specific challenges in this field.
- To integrate diverse similarity scores using machine learning for enhanced performance.
Main Methods:
- Utilizing web search engine page counts for querying individual terms (P, Q) and their conjunction (P AND Q).
- Developing a novel approach combining page counts with lexico-syntactic patterns for semantic similarity computation.
- Integrating multiple similarity scores via support vector machines (SVM) to improve measure robustness.
Main Results:
- Achieved a correlation coefficient of 0.798 on the A. Hliaoutakis dataset.
- Obtained a correlation coefficient of 0.705 with physician scores on the T. Pedersen dataset.
- Reached a correlation coefficient of 0.496 with expert scores on the T. Pedersen et al. dataset.
Conclusions:
- The proposed page-count-based semantic similarity measure demonstrates effectiveness, particularly when integrated with lexico-syntactic patterns and machine learning.
- The approach shows promise for improving semantic similarity calculations in biomedical information retrieval and NLP tasks.
- Experimental results validate the robustness and applicability of the developed measure across different datasets and scoring criteria.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Bioequivalence: Overview
Improving Translational Accuracy

