Related Experiment Video
Updated: Oct 12, 2025

09:20
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
8.9K
An Efficient Parallelized Ontology Network-Based Semantic Similarity Measure for Big Biomedical Document Clustering
Meijing Li1, Tianjie Chen1, Keun Ho Ryu2,3,4
1College of Information Engineering, Shanghai Maritime University, Shanghai 201306, China.
Computational and Mathematical Methods in Medicine
|November 19, 2021
Summary
This study introduces a parallelized semantic similarity method using Hadoop MapReduce for analyzing large biomedical text datasets. The approach efficiently processes big data, overcoming limitations of traditional methods for semantic mining.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Data Science
Background:
- Semantic mining of large biomedical text corpora presents significant computational challenges.
- Existing ontology-based semantic similarity calculations are often too complex for big data analysis.
- Scalability and efficiency are critical for processing extensive biomedical document collections.
Purpose of the Study:
- To develop a parallelized semantic similarity measurement method for big biomedical text data.
- To address the limitations of traditional methods in handling large-scale datasets.
- To enable efficient semantic analysis and clustering of biomedical documents.
Main Methods:
- Document preprocessing and semantic feature extraction.
- Ontology network structure-based semantic similarity calculation within the Hadoop MapReduce framework.
- Document clustering using generated semantic similarity measures.
Main Results:
- Traditional methods fail with datasets exceeding ten thousand biomedical documents.
- The proposed parallelized method demonstrates efficiency and accuracy on large datasets.
- The approach exhibits high parallelism and scalability for big data analysis.
Conclusions:
- The Hadoop MapReduce-based parallelized semantic similarity method effectively handles big biomedical text data.
- This approach overcomes the scalability and complexity issues of traditional semantic analysis techniques.
- The method provides a robust solution for semantic mining and clustering in large biomedical corpora.
Related Concept Videos
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Bioequivalence: Overview
1.3K
Pharmaceutical equivalents, by definition, are drug products with the same active ingredient in the same quantities, encapsulated in identical dosage forms, and intended for the same administration routes. These pharmaceutical equivalents are deemed bioequivalent if the bioavailability of the active entity in the drug preparations is similar. Moreover, pharmaceutical equivalents demonstrating bioequivalence are also regarded as therapeutically equivalent. This means that when used as directed,...
1.3K

