Related Experiment Video
Updated: Oct 18, 2025

09:20
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
8.9K
Classifying domain-specific text documents containing ambiguous keywords
Kamran Karimi1, Sergei Agalakov1, Cheryl A Telmer2
1Department of Biological Sciences, University of Calgary, Calgary, AB T2N 1N4, Canada.
Database : the Journal of Biological Databases and Curation
|September 29, 2021
Summary
Automating literature searches for echinoderm species using machine learning classifiers effectively filters irrelevant PubMed results. This approach saves time and improves accuracy, even with ambiguous common names.
Area of Science:
- Marine Biology
- Bioinformatics
- Computational Biology
Background:
- Keyword-based literature searches in databases like PubMed often yield irrelevant results due to ambiguous terminology.
- Manual curation by domain experts is time-consuming and essential for accurate scientific literature filtering.
- Automating this filtering process requires a solution that is fast, handles limited data, and is domain-neutral.
Purpose of the Study:
- To evaluate various classification algorithms for automating the filtering of domain-specific scientific papers.
- To develop a tool for accurately identifying echinoderm species literature from large databases.
- To assess the performance of different machine learning models in reducing irrelevant search results.
Main Methods:
- A keyword-based search was performed on comprehensive databases such as PubMed.
- Several classification algorithms were tested, including Linear, Naïve Bayes, Nearest Neighbor, Tree, SVM, Bagging, AdaBoost, and Neural Network models.
- The performance of these classifiers was evaluated for their effectiveness in filtering irrelevant articles related to echinoderm species.
Main Results:
- The developed classification tool effectively filters irrelevant articles from PubMed searches.
- The approach demonstrates high accuracy in identifying domain-specific papers, even with ambiguous common names.
- The methodology is adaptable to other fields facing similar literature filtering challenges.
Conclusions:
- Machine learning-based classification offers a practical and efficient solution for automating scientific literature filtering.
- The developed tool successfully addresses challenges posed by ambiguous keywords and limited data availability.
- This approach significantly enhances the efficiency and accuracy of literature reviews in specialized scientific domains.
Related Concept Videos
Force Classification
1.8K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.8K
Classification of Systems-I
366
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
366
Classification of Systems-II
259
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
259
Aggregates Classification
416
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
416
Classification of Signals
1.0K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.0K
Gene Families
3.0K
3.0K

