Biomedical Text Categorization Based on Ensemble Pruning and Optimized Topic Modelling
1Celal Bayar University, Department of Software Engineering, 45400 Turgutlu, Manisa, Turkey.
Computational and Mathematical Methods in Medicine
|August 25, 2018
Summary
This study introduces a novel swarm-optimized topic modeling approach for text categorization, improving Latent Dirichlet allocation (LDA) performance. The method enhances predictive accuracy in biomedical text classification compared to traditional LDA and other methods.
Area of Science:
- Computational linguistics
- Machine learning
- Bioinformatics
Background:
- Text mining, encompassing information retrieval, extraction, and categorization, is crucial in various research fields.
- Latent Dirichlet allocation (LDA) is a powerful tool for text categorization, but its performance is highly dependent on accurate parameter estimation.
- Existing methods for LDA parameter tuning and ensemble classifier construction can be further optimized for improved predictive performance.
Purpose of the Study:
- To propose an efficient multiple classifier system for text categorization using swarm-optimized topic modeling.
- To enhance the performance of Latent Dirichlet allocation (LDA) by optimizing its parameters using swarm intelligence.
- To develop a hybrid ensemble pruning approach that combines diversity measures and clustering for superior predictive accuracy.
Main Methods:
- Swarm optimization algorithms (e.g., genetic algorithms, particle swarm optimization) were used to estimate LDA parameters, including the number of topics.
- A hybrid ensemble pruning technique was developed, combining four diversity measures (disagreement, Q-statistics, correlation coefficient, double fault) and swarm intelligence-based clustering.
- Classifiers were clustered based on diversity, and the best-performing classifier from each cluster was selected for the final system.
Main Results:
- Swarm-optimized LDA demonstrated superior predictive performance over conventional LDA on five biomedical text benchmarks.
- The proposed multiple classifier system significantly outperformed traditional classification algorithms, ensemble learning, and existing ensemble pruning methods.
- Evaluation included comparisons of various metaheuristic algorithms for both swarm-optimized LDA and ensemble pruning.
Conclusions:
- Swarm-optimized topic modeling provides a robust enhancement for LDA in text categorization tasks.
- The proposed hybrid ensemble pruning method effectively constructs diverse and high-performing multiple classifier systems.
- This approach offers a significant advancement in the field of text mining, particularly for biomedical applications.
Related Concept Videos
How Data are Classified: Categorical Data
44.8K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
44.8K
Optimal Foraging
13.9K
How animals obtain and eat their food is called foraging behavior. Foraging can include searching for plants and hunting for prey and depends on the species and environment.
13.9K
Optimization Problems
75
Optimization problems often involve identifying maximum or minimum values under specific constraints. A well-known example is determining the longest horizontal pipe that can be moved around a right-angled corner, where a 3-meter-wide hallway meets a 2-meter-wide hallway. This scenario, common in architectural design and industrial transport, can be understood conceptually through geometric and trigonometric reasoning.To visualize the problem, consider the pipe as a straight line that touches...
75
Stereotype Content Model
15.5K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.5K
Optimal Arousal Theory
858
The optimal arousal theory suggests that performance is maximized when an individual experiences a moderate level of arousal. This theory is closely tied to the Yerkes-Dodson law, which illustrates an inverted U-shaped relationship between arousal and performance. The law, formulated by psychologists Robert Yerkes and John Dodson, implies an ideal arousal level for optimal performance, and deviations from this level can lead to declines in effectiveness.
Inverted U-Shaped Performance Curve
The...
Inverted U-Shaped Performance Curve
The...
858
Optimizing Chromatographic Separations
1.0K
Optimizing chromatographic separations is crucial for obtaining clean separations in a minimum amount of time. Optimization is required for several factors, including kinetic effects related to band broadening, plate height, capacity factor, and separation factor.
Band broadening refers to spreading solute bands as they travel through the column. This broadening can impact resolution. Plate height (H) represents the length required for one theoretical plate. A lower plate height corresponds to...
Band broadening refers to spreading solute bands as they travel through the column. This broadening can impact resolution. Plate height (H) represents the length required for one theoretical plate. A lower plate height corresponds to...
1.0K


