Related Experiment Video
Updated: Aug 2, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Comparative analysis of weka-based classification algorithms on medical diagnosis datasets
Yifeng Dou1,2, Wentao Meng1,2
1Network Information Center, Tianjin Baodi Hospital, Tianjin, China.
Machine learning algorithms show promise in aiding medical diagnosis by accurately classifying diseases. Random Forest and Bagging algorithms demonstrated superior performance across various medical datasets, offering valuable insights for clinical decision-making.
Area of Science:
- Medical Informatics
- Machine Learning
- Data Science
Background:
- The proliferation of Big Data and 5G technology has led to a surge in medical data from electronic health records and digitized equipment.
- Hospitals now possess vast databases containing clinical diagnosis and hospital management information.
- Advancements in medical information technology necessitate efficient data analysis methods.
Purpose of the Study:
- To evaluate the classification performance of diverse machine learning algorithms on medical datasets.
- To explore the potential of machine learning in enhancing medical diagnosis accuracy.
- To identify optimal algorithms for specific medical data classification tasks.
Main Methods:
- Utilized four distinct medical classification datasets from the University of California Irvine machine learning repository.
- Developed six categories of classification models using the Weka platform, based on Bayesian, ensemble, rule-based, and tree-based approaches.
- Conducted between-group and within-group experiments to compare algorithm performance.
Main Results:
- Random Forest algorithm exhibited the highest accuracy on the Indian Liver Patient Dataset (ILPD), Cardiotocography (CADG), and Lymphatic Tractography (LYMP) datasets.
- Bagging algorithm outperformed other ensemble methods across multiple metrics, particularly on binary datasets.
- Logistic Model Tree achieved optimal results on the Mammographic dataset (MAGR), while Random Forest excelled on ILPD, CADG, and LYMP.
Conclusions:
- Machine learning algorithms possess significant potential for disease prediction and diagnostic support.
- The findings provide a valuable reference for applying machine learning in clinical decision-making.
- Specific algorithms like Random Forest and Bagging show strong applicability in medical data classification.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Comparing the Survival Analysis of Two or More Groups
Receiver Operating Characteristic Plot
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Classification of Systems-II
Statistical Software for Data Analysis and Clinical Trials

