Related Experiment Video
Updated: Jan 7, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Distributed Data Classification with Coalition-Based Decision Trees and Decision Template Fusion
Katarzyna Kusztal1, Małgorzata Przybyła-Kasperek1,2
1Institute of Computer Science, University of Silesia in Katowice, Bȩdzińska 39, 41-200 Sosnowiec, Poland.
This study introduces a novel framework for distributed data classification, reducing uncertainty by forming data source coalitions. The method enhances decision trees for improved accuracy, outperforming existing approaches.
Area of Science:
- Data Science
- Machine Learning
- Artificial Intelligence
Background:
- Distributed data environments present classification challenges due to inconsistencies and high informational uncertainty across independent sources.
- Existing methods struggle with data dispersion and maintaining interpretability in complex, multi-source datasets.
Purpose of the Study:
- To propose a novel framework for distributed data classification that reduces entropy and improves decision-making accuracy.
- To integrate conflict analysis, coalition formation, decision tree induction, and decision template fusion for robust classification.
Main Methods:
- Utilized Pawlak's conflict model to identify compatible data sources and form coalitions for aggregating complementary information.
- Developed decision tree classifiers within each coalition and employed decision templates for fusing probabilistic outputs from all models.
- Introduced decision trees for enhanced modeling flexibility and interpretability compared to traditional decision rules.
Main Results:
- The proposed coalition-based framework with decision trees consistently outperformed a non-coalition variant and a rule-based approach.
- Performance improvements were particularly notable under moderate data dispersion scenarios.
- Demonstrated enhanced classification accuracy and robustness across diverse benchmark datasets from the UCI repository.
Conclusions:
- The integration of coalition-based modeling with decision trees offers a significant advancement in distributed data classification.
- Decision templates provide an interpretable mechanism for fusing information from multiple data sources.
- The framework effectively addresses informational uncertainty and improves classification performance in distributed environments.
Related Concept Videos
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Classification of Systems-II
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
