分布式数据分类与基于联盟的决策树和决策模板融合
Katarzyna Kusztal1, Małgorzata Przybyła-Kasperek1,2
1Institute of Computer Science, University of Silesia in Katowice, Bȩdzińska 39, 41-200 Sosnowiec, Poland.
Entropy (Basel, Switzerland)
|December 24, 2025
概括
这项研究引入了分布式数据分类的新框架,通过形成数据源联盟来减少不确定性. 该方法增强了决策树,提高了准确性,优于现有的方法.
科学领域:
- 数据科学数据科学数据科学
- 机器学习 机器学习
- 人工智能的人工智能
背景情况:
- 分布式数据环境存在分类挑战,原因是独立来源之间存在不一致性和高度信息不确定性.
- 现有的方法在复杂的多源数据集中扎着数据分散和保持可解释性.
研究的目的:
- 为分布式数据分类提出一个新的框架,以减少和提高决策准确性.
- 整合冲突分析,联盟形成,决策树诱导和决策模板融合以进行强大的分类.
主要方法:
- 利用Pawlak的冲突模型来识别兼容的数据源,并形成联盟来聚合补充信息.
- 在每个联盟中开发了决策树分类器,并使用决策模板来融合所有模型的概率输出.
- 与传统决策规则相比,引入了决策树,以提高建模灵活性和可解释性.
主要成果:
- 提出的以决策树为基础的联盟框架始终优于非联盟变体和基于规则的方法.
- 在中度数据分散场景下,性能改善尤其显著.
- 在来自UCI存储库的各种基准数据集中证明了增强的分类准确性和稳定性.
结论:
- 将基于联盟的建模与决策树集成为分布式数据分类提供了显著的进步.
- 决策模板提供了一个可解释的机制,用于融合来自多个数据源的信息.
- 该框架有效地解决了信息不确定性,并改善了分布式环境中的分类性能.
相关概念视频
Classification of Systems-I
533
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
533
Aggregates Classification
947
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
947
Force Classification
2.2K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
2.2K
Classification of Systems-II
445
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
445
How Data are Classified: Categorical Data
42.4K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
42.4K
Decision Making: Traditional Method
5.0K
The process of hypothesis testing based on the traditional method includes calculating the critical value, testing the value of the test statistic using the sample data, and interpreting these values.
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
5.0K

