Related Experiment Video
Updated: Jun 18, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Exploiting Data Distribution: A Multi-Ranking Approach
Beata Zielosko1, Kamil Jabloński1, Anton Dmytrenko1
1Institute of Computer Science, University of Silesia in Katowice, Bȩdzińska 39, 41-200 Sosnowiec, Poland.
Abstract:
Data heterogeneity is the result of increasing data volumes, technological advances, and growing business requirements in the IT environment. It means that data comes from different sources, may be dispersed in terms of location, and may be stored in different structures and formats. As a result, the management of distributed data requires special integration and analysis techniques to ensure coherent processing and a global view. Distributed learning systems often use entropy-based measures to assess the quality of local data and its impact on the global model. One important aspect of data processing is feature selection. This paper proposes a research methodology for multi-level attribute ranking construction for distributed data. The research was conducted on a publicly available dataset from the UCI Machine Learning Repository. In order to disperse the data, a table division into subtables was applied using reducts, which is a very well-known method from the rough sets theory. So-called local rankings were constructed for local data sources using an approach based on machine learning models, i.e., the greedy algorithm for the induction of decision rules. Two types of classifiers relating to explicit and implicit knowledge representation, i.e., gradient boosting and neural networks, were used to verify the research methodology. Extensive experiments, comparisons, and analysis of the obtained results show the merit of the proposed approach.
Related Concept Videos
Ordinal Level of Measurement
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
Review and Preview
Percentiles are a type of fractile that partition data into...
Ranks
Wilcoxon Rank-Sum Test
Quantifying and Rejecting Outliers: The Grubbs Test
Friedman Two-way Analysis of Variance by Ranks

