Small Stochastic Data Compactification Concept Justified in the Entropy Basis
Viacheslav Kovtun1, Elena Zaitseva2, Vitaly Levashenko2
1Internet of Things Group, Institute of Theoretical and Applied Informatics Polish Academy of Sciences, Bałtycka 5, 44-100 Gliwice, Poland.
Entropy (Basel, Switzerland)
|December 23, 2023
Summary
This study introduces a novel method for data compactification using relative entropy, effectively reducing data dimensionality while preserving information. The approach proves stable and efficient for assessing data reliability and handling stochastic parameters.
Area of Science:
- Data Science
- Machine Learning
- Information Theory
Background:
- Data dimensionality is a significant challenge in machine learning, impacting tasks like classification and clustering.
- Data compactification aims to reduce dimensionality while minimizing information loss, a process complicated by stochastic parameters.
Purpose of the Study:
- To propose a new model for structured stochastic data collections using relative entropy.
- To develop an iterative procedure for compacting such data by maximizing relative entropy.
- To assess the compactification procedure's effectiveness and its relevance to data reliability.
Main Methods:
- Modeling structured stochastic data collections in terms of relative entropy.
- Formalizing compactification as an iterative procedure maximizing relative entropy of data projections.
- Developing an approximation for the relative entropy function to reduce computational complexity.
- Assessing compactification using information capacity and information loss metrics.
Main Results:
- A stable and efficient iterative procedure for data compactification of stochastic data.
- Demonstrated effectiveness of the proposed method compared to Principal Component Analysis and Random Projection.
- The proposed metrics are relevant for assessing data reliability and completeness.
Conclusions:
- The proposed relative entropy-based compactification method offers a robust solution for high-dimensional stochastic data.
- This approach enhances data management and reliability assessment in machine learning.
- The method shows superior stability and efficiency over existing techniques.
Related Concept Videos
Entropy
30.2K
Salt particles that have dissolved in water never spontaneously come back together in solution to reform solid particles. Moreover, a gas that has expanded in a vacuum remains dispersed and never spontaneously reassembles. The unidirectional nature of these phenomena is the result of a thermodynamic state function called entropy (S). Entropy is the measure of the extent to which the energy is dispersed throughout a system, or in other words, it is proportional to the degree of disorder of a...
30.2K
Third Law of Thermodynamics
18.9K
A pure, perfectly crystalline solid possessing no kinetic energy (that is, at a temperature of absolute zero, 0 K) may be described by a single microstate, as its purity, perfect crystallinity,and complete lack of motion means there is but one possible location for each identical atom or molecule comprising the crystal (W = 1). According to the Boltzmann equation, the entropy of this system is zero.
18.9K
Entropy and the Second Law of Thermodynamics
2.8K
The second law of thermodynamics can be stated quantitatively using the concept of entropy. Entropy is the measure of disorder of the system.
The relation between entropy and disorder can be illustrated with the example of the phase change of ice to water. In ice, the molecules are located at specific sites giving a solid state, whereas, in a liquid form, these molecules are much freer to move. The molecular arrangement has therefore become more randomized. Although the change in average...
The relation between entropy and disorder can be illustrated with the example of the phase change of ice to water. In ice, the molecules are located at specific sites giving a solid state, whereas, in a liquid form, these molecules are much freer to move. The molecular arrangement has therefore become more randomized. Although the change in average...
2.8K
Compacting Factor test
163
The compacting factor test is a method used to assess the workability of concrete. It is especially suitable for concrete mixes containing aggregates up to one and a half inches in size. This test involves specialized equipment consisting of two truncated cone-shaped hoppers and a cylinder, all with polished interior surfaces to minimize friction.
The procedure begins by placing concrete into the upper hopper without any compaction. Once filled, the bottom door of this hopper is opened,...
The procedure begins by placing concrete into the upper hopper without any compaction. Once filled, the bottom door of this hopper is opened,...
163
Sampling Distribution
12.7K
Given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability distribution of all the measured characteristics is called a sampling distribution. How much the statistic varies from one sample to another is known as the sampling variability of a statistic. You typically measure the sampling variability of a statistic by its standard error. The standard error of the mean is an example...
12.7K
Chebyshev's Theorem to Interpret Standard Deviation
4.2K
Chebyshev’s theorem, also known as Chebyshev’s Inequality, states that the proportion of values of a dataset for K standard deviation is calculated using the equation:
4.2K


