Renormalized Mutual Information for Artificial Scientific Discovery.
Leopoldo Sarra1, Andrea Aiello1, Florian Marquardt1,2
1Max Planck Institute for the Science of Light, 91058 Erlangen, Germany.
Physical Review Letters
|June 10, 2021
Summary
We developed a new method to measure how continuous variables are related, even when one perfectly predicts the other. This helps in feature extraction and understanding complex systems.
Area of Science:
- Information theory
- Machine learning
- Statistical physics
Background:
- Estimating dependence between continuous random variables is crucial for data analysis.
- Existing methods struggle when one variable deterministically depends on another, a common scenario in feature extraction.
- Discovering collective variables and analyzing information flow in neural networks are key challenges.
Purpose of the Study:
- To develop a robust method for estimating mutual information between continuous random variables.
- To address the challenge of deterministic relationships in dependence estimation.
- To provide a tool for feature extraction and understanding complex systems.
Main Methods:
- Derivation of a renormalized mutual information measure.
- Application to scenarios with deterministic dependencies between continuous random variables.
- Utilizing the method for feature extraction and system analysis.
Main Results:
- A well-defined renormalized mutual information was derived.
- The method successfully estimates dependence even with deterministic relationships.
- The approach facilitates the discovery of collective variables and analysis of information flow.
Conclusions:
- The renormalized mutual information offers a powerful tool for analyzing complex dependencies.
- This method enhances artificial scientific discovery by enabling collective variable identification.
- It improves the analysis of information flow in artificial neural networks.
Related Concept Videos
Correlation of Experimental Data
358
Dimensional analysis simplifies complex physical problems and guides experimental investigations, but it does not provide complete solutions. It identifies the dimensionless groups that influence a phenomenon, but experimental data is needed to establish the specific relationships and validate theoretical predictions.
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...
358
Sign Test for Matched Pairs
254
The sign test for matched pairs offers a robust method for comparing two paired samples, often for the effects of an intervention in one of them. This method is very useful in situations where the underlying distribution of the data is unknown. The test compares two related samples—often pre- and post-treatment measurements on the same subjects—to determine if there are significant differences in their median values.
To conduct the sign test, we first calculate the differences in...
To conduct the sign test, we first calculate the differences in...
254
Statistical Significance
20.6K
Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
20.6K
Quantifying and Rejecting Outliers: The Grubbs Test
2.9K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.9K
Unusual Results
3.6K
Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
3.6K


