在实验数据中隐藏着基本变量的自动发现
Boyuan Chen1, Kuang Huang2, Sunand Raghupathi2
1Department of Computer Science, Columbia University, New York, USA. bchen@cs.columbia.edu.
Nature computational science
|January 4, 2024
概括
研究人员开发了一种新原理,可以自动识别从观测数据中隐藏的状态变量. 这种方法确定了变量的数量和性质,推进了基于数据的物理建模.
科学领域:
- 物理 物理学 物理
- 数据科学数据科学数据科学
- 动态系统 动态系统
背景情况:
- 物理定律是状态变量之间的数学关系,对于系统描述至关重要.
- 从观测数据中识别这些隐藏状态变量仍然是自动化方法的一个重大挑战.
- 当前数据驱动的建模通常需要对系统的状态变量有事先的了解.
研究的目的:
- 提出从高维观测数据中确定状态变量的数量和身份的原则.
- 开发一种自动化方法,在没有先前的物理知识的情况下发现隐藏的状态变量.
- 解决长期存在的问题,即仅从观察到的数据中识别状态变量.
主要方法:
- 开发一种新的原理来推断观察到的动态的内在维度.
- 应用一种算法对各种物理动态系统的视频记录.
- 在没有先前存在的物理模型的情况下利用高维观测数据.
主要成果:
- 成功确定了各种观察到的动态系统的内在维度.
- 确定了不同物理现象相关状态变量的候选集.
- 证明了算法的有效性在诸如弹性双和火焰等系统上.
结论:
- 提出的原则为自动化状态变量发现提供了一种可行的方法.
- 这种方法可以显著增强物理和其他科学中的数据驱动建模.
- 能够从原始观测中发现复杂系统的基本描述变量.
更多相关视频
20:24Characterization of Complex Systems Using the Design of Experiments Approach: Transient Protein Expression in Tobacco as a Case Study
Published on: January 31, 2014
16.5K
07:34Large Scale Non-targeted Metabolomic Profiling of Serum by Ultra Performance Liquid Chromatography-Mass Spectrometry UPLC-MS
Published on: March 14, 2013
12.8K
相关概念视频
Correlation of Experimental Data
231
Dimensional analysis simplifies complex physical problems and guides experimental investigations, but it does not provide complete solutions. It identifies the dimensionless groups that influence a phenomenon, but experimental data is needed to establish the specific relationships and validate theoretical predictions.
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...
231
Biostatistics: Overview
246
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
246
Variability: Analysis
143
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
143
Statistical Methods to Analyze Parametric Data: ANOVA
378
Analysis of Variance, or ANOVA, is a powerful statistical technique used to analyze parametric data, primarily in research and experimental studies. It's designed to compare the means of two or more groups, assisting researchers in identifying any significant differences between these group means. There are two main types of ANOVA based on the complexity of the analysis: one-way and two-way.
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
378
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Statistical Analysis: Overview
6.6K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.6K
