使用DeepInsight进行零膨胀的高维组合数据.
1Department of Statistics, Kyungpook National University, Daegu, South Korea.
PloS one
|April 16, 2025
概括
这项研究引入了一种新的方法来分析复杂的微生物组数据,解决零通货膨胀和高维度问题. 这种方法可以提高儿童炎症性肠病 (IBD) 等疾病的诊断准确度.
科学领域:
- 微生物组研究 微生物组研究
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 人类微生物组研究因下一代测序 (NGS) 和高通量测序 (HTS) 等先进的测序技术而扩大.
- 微生物组数据分析经常面临零通胀和高维度的挑战,复杂化了构成性解释.
- 操作分类单位 (OTU) 和扩展序列变异 (ASV) 是微生物组研究中常用的数值代理.
研究的目的:
- 开发一种分析零膨胀和高维微生物组数据的方法,并进行组合解释.
- 解决现有方法在准确分析复杂微生物组数据集方面的局限性.
主要方法:
- 将平方根转换应用于组合微生物组数据,以解决零通胀问题,将其映射到超球上.
- 修改了使用卷积神经网络 (CNN) 的DeepInsight图像生成方法,以适应高维数据的超球空间.
- 引入了一种策略,通过添加一个小值来区分零膨胀数据中的真零和假零.
主要成果:
- 提出的方法在分析儿科炎性肠病 (IBD) 便样本数据方面表现出有效性.
- 实现了0.847的曲线下面积 (AUC),超过了之前研究的0.83.3的AUC.
结论:
- 开发的方法成功地处理零膨胀和高维微生物组数据.
- 这种方法通过增强微生物组数据分析,为IBD等疾病提供了更好的诊断潜力.
相关概念视频
Collisions in Multiple Dimensions: Problem Solving
3.5K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
3.5K
Collisions in Multiple Dimensions: Introduction
4.4K
It is far more common for collisions to occur in two dimensions; that is, the initial velocity vectors are neither parallel nor antiparallel to each other. Let's see what complications arise from this. The first idea is that momentum is a vector. Like all vectors, it can be expressed as a sum of perpendicular components (usually, though not always, an x-component and a y-component, and a z-component if necessary). Thus, when the statement of conservation of momentum is written for a...
4.4K
Outliers and Influential Points
3.9K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
3.9K
Dimensional Analysis
809
Dimensional analysis is a powerful tool that is used in physics and engineering to understand and predict the behavior of physical systems. The basic idea behind dimensional analysis is to express physical quantities in terms of fundamental dimensions such as the mass, length, and time. Derived dimensions like the velocity, acceleration, and force are derived from the combinations of these fundamental dimensions.
Dimensional analysis allows us to analyze and compare physical quantities on a...
Dimensional analysis allows us to analyze and compare physical quantities on a...
809
Upsampling
168
Managing signal sampling rates is essential in digital signal processing to maintain signal integrity. A decimated signal, characterized by a reduced frequency range due to its lower sampling rate, can be upsampled by inserting zeros between each sample. This upsampling process expands the original spectrum and introduces repeated spectral replicas at intervals dictated by the new Nyquist frequency. To refine this zero-inserted sequence, it is passed through a lowpass filter with a cutoff...
168
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K


