不确定性建模优于微生物组数据分析的机器学习
Maxwell A Konnaris1, Manan Saxena2, Nicole Lazar3
1Program in Bioinformatics and Genomics, Pennsylvania State University, University Park, PA, USA.
bioRxiv : the preprint server for biology
|September 26, 2025
概括
微生物组测序缺乏总微生物负载数据. 机器学习模型无法准确预测负载,但贝叶斯方法为微生物组分析提供了可靠的解决方案.
科学领域:
- 微生物学 微生物学
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 微生物组测序量化了微生物的相对,而不是绝对丰度.
- 现有的规范化方法依赖于可能引入偏差的假设.
- 直接测量微生物负荷是准确的,但昂贵和不频繁.
研究的目的:
- 评估机器学习在预测微生物负载时仅仅从测序数据的有效性.
- 评估机器学习模型在各种微生物组研究中的通用性.
- 将机器学习方法与处理微生物负载不确定性的替代方法进行比较.
主要方法:
- 组装了"mutt",这是最大的配对测序和微生物负载测量数据库 (35项研究,>15,000个样本).
- 在"mutt"数据库和基准数据集上评估已发布的机器学习模型.
- 实施并比较贝叶斯的部分识别模型来传播规模不确定性.
主要成果:
- 机器学习模型的概括性很差,平均表现比纯粹的基线差.
- 模型失败归因于共变量转移,有限的共享种类,组成差异和预处理变化.
- 在30个基准数据集中,贝叶斯部分识别模型的表现始终优于规范化和机器学习方法.
结论:
- 机器学习方法对于从微生物组测序数据中预测微生物负载是不可靠的.
- 贝叶斯部分识别模型提供了一个原则和可重复的方法,用于计算微生物组推断中的规模不确定性.
- "马特"数据库是评估微生物组分析方法的宝贵资源.
相关概念视频
Uncertainty: Overview
1.6K
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
1.6K
Uncertainty: Confidence Intervals
10.2K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
10.2K
Mechanistic Models: Compartment Models in Individual and Population Analysis
249
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
249
Modern Molecular Taxonomy
590
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
590
Uncertainty in Measurement: Accuracy and Precision
99.8K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
99.8K
Steps in Outbreak Investigation
492
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
492


