Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Survival Tree01:19

Survival Tree

389
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
389
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

3.6K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.6K
RNA Stability01:53

RNA Stability

35.6K
Intact DNA strands can be found in fossils, while scientists sometimes struggle to keep RNA intact under laboratory conditions. The structural variations between RNA and DNA underlie the differences in their stability and longevity. Because DNA is double-stranded, it is inherently more stable. The single-stranded structure of RNA is less stable but also more flexible and can form weak internal bonds. Additionally, most RNAs in the cell are relatively short, while DNA can be up to 250 million...
35.6K
Improving Translational Accuracy02:07

Improving Translational Accuracy

14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.6K
3.6K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Addressing biases and limitations in feature attribution for circRNA modification profiling.

Briefings in bioinformatics·2026
Same author

Beyond Parametric Assumptions: Unsupervised Methods Enhance DNA Repair Gene Discovery in <i>SMARCAL1</i> Identification.

Journal of clinical oncology : official journal of the American Society of Clinical Oncology·2026
Same author

Beyond linear assumptions in pulmonary function analysis: a methodological critique of veteran health outcome models.

Annals of the American Thoracic Society·2026
Same author

Revisiting AI Interpretability in Precision Oncology: Why Predictive Accuracy Does Not Ensure Stable Feature Importance.

Cancers·2026
Same author

The Diagnostic Trap in Radiation-Induced Mesothelioma: Kinetic-Morphological Decoupling Masks Molecular Aggression.

Cancers·2026
Same author

Beyond predictive accuracy: interpreting feature importances in pulmonary endarterectomy risk models with total correlation analysis.

The European respiratory journal·2025

相关实验视频

Updated: Jan 17, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K

对应:准确性是不够的:稳定性意识特征选择可重复生物标志物发现的重复性生物标志物发现.

Yoshiyasu Takefuji1

  • 1Faculty of Data Science, Musashino University, Tokyo, Japan.

Allergy
|September 23, 2025
PubMed
概括

随机森林模型提供了高精度但不稳定的特征重要性. 像FA,HVGS和斯皮尔曼相关性这样的稳定性意识方法为可复制生物标志物发现提供了更可靠的特征选择.

科学领域:

  • 生物信息学是一种生物信息学.
  • 机器学习 机器学习
  • 计算生物学 计算生物学

背景情况:

  • 随机森林 (RF) 模型被广泛用于高预测准确性.
  • 然而,RF的模型特定特征的重要性可能是不稳定的和误导性的,阻碍可复制生物标志物的发现.

研究的目的:

  • 为了比较机器学习模型的特征选择策略的稳定性和可靠性.
  • 评估特征选择稳定性对预测准确性和生物标志物发现的影响.

主要方法:

  • 我们比较了五种特征选择策略:随机森林 (RF),后勤回归,特征聚合 (FA),高度可变的基因选择 (HVGS) 和斯皮尔曼相关性.
  • 使用前五个特征和删除前两个后的前三个特征评估交叉验证的准确性.
  • 使用了一套过敏基准数据集,包含1万个实例和11个特征.

主要成果:

  • 在前五个特征中,RF实现了近乎完美的精度 (0.9999) ,但在前五个特征被减少时,精度明显下降 (0.8836),排名不稳定.
  • 后勤回归也显示了不稳定的排名.
  • FA,HVGS和斯皮尔曼相关性在前五个特征中实现了近乎完美的精度 (0.9999) 并保持了高精度 (0.9076-0.9116) 在特征减少后保持稳定的排名.
关键词:
生物标志物发现发现功能选择 功能选择随机的森林随机的森林可复制性的可复制性稳定的稳定性 稳定的稳定性

更多相关视频

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.2K
Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data
04:57

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data

Published on: May 16, 2022

17.3K

相关实验视频

Last Updated: Jan 17, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.2K
Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data
04:57

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data

Published on: May 16, 2022

17.3K

结论:

  • 高度的预测准确性并不能保证可靠的特征重要性.
  • 稳定性意识,模型不可知或无监督的特征选择方法对于可重现的生物标志物发现是优越的.
  • 与RF和物流回归相比,FA,HVGS和Spearman相关性为此数据集提供了更稳定和可靠的特征选择.