伪造ML:一种伪造机器学习框架,用于控制变量选择和电子健康记录数据中的风险分层
Qi Wang1, Linyan Li2,3, Yi Yang4
1Department of Data Science, City University of Hong Kong, Kowloon, Hong Kong SAR.
NPJ digital medicine
|November 26, 2025
概括
Knockoff-ML是一种新的机器学习框架,有效地从电子健康记录中识别患者的风险因素,以改善临床决策. 它比现有的评分系统提供了更高的预测能力和可解释性.
科学领域:
- 计算生物学和生物信息学
- 临床信息学 临床信息学
- 医疗保健中的机器学习
背景情况:
- 在临床实践中,有效的风险分层对于优化资源配置和患者结果至关重要.
- 电子健康记录 (EHR) 中的机器学习模型有助于风险预测,但往往缺乏临床医生可以解释的决策规则.
- 现有的可解释性指标难以识别影响结果的特定患者特征.
研究的目的:
- 引入Knockoff-ML,一种无模型的机器学习框架,用于同时预测结果和识别风险特征.
- 整合一个仿制框架与预测机器学习算法,以增强变量选择.
- 为了使错误发现率 (FDR) 控制能够识别EHR数据中的复杂,非线性关联.
主要方法:
- 开发了Knockoff-ML,通过将传统的机器学习模型与一个Knockoff框架进行增强.
- 使用FDR控制的变量选择来识别显著的风险特征.
- 在MIMIC-IV数据库上通过模拟和现实应用评估性能.
主要成果:
- 在模拟中控制FDR的同时, Knockoff-ML在识别风险特征方面表现出高的统计能力,优于传统方法.
- 在50591名重症监护室 (ICU) 患者中,确定了与短期和长期死亡相关的重大风险特征.
- 与SOFA和SAPS II评分系统相比,实现了与完整模型相比的预测准确度和更高的预测能力和临床实用性.
结论:
- Knockoff-ML通过识别关键患者风险因素,为临床决策提供了强大且可解释的工具.
- 该框架提高了预测准确性和临床实用性,为患者的治疗结果和医疗保健提供提供了潜在的改善.
- Knockoff-ML 在利用 EHR 数据进行个性化风险分层方面取得了重大进展.
相关概念视频
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
387
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
387
Statistical Software for Data Analysis and Clinical Trials
1.4K
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
1.4K
Randomized Experiments
8.8K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
8.8K
Study Designs in Epidemiology
860
Epidemiological study designs are fundamental tools for investigating the distribution, determinants, and control of health conditions in populations. They help researchers understand the relationships between exposures and outcomes, and they broadly fall into two categories: "observational" and "experimental" studies.
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
860
Statistical Methods for Analyzing Epidemiological Data
885
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
885
Study Design in Statistics
9.9K
A study design is a set of techniques that allow a researcher to collect and analyze data from different variables defined for a specific research problem. Statistics is commonly for effective study design and more robust experiments,
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
9.9K


