在使用医院数据的自动化变量选择中对集群的计算:不同 LASSO 方法的比较
Stella Bollmann1,2, Andreas Groll3, Michael M Havranek4
1Competence Center for Health Data Science, Faculty of Health Sciences and Medicine, University Lucerne, Frohburgstrasse 3, 6002, Lucerne, Switzerland. stella.bollmann@ife.uzh.ch.
BMC medical research methodology
|November 25, 2023
概括
对医院患者数据的多层结构进行核算,可以提高使用最小绝对收缩和选择运算符 (LASSO) 进行预测的一些结果. 这种多层次的方法对于在医疗保健质量评估中进行公正的预测因素选择至关重要.
科学领域:
- 医疗保健服务研究 医疗服务研究
- 生物统计学 生物统计学
- 医疗信息学 医疗信息学
背景情况:
- 自动化特征选择,包括最小绝对收缩和选择操作符 (LASSO),对于预测医疗保健质量结果和风险调整至关重要.
- 现有的 LASSO 方法往往忽视了医院内嵌入的患者数据的等级性质.
研究的目的:
- 为了证明将医院数据的多层结构纳入LASSO.
- 为了比较多层 LASSO 与标准 LASSO 变体的预测性能和变量重要性.
主要方法:
- 在三个患者数据集 (急性心肌梗塞,COPD,中风) 上应用了各种 LASSO 技术,并没有考虑数据的多层结构.
- 利用两个依赖变量 (数值和二进制) 和20倍的子样本程序来评估预测性能和变量重要性.
主要成果:
- 多层 LASSO 改善了数字结果"停留时间"的预测,但对二进制结果"死亡率"没有显示差异.
- 在某些情况下,在多层 LASSO 方法和标准 LASSO 方法之间观察到变量重要性的显著差异.
结论:
- 使用LASSO将多层数据结构集成到自动预测器选择中是可行的,这可能会提高预测性能.
- 考虑多层结构对于无偏见的预测因素选择至关重要,尤其是在考虑医院特定变异时.
相关概念视频
Comparing the Survival Analysis of Two or More Groups
197
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
197
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Statistical Analysis: Overview
6.6K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.6K
Statistical Software for Data Analysis and Clinical Trials
565
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
565


