在表式临床数据中对流行分类算法的样本大小要求:实证研究
1Department of Biostatistics, School of Public Health, Virginia Commonwealth University, Richmond, VA, United States.
Journal of medical Internet research
|December 17, 2024
概括
为分类算法确定最佳样本大小至关重要. 数据集特征,如类平衡,显著影响稳定性能所需的样本大小,影响研究效率.
科学领域:
- 机器学习 机器学习
- 生物统计学 生物统计学
- 临床信息学 临床信息学
背景情况:
- 分类算法性能高原,需要最佳样本大小的确定.
- 为了高效的研究,平衡性能收益与计算成本至关重要.
研究的目的:
- 为了确定二进制分类算法的最佳样本大小.
- 在各种算法中调查样本大小和数据集特征之间的关系.
主要方法:
- 在16个不同的数据集上评估了4个算法 (XGBoost,随机森林,后勤回归,神经网络).
- 在增加样本大小以适应学习曲线时计算交叉验证的曲线下面面积 (AUC).
- 使用回归模型,量化数据集特征与所需样本大小之间的关系.
主要成果:
- 达到AUC稳定的平均样本大小各不相同:XGBoost (9960),随机森林 (3404),物流回归 (696),神经网络 (12,298).
- 增加的类平衡和数据集复杂性通常分别减少或增加了样本大小要求.
- 少数阶级比例是所有算法的关键预测因素;数据集AUC和非线性对XGBoost,随机森林和神经网络很重要.
结论:
- 对分类算法的最佳样本大小取决于方法和数据集.
- 数据集的特征,如类平衡和特征复杂性,可以预测并潜在地被利用来优化研究研究的样本大小选择.
相关概念视频
Statistical Software for Data Analysis and Clinical Trials
492
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
492
Sample Size Calculation
3.2K
Knowledge of the sample size is the first requirement to conduct random sampling or an experiment. The sample size is the total number of units, observations, or groups (in some cases) used to get the data to estimate a population parameter. As the name suggests, the sample size is that of the sample drawn from the population and differs from the population size.
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
3.2K
Comparing the Survival Analysis of Two or More Groups
149
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
149
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K
How Data are Classified: Categorical Data
31.6K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
31.6K
Survival Tree
58
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
58


