用于预测建模和分析员工 attrition 和 retention 的数据集
Haya Alqahtani1, Hana Almagrabi1, Amal Alharbi1
1Department of Information Systems, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah, Saudi Arabia.
Data in brief
|December 3, 2025
概括
本研究分析了影响沙特阿拉伯私营部门员工 attrition 的因素. 了解这些驱动因素有助于组织制定有效的员工管理策略和减少流动.
科学领域:
- 商业和管理的管理.
- 数据科学数据科学数据科学
- 人力资源 人力资源 人力资源
背景情况:
- 在私营部门,员工 attrition 是一个重大的挑战.
- 了解消耗驱动因素对于组织稳定性和生产力至关重要.
- 沙特阿拉伯的私营部门面临着独特的劳动力动力.
研究的目的:
- 确定影响沙特阿拉伯私营部门员工 attrition 的关键因素.
- 为分析员工流动提供数据集.
- 支持制定明智的劳动力管理策略.
主要方法:
- 在线调查分发给沙特阿拉伯私营部门的1191名参与者.
- 收集关于人口统计,与工作相关的,以及心理/满意度变量的数据.
- 机器学习分类技术应用于具有34个属性和1191行数据集.
主要成果:
- 确定影响员工磨损的主要因素.
- 数据集可以对员工流动进行复杂的分析.
- 洞察工作满意度与劳动疲劳之间的关系.
结论:
- 数据集是了解和预测员工流动的宝贵资源.
- 调查结果可以为沙特阿拉伯私营部门的劳动力管理政策提供信息.
- 可以使用提供的属性和目标变量开发预测模型.
相关概念视频
Survival Tree
362
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
362
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Regression Analysis
7.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.7K
Applications of Life Tables
307
Life tables are versatile across various fields, providing a quantitative basis for analyzing mortality and survival rates. Whether used by demographers, actuaries, epidemiologists, or sociologists, life tables offer valuable insights into the dynamics of life and death, facilitating informed decisions in public health, insurance, conservation, and beyond. Their broad applicability highlights the interconnectedness of demographic data with practical outcomes in everyday life and strategic...
307
Hazard Rate
375
The hazard rate, also known as the hazard function or failure rate, is a statistical measure used to describe the instantaneous rate at which an event occurs, given that the event has not yet happened. From a probabilistic perspective, it represents the likelihood that a subject will experience the event in a very small time interval, conditional on surviving up to the beginning of that interval. In terms of frequency, the hazard rate can be viewed as the ratio of the number of events to the...
375
Prediction Intervals
3.1K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.1K

