使用新生入学心理调查预测大学生中心理风险:基于LASSO-逻辑回归的机器学习模型
Chang-Zheng Ma1, Fang Xiao1, Jing Zhang1
1School of Mental Health, Bengbu Medical University, Bengbu, Anhui, 233030, China.
Journal of affective disorders
|December 21, 2025
概括
机器学习使用查数据准确预测大学新生的心理风险. 这种模式有助于早期识别和干预面临心理健康挑战的学生.
科学领域:
- 心理学 心理学 心理学
- 机器学习 机器学习
- 公共卫生 公共卫生
背景情况:
- 大学的新生面临着相当大的心理压力.
- 早期发现心理健康风险对于及时干预至关重要.
- 现有的选方法需要提高预测准确度.
研究的目的:
- 应用机器学习来预测大学新生心理风险.
- 开发一个预测模型来识别有风险的学生.
- 提高在高等教育中早期发现心理健康问题的能力.
主要方法:
- 对7211名新生进行心理查 (UPI和SDS尺度).
- 使用LASSO回归来进行特征选择和物流回归来进行模型开发.
- 使用ROC曲线,校准曲线,H-L测试和DCA验证的模型有效性.
主要成果:
- 确定了关键的影响因素:身体和精神的满意度,咨询/精神病史,孕产妇死亡,自杀念头,UPI和SDS分数.
- 获得的AUC为0.807 (培训) 和0.757 (验证).
- 通过校准,H-L测试和DCA证明了良好模型适合性和高临床实用性.
结论:
- 开发了一种机器学习模型,用于预测大学新生在选后的心理风险.
- 模型和名图为早期风险识别提供了有价值的工具.
- 促进针对学生心理健康的有针对性的支持.
相关概念视频
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Surveys
16.6K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
16.6K
Correlations
35.7K
Correlation means that there is a relationship between two or more variables (such as ice cream consumption and crime), but this relationship does not necessarily imply cause and effect. When two variables are correlated, it simply means that as one variable changes, so does the other. We can measure correlation by calculating a statistic known as a correlation coefficient. A correlation coefficient is a number from -1 to +1 that indicates the strength and direction of the relationship between...
35.7K
Longitudinal Research
13.0K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
13.0K
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Regression Analysis
7.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.8K


