建立和验证甲状腺结节风险因素的多变量物流模型,使用拉索回归查
Jianning Liu1,2, Zhuoying Feng3, Ru Gao1,2
1Center for Endemic Disease Control, Chinese Center for Disease Control and Prevention, Harbin Medical University, Harbin, Heilongjiang, China.
Frontiers in endocrinology
|April 17, 2024
概括
晚年,女性性别,超重状态,禁食葡萄糖受损和脂质失调是发展甲状腺结节的关键危险因素. 这项研究确定了甲状腺结节疾病预测的关键因素.
科学领域:
- 内分泌学 在内分泌学.
- 公共卫生 公共卫生
- 医疗信息学 医疗信息学
背景情况:
- 甲状腺结节很常见,识别相关的风险因素对于早期检测和管理至关重要.
- 了解甲状腺结节的患病率和预测因素可以为公共卫生策略和临床查协议提供信息.
研究的目的:
- 调查各种人口和临床因素与甲状腺结节的发生之间的关联.
- 开发和验证甲状腺结节疾病的预测风险因子模型.
主要方法:
- 利用最小绝对缩小和选择运算符 (Lasso) 回归来从一个全面的数据集中选择初始变量.
- 采用二进制物流回归来分析确定因素与甲状腺结节患病率之间的关系.
- 使用接收器操作特征 (ROC) 曲线的曲线下的面积 (AUC) 验证了预测模型.
主要成果:
- 高龄 (OR=1.046),女性性别 (OR=1.709),超重状况 (OR=1.546),禁食葡萄糖受损 (OR=1.590),以及脂质失调 (OR=1.588) 被确定为显著的危险因素 (p<0.05).
- 开发的二进制物流回归模型表现出具有0.68的AUC (95%CI:0.64-0.72) 的预测能力.
结论:
- 高龄,女性性别,超重,禁食葡萄糖受损和脂质障碍是与甲状腺结节疾病相关的重要危险因素.
- 已建立的风险因素模型为预测甲状腺结节的发生提供了基础,有助于临床风险分层.
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K


