通过随机森林和可解释机器学习在解释性项目响应模型中使用个人和项目级预测器来解释个人对项的响应
Sun-Joo Cho1, Goodwin Amanda1, Jorge Salas1
1https://ror.org/02vm5rt34Vanderbilt University's Peabody College, United States.
Psychometrika
|July 31, 2025
概括
本研究引入了一种混合解释性项目响应模型与随机森林 (RF) 集成,以更好地解释项目响应. EIRM-RF模型有效地捕捉复杂的预测相互作用和非线性,优于独立的RF或EIRM方法.
科学领域:
- 心理测量 心理测量 心理测量
- 机器学习 机器学习
- 教育测量教育的测量
背景情况:
- 项目响应模型 (IRT) 传统上与复杂的预测器相互作用作斗争.
- 随机森林 (RF) 擅长捕捉非线性,但缺乏可解释性.
- 现有的模型可能无法同时充分考虑个人和项目级预测效应.
研究的目的:
- 开发和评估一个混合模型 (EIRM-RF),结合解释性项目响应模型 (EIRM) 和随机森林 (RF).
- 通过建模非线性和相互作用效应来改进对象响应的解释.
- 在这种混合框架内评估可解释机器学习 (ML) 方法的性能.
主要方法:
- 在EIRM框架 (EIRM-RF) 中将射频预测值作为预测器集成.
- 可解释的ML技术的应用:特征重要性,部分依赖图,累积局部效应图和H统计.
- 用经验数据集对阅读理解差异 (数字与纸质) 和模拟研究进行模型比较的插图.
主要成果:
- 与独立的EIRM或RF相比,EIRM-RF模型在解释项目响应方面表现出卓越的性能.
- 可解释的ML方法为EIRM-RF捕获的非线性和相互作用效应提供了洞察力.
- 经验和模拟结果证实了混合方法在建模复杂的预测器关系和随机效应方面的优势.
结论:
- EIRM-RF混合模型提供了一种强大的方法来分析复杂的个体对项目响应数据.
- 将IRT与射频和可解释的ML相结合,可以提高模型的准确性和对预测效应的理解.
- 这种混合方法推进了教育测量和心理测量建模领域.
更多相关视频
04:35Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
3.4K
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
8.3K
相关概念视频
Response Surface Methodology
267
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
267
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Mechanistic Models: Compartment Models in Individual and Population Analysis
87
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
87
Variation
7.2K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
7.2K
Survival Tree
160
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
160
Factorial Design
13.3K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.3K
