使用深度强化学习来决定测试长度
James Zoucha1, Igor Himelfarb2, Nai-En Tang2
1University of Northern Colorado, Greeley, CO, USA.
Educational and psychological measurement
|May 7, 2025
概括
深度强化学习 (DRL) 可以优化测试长度,但建议使用当前的手术考试长度. 较短的形式保持了准确性,但没有结构完整性,显示了DRL.
科学领域:
- 教育中的人工智能
- 心理测量和教育测量方法
- 计算优化计算优化
背景情况:
- 优化标准化测试长度对于效率和有效性至关重要.
- 当前的测试构造方法可能无法充分利用先进的计算方法.
- 深度强化学习 (DRL) 为复杂的优化问题提供了一个新的框架.
研究的目的:
- 为了研究DLR在优化测试长度中的有效性,为全国手术检查员委员会进行了第I部分考试.
- 为了确定更短的测试形式是否可以保持心理测量完整性和结构约束.
- 探索DRL对个性化测试和适应性项目选择的潜力.
主要方法:
- 在马尔科夫决策过程中,建模测试形式构造作为组合优化问题.
- 开发和应用DRL算法,从定义的项目库中生成测试表格.
- 根据能力估计的准确性,内容表示和项目难度分布来评估测试表格.
主要成果:
- DRL成功地确定了较短的测试形式,具有可比能力估计准确度.
- 由DRL生成的较短的测试表格并不始终保持关键结构约束.
- 现有的240项测试长度被认为是值得推的,因为它符合约束.
结论:
- DRL是探索测试长度优化的强大工具,但需要仔细考虑结构约束.
- DRL的适应能力使其适合未来的个性化和适应性测试环境.
- 进一步的研究应该集中在扩大项目库和计算资源,以提高DRL性能.
相关概念视频
Reinforcement Schedules
116
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
116
Reinforcement
154
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
154
Timing and Consequences on Behavior
57
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
57


