关于自动评分与大型语言模型的一致性
Mingfeng Xue1, Xingyao Xiao2, Yunting Liu3
1University of North Carolina Greensboro, USA.
Educational and psychological measurement
|February 19, 2026
概括
大型语言模型 (LLM) 显示得分在LLM内部的一致性很高,但在LLM之间的一致性是适度的. 结合LLM输出的投票策略可以提高得分准确度.
科学领域:
- 人工智能的人工智能
- 教育测量教育的测量
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 显示了自动化评分任务的潜力.
- 在LLM评分中的不一致性可能来自模型变化和培训数据差异.
- 了解LLM评分的一致性对于可靠的自动评估至关重要.
研究的目的:
- 调查五个LLM的LLM内部和LLM之间的得分一致性.
- 检查温度设置对LLM得分一致性的影响.
- 评估评分一致性和准确性之间的关系.
- 提出和评估一个投票策略,以改善LLM分数.
主要方法:
- 评估了五个LLM (Claude,DeepSeek,Gemini,GPT,Qwen) 的评分一致性.
- 在不同的温度设置下评估一致性.
- 利用了来自科学教育和ASAP数据集的构建响应项目.
- 在LLMs中实施多数投票策略.
主要成果:
- 无论温度如何,LLM显示出几乎完美的LLM内部的一致性.
- 跨LLM的一致性是适度的,更容易的项目更高.
- 在LLM内部的一致性超过了LLM内部的一致性.
- 在LLM内部的一致性与准确性没有相关性;在LLM内部的一致性显示出正相关性.
- 多数投票通过结合不同的LLM优势来提高得分准确度.
结论:
- LLM提供高的内部评分可靠性,但外部协议不同.
- 跨LLM的一致性比LLM内部的一致性更好地预测得分准确性.
- 组合方法,如多数投票,可以减轻LLM得分不一致性,提高教育评估的准确性.
相关概念视频
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
Improving Translational Accuracy
3.7K
3.7K
Automatic Processing and Automatic Social Behavior
276
Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
276
Language and Cognition
836
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
836
Language Development
973
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
973
Accuracy, limits, and approximation
1.3K
Accuracy, limits, and approximations are common in many fields, especially in engineering calculations. These concepts are imperative for ensuring that a given value is as close as possible to its true value.
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
1.3K
