利用LLM进行多维写作评估:可靠性和与人类判断一致
Xiaoyi Tang1, Hongwei Chen1, Daoyu Lin2
1School of Foreign Studies, University of Science and Technology Beijing, Beijing 100083, China.
Heliyon
|August 8, 2024
概括
大型语言模型 (LLM) 显示了自动化作文评分 (AES) 的前景,GPT-4显示出卓越的准确性和一致性. 快速的工程和较低的温度设置提高了LLM在评估写作质量的可靠性.
科学领域:
- 自然语言处理自然语言处理.
- 计算语言学 计算语言学
- 人工智能 (AI) 是一种人工智能.
背景情况:
- 由于人工智能的进步,大型语言模型 (LLM) 越来越多地被用于自动化作文评分 (AES).
- LLM提供了高效和公正的写作评估的潜力.
- 在AES中,LLM与人类评分器的可靠性和对齐性需要进行彻底的调查.
研究的目的:
- 评估LLMs在自动化作文评分 (AES) 的可靠性.
- 探索快速工程,温度设置和多层次评级维度对LLM评分表现的影响.
- 评估LLM成绩与人类评价的一致性.
主要方法:
- 研究了快速工程策略 (标准和样本参考证明) 对LLM绩效的影响.
- 分析了温度设置对LLM输出一致性的影响.
- 通过使用二次加权卡帕 (QWK) 对多个写作维度 (想法,组织) 的LLM绩效进行评估.
主要成果:
- 快速工程显著提高了LLM的可靠性,GPT-4的性能优于GPT-3.5和Claude 2.
- 较低的温度设置导致LLM得分与人类评估更一致.
- 通过优化提示,GPT-4在"想法" (QWK=0.551) 和"组织" (QWK=0.584) 中表现出强的表现.
结论:
- 法学学位,特别是GPT-4,显示出可靠和准确的自动化作文评分的巨大潜力.
- 谨慎的快速工程和温度控制对于最大限度地提高AES的LLM性能和公平性至关重要.
- 研究结果表明,LLM可以在人工智能驱动的教育环境中增强写作指令和反.
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Guidelines for Writing Outcome
2.7K
When developing expected outcomes for a patient care plan, the nurse should adhere to the following recommendations:
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
2.7K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Group Design
8.9K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
8.9K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Measures of Intelligence
7.1K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.1K


