Related Experiment Video
Updated: Jun 17, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
533
Harnessing LLMs for multi-dimensional writing assessment: Reliability and alignment with human judgments
Xiaoyi Tang1, Hongwei Chen1, Daoyu Lin2
1School of Foreign Studies, University of Science and Technology Beijing, Beijing 100083, China.
Heliyon
|August 8, 2024
Summary
Large Language Models (LLMs) show promise for Automated Essay Scoring (AES), with GPT-4 demonstrating superior accuracy and consistency. Prompt engineering and lower temperature settings enhance LLM reliability in evaluating writing quality.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Artificial Intelligence (AI)
Background:
- Large Language Models (LLMs) are increasingly utilized in Automated Essay Scoring (AES) due to advancements in AI.
- LLMs offer potential for efficient and unbiased writing assessment.
- The reliability and alignment of LLMs with human raters in AES require thorough investigation.
Purpose of the Study:
- To assess the reliability of LLMs in Automated Essay Scoring (AES).
- To explore the impact of prompt engineering, temperature settings, and multi-level rating dimensions on LLM scoring performance.
- To evaluate the alignment of LLM scores with human evaluations.
Main Methods:
- Investigated the effect of prompt engineering strategies (criteria and sample-referenced justification) on LLM performance.
- Analyzed the influence of temperature settings on LLM output consistency.
- Evaluated LLM performance across multiple writing dimensions (Ideas, Organization) using Quadratic Weighted Kappa (QWK).
Main Results:
- Prompt engineering significantly improved LLM reliability, with GPT-4 outperforming GPT-3.5 and Claude 2.
- Lower temperature settings resulted in LLM scores more consistent with human evaluations.
- GPT-4 demonstrated strong performance in 'Ideas' (QWK=0.551) and 'Organization' (QWK=0.584) with optimized prompts.
Conclusions:
- LLMs, particularly GPT-4, show significant potential for reliable and accurate Automated Essay Scoring.
- Careful prompt engineering and temperature control are crucial for maximizing LLM performance and fairness in AES.
- Findings suggest LLMs can enhance writing instruction and feedback in the AI-driven educational landscape.
Keywords:
Automated essay scoring (AES)Generative pre-trained transformer (GPT)Large language models (LLMs)Multi-dimensional writing assessmentPrompt engineeringMore Related Videos
Related Concept Videos
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Guidelines for Writing Outcome
2.7K
When developing expected outcomes for a patient care plan, the nurse should adhere to the following recommendations:
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
2.7K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Group Design
8.9K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
8.9K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Measures of Intelligence
7.1K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.1K

