Improving Automated Essay Scoring by Prompt Prediction and Matching
Jingbo Sun1, Tianbao Song2, Jihua Song1
1School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China.
Entropy (Basel, Switzerland)
|September 23, 2022
Summary
This study introduces a prompt feature fusion method for automated essay scoring, enhancing natural language processing models. Multi-task learning with auxiliary tasks significantly improves scoring accuracy on the HSK dataset.
Area of Science:
- Natural Language Processing
- Educational Technology
- Artificial Intelligence
Background:
- Automated essay scoring (AES) is a key application of NLP in education.
- Pre-trained models are increasingly used for AES, but prompt feature extraction needs improvement.
- Current methods lack sufficient focus on optimizing prompt features for fine-tuning.
Purpose of the Study:
- To develop an effective prompt feature fusion method for AES.
- To enhance feature extraction from pre-trained encoders for better essay scoring.
- To investigate the impact of multi-task learning on AES performance.
Main Methods:
- A novel prompt feature fusion method was created for fine-tuning.
- Multi-task learning was employed with two auxiliary tasks: prompt prediction and prompt matching.
- The NEZHA pre-trained encoder was utilized and evaluated.
Main Results:
- Both auxiliary tasks individually improved model performance in AES.
- The combination of auxiliary tasks with the NEZHA encoder yielded the best results.
- Quadratic Weighted Kappa improved by 2.5% and Pearson's Correlation Coefficient by 2% on average.
Conclusions:
- Multi-task learning and prompt feature fusion are effective strategies for enhancing AES.
- The proposed method significantly boosts the performance of pre-trained models in educational applications.
- This research offers a promising direction for more accurate automated essay evaluation.
Keywords:
automated essay scoringhierarchical structure modelmulti-task learningnatural language processingpre-trained language modelMore Related Videos
Related Concept Videos
Reliability and Validity
12.9K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.9K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Predicting Reaction Outcomes
8.6K
Kinetics describes the rate and path by which a reaction occurs. In contrast, thermodynamics deals with state functions and describes the properties, behavior, and components of a system. It is not concerned with the path taken by the process and cannot address the rate at which a reaction occurs. Although it does provide information about what can happen during a reaction process, it does not describe the detailed steps of what appears on an atomic or a molecular level. On the other hand,...
8.6K
Predicting Products: Substitution vs. Elimination
12.0K
When a nucleophile and an alkyl halide react, nucleophilic substitution and β-elimination reactions compete to generate products.
The following factors can influence the mechanisms competing against each other:
The following factors can influence the mechanisms competing against each other:
12.0K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Wilcoxon Signed-Ranks Test for Matched Pairs
202
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
202


