Related Experiment Video
Updated: May 7, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating large language models for criterion-based grading from agreement to consistency
Da-Wei Zhang1, Melissa Boey2, Yan Yu Tan2
1Department of Psychology, Jeffrey Cheah School of Medicine and Health Sciences, Monash University Malaysia, Bandar Sunway, 475000, Malaysia. daweizhang.edu@gmail.com.
Large language models (LLMs) can perform criterion-based grading effectively. Prompt engineering with specific criteria enhances LLM grading accuracy, showing domain knowledge is key for educational feedback.
Area of Science:
- Artificial Intelligence
- Educational Technology
- Natural Language Processing
Background:
- Large language models (LLMs) show promise in various applications.
- Automated grading systems are crucial for educational efficiency.
- The efficacy of LLMs in criterion-based grading requires thorough evaluation.
Purpose of the Study:
- To assess the capability of LLMs in criterion-based grading.
- To investigate the effect of detailed prompt engineering on grading performance.
- To understand the role of domain-specific knowledge in LLM grading.
Main Methods:
- Quantitative analysis comparing LLM performance against human benchmarks.
- Evaluation of LLMs using well-established grading criteria.
- Experimentation with prompt engineering techniques to refine LLM instructions.
Main Results:
- Free LLMs demonstrated proficiency in criterion-based grading.
- LLMs exhibited a nuanced understanding of grading criteria.
- Domain-specific understanding proved more critical than model complexity for accurate grading.
Conclusions:
- LLMs are capable of delivering criterion-based educational feedback.
- Prompt engineering significantly impacts LLM grading quality.
- LLMs offer a scalable solution for providing educational feedback.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Related Concept Videos
Sieve Analysis and Grading Curves
Improving Translational Accuracy
Types of Aggregate Grading
Well-graded aggregates include a complete range of necessary size fractions that fit together to create a dense matrix with minimal voids, represented by a smooth, continuous gradation curve. This type of grading ensures good...
Accuracy, limits, and approximation
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Reliability and Validity
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...