Related Experiment Video
Updated: Feb 20, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.2K
On the Consistency of Automatic Scoring with Large Language Models.
Mingfeng Xue1, Xingyao Xiao2, Yunting Liu3
1University of North Carolina Greensboro, USA.
Educational and Psychological Measurement
|February 19, 2026
Summary
Large language models (LLMs) show high intra-LLM consistency in scoring, but inter-LLM consistency is moderate. A voting strategy combining LLM outputs improves scoring accuracy.
Area of Science:
- Artificial Intelligence
- Educational Measurement
- Natural Language Processing
Background:
- Large language models (LLMs) demonstrate potential for automated scoring tasks.
- Inconsistency in LLM scoring can arise from model variations and training data differences.
- Understanding LLM scoring consistency is crucial for reliable automated assessment.
Purpose of the Study:
- To investigate intra-LLM and inter-LLM scoring consistency across five LLMs.
- To examine the impact of temperature settings on LLM scoring consistency.
- To evaluate the relationship between scoring consistency and accuracy.
- To propose and assess a voting strategy for improving LLM scoring.
Main Methods:
- Evaluated scoring consistency of five LLMs (Claude, DeepSeek, Gemini, GPT, Qwen).
- Assessed consistency under varying temperature settings.
- Utilized constructed-response items from science education and ASAP datasets.
- Implemented a majority voting strategy across LLMs.
Main Results:
- LLMs exhibited near-perfect intra-LLM consistency, irrespective of temperature.
- Inter-LLM consistency was moderate, higher for easier items.
- Intra-LLM consistency surpassed inter-LLM consistency.
- Intra-LLM consistency did not correlate with accuracy; inter-LLM consistency showed a positive correlation.
- Majority voting enhanced scoring accuracy by combining diverse LLM strengths.
Conclusions:
- LLMs offer high internal scoring reliability but varying external agreement.
- Inter-LLM consistency is a better predictor of scoring accuracy than intra-LLM consistency.
- Ensemble methods, like majority voting, can mitigate LLM scoring inconsistencies and improve accuracy in educational assessments.
Related Concept Videos
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
Improving Translational Accuracy
3.7K
3.7K
Automatic Processing and Automatic Social Behavior
276
Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
276
Language and Cognition
836
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
836
Language Development
973
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
973
Accuracy, limits, and approximation
1.3K
Accuracy, limits, and approximations are common in many fields, especially in engineering calculations. These concepts are imperative for ensuring that a given value is as close as possible to its true value.
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
1.3K