Related Experiment Video
Updated: Jun 9, 2026

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
Published on: August 29, 2025
A survey on LLM-as-a-judge
Jiawei Gu1,2, Xuhui Jiang1,3, Zhichao Shi1,3,4
1IDEA Research, Shenzhen, China.
Abstract:
Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains challenging due to inherent subjectivity, variability, and scale. Large language models (LLMs) have achieved remarkable success, leading to "LLM-as-a-judge," where LLMs serve as evaluators for complex tasks. With their ability to process diverse data types and provide scalable assessments, LLMs present a compelling alternative to traditional expert-driven evaluations. However, ensuring the reliability of LLM-as-a-judge systems remains a significant challenge requiring careful design and standardization. This paper provides a comprehensive survey of LLM-as-a-judge, offering a formal definition and detailed classification while addressing the core question of how to build reliable LLM-as-a-judge systems. We explore strategies to enhance reliability, including improving consistency, mitigating biases, and adapting to diverse scenarios. We propose methodologies for evaluating reliability, supported by a novel benchmark. To advance development and deployment, we discuss practical applications, challenges, and future directions. Our contributions span multiple levels: we establish conceptual boundaries, reorganize fragmented literature into a unified framework, and propose a reliability-oriented benchmark. We articulate a forward-looking research agenda, offering theoretical foundations and practical guidance for constructing reliable and trustworthy LLM-as-a-judge systems.
Related Concept Videos
Sources of Law
Constitutional law is foundational, deriving from federal and state constitutions, and...
Types of Surveys
Surveys
Torts III
Quasi-intentional torts in healthcare involve acts where intent is not directed to harm an individual but results in harm due to careless or reckless speech.
Torts II
Survey Safety
