Related Experiment Video
Updated: Jan 16, 2026

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
Published on: August 29, 2025
Evaluating AI-Powered Q&A Systems: A Simple Approach to Determining the Need for Expert Ratings
Dorian Zwanzig1, Luca Kreibich1, Uta Binder1
1HTW Berlin, Faculty 4 (Computing, Communication and Business).
None:
This paper introduces a simple approach for assessing whether laypeople or AI-based automations can adequately substitute for expert ratings in the evaluation of AI-powered Q&A systems It employs weighted Cohen's Kappa to assess inter-rater reliability, establishing an expert agreement benchmark and comparing this to individual alternative rater-expert agreements. By visualizing these results in an inter-rater reliability matrix, it is a transparent and structured way to determine the adequacy of non-expert raters. Our findings, based on a real project, suggest that laypeople or AI, in some cases, can match or exceed expert agreement, particularly when risk aversion is a factor. The approach can be adapted to different contexts and rating attributes.
Related Concept Videos
Reason and Intuition
Self-Evaluation: Self-Enhancement and Self-Verification
Self-Evaluation Maintenance Model
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Cochran's Q Test
Friedman Two-way Analysis of Variance by Ranks