Related Experiment Video
Updated: Aug 22, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
QAScore-An Unsupervised Unreferenced Metric for the Question Generation Evaluation
Tianbo Ji1, Chenyang Lyu2, Gareth Jones1
1ADAPT Centre, School of Computing, Dublin City University, 9 Dublin, Ireland.
A new evaluation metric, QAScore, is proposed for Question Generation (QG) systems. This reference-free metric better aligns with human judgment than existing methods, improving evaluation accuracy for automated question generation.
Area of Science:
- Natural Language Processing
- Artificial Intelligence
Background:
- Automated Question Generation (QG) has advanced with neural models, but evaluation remains a challenge.
- Current metrics like BLEU and BERTScore rely on references and show low agreement with human judgment.
- Existing metrics for QG systems do not consider the passage or answer context.
Purpose of the Study:
- To introduce QAScore, a novel reference-free evaluation metric for Question Generation.
- To provide a more accurate and human-aligned method for assessing QG system performance.
- To address the limitations of current QG evaluation metrics.
Main Methods:
- QAScore evaluates generated questions by measuring a language model's ability to predict masked answer words.
- The metric computes cross-entropy based on the probability of correctly generating masked words within the answer.
- A human evaluation experiment was conducted to compare QAScore with existing metrics.
Main Results:
- QAScore demonstrates a stronger correlation with human judgment compared to BLEU and BERTScore.
- The proposed metric offers improved evaluation accuracy for Question Generation systems.
- Human evaluation confirmed the superiority of QAScore in assessing question quality.
Conclusions:
- QAScore provides a more reliable and accurate evaluation for Question Generation systems.
- Reference-free evaluation using QAScore enhances the assessment of automated question quality.
- This metric facilitates better development and benchmarking of QG technologies.
Related Concept Videos
Detection of Gross Error: The Q Test
Cochran's Q Test
Quantifying and Rejecting Outliers: The Grubbs Test
Random Sampling Method
Confidence Coefficient
Randomized Experiments
Simple randomization
Simple...

