Related Experiment Video
Updated: Jun 5, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations
Lucy Lu Wang1,2, Yulia Otmakhova3, Jay DeYoung4
1University of Washington.
Evaluating multi-document summarization (MDS) quality for biomedical literature reviews is challenging. Current automated metrics often fail to align with human assessments, necessitating new evaluation methods.
Area of Science:
- Natural Language Processing
- Biomedical Informatics
- Information Retrieval
Background:
- Evaluating multi-document summarization (MDS) quality, particularly for biomedical literature reviews, is complex due to the need to synthesize conflicting evidence.
- Existing automated metrics like ROUGE may not accurately reflect summary quality and can be exploited by models through unintended shortcuts.
- There is a lack of resources to assess the validity of proposed automated evaluation metrics for MDS.
Purpose of the Study:
- To introduce a novel dataset of human-assessed summary quality facets and pairwise preferences for literature review MDS.
- To facilitate the development and benchmarking of improved automated evaluation methods for biomedical literature review summarization.
- To analyze the correlation between existing automated metrics, proposed metrics, and human judgments of summary quality.
Main Methods:
- Compiled a diverse dataset of generated summaries from the Multi-document Summarization for Literature Review (MSLR) shared task.
- Collected human assessments of summary quality, including specific facets and pairwise preferences.
- Analyzed correlations between automated metrics (standard and novel) and human-assessed quality aspects.
Main Results:
- Automated metrics frequently fail to capture summary quality aspects as perceived by human annotators.
- In many instances, automated metrics produce system rankings that are inversely correlated with human judgments.
- Proposed novel metrics and analyzed their correlation with human assessments.
Conclusions:
- Current automated evaluation metrics are insufficient for assessing the quality of biomedical literature review summaries.
- The developed dataset provides a crucial resource for advancing research in MDS evaluation.
- Further development of automated metrics that align with human quality perception is essential for reliable MDS in the biomedical domain.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Improving Translational Accuracy
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...