Related Experiment Video
Updated: Feb 24, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Exploring evaluation measures of large language models for family caregiver use: A scoping review
Soojeong Han1, Hannah Cho2, Yong K Choi3
1School of Nursing, Columbia University, New York, NY, USA.
Background:
Large language models have a huge positive impact on various disciplines, including healthcare. As family caregivers are an essential part of the healthcare system, they need support and can benefit from the technology. However, there is no consensus on reliable and valid measures to evaluate large language models.
Objective:
This study aims to review the literature on the evaluation measures of large language models for caregivers.
Methods:
We conducted a scoping review guided by Arksey and O'Malley methodology and the PRISMA-ScR checklist. A literature search on PubMed, EMBASE, CINAHL, and PsycINFO, from 2018 through July 2024, was carried out. An additional rapid review was conducted for the recent literature update from July 2024 through November 2025.
Results:
All 10 final publications that met the inclusion criteria out of 1812 focused on ChatGPT, whereas three of them also addressed other large language models, such as Google Bard and Bing AI. The most commonly assessed core conceptual components of evaluation measures were accuracy, reliability, readability, and comprehensiveness. Overall, the included studies reported that large language models' responses were somewhat accurate and reliable and mixed results in readability and comprehensiveness. The final 14 publications from a rapid review offered additional evidence on ChatGPT-centrism.
Conclusions:
This review provides a comprehensive overview of the measures for evaluating large language models and highlights the need for their improvement using reliable and valid measures. The findings guide the direction of future research and practice to maximize the benefits through continuous quality improvement.
More Related Videos
07:56Assessing the Coherence of Parents' Short Narratives Regarding their Child Using the Five-Minute Speech Sample Procedure
Published on: September 19, 2019
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020