Related Experiment Video
Updated: Sep 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Exploiting survey and online patient experience comments in general practice: validating a fine-tuned language model
Rebecka Maria Norman1, Lilja Charlotte Storset2, Petter Mæhlum2
1Division for Health Services, Norwegian Institute of Public Health, N-0213 Oslo, Norway.
Background:
Automated sentiment analysis can be used to analyse free-text patient experience comments, but in health services research, it requires technical performance and methodological validity. Traditional models lack contextual understanding, whereas masked language models (MLMs) enable nuanced sentiment classification. This study is the first to fine-tune an MLM (NorBERT3large) on human-annotated data and evaluate the validity of four-category sentiment classifications across survey and online feedback in general practice.
Objective(S):
To evaluate a fine-tuned MLM against human-annotated sentiment labels, compare its performance with traditional methods and a large language model, and assess convergent and known-groups validity using survey and online comments.
Methods:
NorBERT3large was fine-tuned on human-annotated free-text survey comments. Performance was evaluated using F1-scores. The model classified unseen online data. Construct validity was assessed per COSMIN guidelines through correlations with patient experience scores and subgroup analyses.
Results:
The fine-tuned NorBERT3large achieved high F1-scores on held-out survey test data and outperformed traditional models used in prior patient experience research, as well as a zero-shot large language model reference. On unseen online reviews, performance was high in the random sample and strongest for mixed and positive sentiment in the balanced evaluation set, while the rare neutral comments were difficult to classify reliably. Construct validity was supported by expected correlations with patient experience scores and known subgroup differences.
Conclusion:
Beyond good classification performance, the findings support automated sentiment analysis as a valid patient experience measure in general practice. The sentiment classifications may complement traditional patient experience measures and help identify areas for quality improvement.