Related Experiment Video
Updated: Jan 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Ensemble learning for improved sentiment analysis in doctor-patient communication
Yufan Ge1, Lingling Dai1, Bingding Huang2
1International College, Anhui Medical University, Hefei, China.
An ensemble model achieved the highest accuracy in classifying clinician-patient sentiment, outperforming deep learning and transformer models. This approach enhances understanding of patient-doctor interactions for improved healthcare.
Area of Science:
- Natural Language Processing (NLP)
- Machine Learning in Healthcare
- Computational Linguistics
Background:
- Accurate sentiment analysis of clinician-patient interactions is crucial for evaluating healthcare quality and patient experience.
- Existing research lacks comprehensive benchmarking of advanced machine learning models for this specific task.
- Sentiment classification in clinical dialogue presents unique challenges due to nuanced language and context.
Purpose of the Study:
- To benchmark deep learning, transformer, and ensemble models for three-class sentiment classification (low/medium/high) in doctor-patient consultations.
- To address the gap in standardized evaluation of sentiment analysis models within the clinical domain.
- To identify the most effective model architecture for analyzing sentiment in anonymized doctor-patient dialogues.
Main Methods:
- Utilized a publicly available dataset of 3325 anonymized doctor-patient consultations.
- Evaluated Long Short-Term Memory (LSTM), Bidirectional LSTM (BiLSTM), Convolutional Neural Networks (CNN), CNN-LSTM, and Bidirectional Encoder Representations from Transformers (BERT).
- An ensemble model (hard voting over Logistic Regression, Random Forest, and Support Vector Classifier) was also tested using stratified five-fold cross-validation.
Main Results:
- The ensemble model achieved the highest accuracy (75.5% ± 0.5%), outperforming individual models including BERT (66.98% ± 0.6%).
- The ensemble demonstrated robust detection of high-severity interactions (F1 score: 90.8% ± 1.3%), though low-severity classification remained challenging.
- BERT offered the highest precision for low-severity interactions (65.5% ± 1.0%), while the ensemble improved recall (58.7% ± 1.0%).
Conclusions:
- Ensemble learning provides the strongest and most balanced performance for three-class sentiment classification in clinician-patient dialogue.
- Transformer models like BERT offer valuable precision for challenging low-severity cases, complementing ensemble approaches.
- Interpretability analyses enhance transparency, supporting clinical application; future work should explore multimodal and privacy-preserving models.
Related Concept Videos
Patient-centered Care
Techniques of Therapeutic Communication II: Focusing, Paraphrasing, and Summarizing
This therapeutic technique can also be used when a patient brings up pertinent information during a health-related conversation. The...
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Improving Translational Accuracy
Improving Translational Accuracy
Current Trends in Nursing II
