Related Experiment Video
Updated: Oct 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Natural Language Processing and Machine Learning Methods to Characterize Unstructured Patient-Reported Outcomes:
Zhaohua Lu1, Jin-Ah Sim2,3, Jade X Wang1
1Department of Biostatistics, St. Jude Children's Research Hospital, Memphis, TN, United States.
Natural language processing (NLP) and machine learning (ML) algorithms, particularly BERT, accurately assess patient-reported outcomes (PROs) in pediatric cancer survivors. This approach offers a valid alternative to traditional PRO surveys for evaluating symptoms like pain and fatigue.
Area of Science:
- Computational linguistics
- Medical informatics
- Oncology
Background:
- Patient-reported outcomes (PROs) are crucial for assessing survivorship in pediatric cancer patients.
- Traditional PRO surveys can be supplemented by analyzing unstructured data from clinical interviews.
Purpose of the Study:
- To validate natural language processing (NLP) and machine learning (ML) algorithms for identifying pain interference and fatigue attributes in child and adolescent cancer survivors.
- To compare the performance of NLP/ML algorithms against expert-adjudicated PRO data.
Main Methods:
- A cross-sectional study involving 391 meaning units for pain interference and 423 for fatigue from interviews with pediatric cancer survivors (aged 8-17) and caregivers.
- Semantic labeling of meaning units by content experts into physical, cognitive, and social attributes.
- Validation of two NLP/ML methods: bidirectional encoder representations from transformers (BERT) and Word2vec with support vector machine or extreme gradient boosting.
Main Results:
- BERT demonstrated superior accuracy in identifying cognitive and social attributes of pain interference and fatigue compared to Word2vec-based methods.
- BERT achieved high accuracy scores (e.g., 0.931 for pain interference cognitive attributes) and superior areas under the receiver operating characteristic and precision-recall curves.
Conclusions:
- The BERT model shows high validity and accuracy for assessing patient-reported outcomes in pediatric cancer survivors.
- NLP/ML methods, especially BERT, provide a viable and effective alternative to standard PRO surveys for capturing symptom experiences in clinical settings.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Nursing Evaluation
Formulating and Validating Nursing Diagnosis II
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...
Analysis of Population Pharmacokinetic Data