Related Experiment Video
Updated: Jan 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Performance of LLMs in ICF Classification: Insights from Medical and General Models
Abstract:
In the medical field, text data comprising personal anecdotes and detailed patient insights are often underutilized due to their unstructured nature and variability among clinicians. However, recent advances in Large Language Models (LLMs) present an opportunity to harness this data effectively. This paper explores the use of the International Classification of Functioning, Disability, and Health (ICF) framework recommended by the World Health Organization (WHO), which offers a holistic approach considering personal and environmental factors along with impairments, to structure textual descriptions systematically. The study investigates the application of medically fine-tuned LLMs, such as MedAlpaca and Meditron, for automated ICF creation, comparing their efficiency in processing real medical cases from two distinct contexts: rehabilitation and intensive care units. Additionally, we benchmark medical LLMs against general-purpose LLMs, including ChatGPT and Claude, to assess whether specialized models truly offer an advantage in medical classification tasks. Preliminary findings indicate that while medical LLMs show potential for ICF classification tasks, they may not necessarily outperform general-purpose models, as the complexity of ICF requires a deeper level of contextual understanding.International classification of functioning, disability and health (ICF), Intensive care, large language models, AI, medical, LLM, Rehabilitation.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Improving Translational Accuracy
Improving Translational Accuracy

