Related Experiment Video
Updated: Jan 9, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
994
Clinical Large Language Model Evaluation by Expert Review (CLEVER): Framework Development and Validation
Veysel Kocaman1, Mustafa Aytuğ Kaya2, Andrei Marian Feier1
1John Snow Labs Inc, 16192 Coastal Highway, Lewes, DE, 19958, United States, +1 (302) 786-5227.
JMIR AI
|December 4, 2025
Summary
A new evaluation method, CLEVER, shows that a specialized small LLM outperforms GPT-4o in clinical tasks. This highlights the potential of healthcare-specific large language models (LLMs) for medical applications.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing
- Clinical Informatics
Background:
- Evaluating large language models (LLMs) is challenging due to data contamination and the gap between benchmark tasks and clinical practice.
- Existing LLM evaluation methods like public benchmarks and LLM-as-a-judge are limited by data issues and self-preference bias.
- There is a critical need for robust evaluation frameworks that reflect real-world clinical utility.
Purpose of the Study:
- To introduce CLEVER (Clinical Large Language Model Evaluation-Expert Review), a novel methodology for evaluating LLMs in healthcare.
- To conduct a blind, randomized, preference-based evaluation of LLMs using practicing medical doctors.
- To compare the performance of a general-purpose LLM against healthcare-specific LLMs on clinical tasks.
Main Methods:
- The CLEVER methodology was employed to compare GPT-4o with two healthcare-specific LLMs (8B and 70B parameters).
- Evaluations were performed on three distinct clinical tasks: text summarization, information extraction, and question answering.
- Practicing medical doctors provided preference-based evaluations focusing on factuality, clinical relevance, and conciseness.
Main Results:
- Medical doctors preferred the smaller, healthcare-specific LLM over GPT-4o in 45% to 92% of cases across key clinical dimensions.
- The healthcare-specific LLM demonstrated superior performance in factuality, clinical relevance, and conciseness compared to GPT-4o.
- Performance was comparable in open-ended medical question answering, indicating specialized LLMs excel in context-dependent tasks.
Conclusions:
- Healthcare-specific LLMs can outperform larger, general-purpose LLMs in tasks requiring clinical context understanding.
- The CLEVER methodology provides a valid and reliable approach for evaluating clinical LLMs, confirmed by interannotator agreement and correlation analyses.
- This study underscores the importance of specialized LLMs and expert review for advancing AI in medicine.
Related Concept Videos
Data Validation
6.3K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
6.3K
Language Development
806
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
806
Language and Cognition
693
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
693
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Components of Language
723
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
723

