Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Language Development01:22

Language Development

444
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
444
Higher Mental Functions of the Brain: Language01:10

Higher Mental Functions of the Brain: Language

1.0K
Language is a system of communication that allows the expression of thoughts, ideas, and feelings. The brain processes language in both hemispheres.
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
1.0K
Components of Language01:24

Components of Language

388
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
388
Language and Cognition01:27

Language and Cognition

436
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
436
Genetic Lingo01:11

Genetic Lingo

104.5K
Overview
104.5K
Language01:16

Language

420
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
420

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A dynamic risk prediction framework for Alzheimer's disease and related dementias with interpretability.

NPJ digital medicine·2026
Same author

Clinical document metadata extraction: A scoping review.

Journal of biomedical informatics·2026
Same author

Reproducibility and Robustness of Large Language Models for Mobility Functional Status Extraction.

medRxiv : the preprint server for health sciences·2026
Same author

Understanding Primary and Secondary Concerns from Patient Portal Messages through Clinical Data Annotation, Analysis, and Modeling.

AMIA ... Annual Symposium proceedings. AMIA Symposium·2026
Same author

Mobility functional status ascertainment in electronic health records using large language models.

Scientific reports·2026
Same author

Context matching is not reasoning when performing generalized clinical evaluation of generative language models.

NPJ digital medicine·2025

Related Experiment Video

Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

680

Context Matching is not Reasoning: Assessing Generalized Evaluation of Generative Language Models in Clinical

Andrew Wen1, Qiuhao Lu1, Yu-Neng Chuang2

  • 1The University of Texas Health Science Center at Houston.

Research Square
|September 5, 2025
PubMed
Summary

Generative language models (GLMs) struggle with clinical benchmark validity. Current assessments fail to ensure knowledge application, consistent answers, or recognition of null-answer scenarios, necessitating improved benchmark designs.

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

565
Portable Intermodal Preferential Looking IPL: Investigating Language Comprehension in Typically Developing Toddlers and Young Children with Autism
10:11

Portable Intermodal Preferential Looking IPL: Investigating Language Comprehension in Typically Developing Toddlers and Young Children with Autism

Published on: December 14, 2012

18.5K

Related Experiment Videos

Last Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

680
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

565
Portable Intermodal Preferential Looking IPL: Investigating Language Comprehension in Typically Developing Toddlers and Young Children with Autism
10:11

Portable Intermodal Preferential Looking IPL: Investigating Language Comprehension in Typically Developing Toddlers and Young Children with Autism

Published on: December 14, 2012

18.5K

Area of Science:

  • Artificial Intelligence in Medicine
  • Natural Language Processing
  • Clinical Assessment Validation

Background:

  • Generative Language Models (GLMs) are increasingly evaluated using clinical Multiple-Choice Question-Answer (MCQA) benchmarks.
  • Existing MCQA benchmarks, designed for human examinees, may not accurately reflect GLM capabilities due to unique model characteristics.

Purpose of the Study:

  • To validate the generalizability of MCQA-based assessments for GLMs.
  • To investigate the assumptions of knowledge application, semantic consistency, and null-answer recognition in GLMs.
  • To identify limitations of current benchmarks and propose more robust assessment designs.

Main Methods:

  • Four clinical MCQA benchmarks were validated using eight GLMs.
  • Model performance was analyzed by ablating parameter size and reasoning capabilities.
  • Prompt permutation was employed to test three core assumptions of MCQA generalizability.

Main Results:

  • All tested GLMs demonstrated significant failures in recognizing null-answer scenarios.
  • Large GLMs showed greater resilience to prompt perturbations than smaller models.
  • Despite retaining knowledge, smaller GLMs exhibited a propensity for answer memorization rather than application.

Conclusions:

  • Current MCQA benchmarks have questionable validity for assessing GLMs in clinical settings.
  • Fundamental assumptions regarding knowledge application, semantic consistency, and null-answer recognition are globally invalidated for GLMs.
  • Adaptations to benchmark design are crucial for more accurate and reliable evaluation of GLM clinical capabilities.