Related Experiment Video
Updated: May 9, 2025

06:48
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.0K
Assessing large language model performance related to aging in genetic conditions
Amna A Othman1, Kendall A Flaharty2, Suzanna E Ledgister Hanchard2
1Medical Genomics Unit, National Human Genome Research Institute, National Institutes of Health, Bethesda, MD, USA. amna.othman@nih.gov.
Npj Aging
|May 3, 2025
Summary
Large language models (LLMs) like GPT-3.5 and Llama-2-70b-chat can generate age-appropriate medical information for genetic conditions. However, their clinical management plans require further improvement for accuracy and completeness.
Area of Science:
- Medical Genetics
- Artificial Intelligence in Medicine
- Clinical Informatics
Background:
- Genetic conditions are often described in pediatric populations, creating a knowledge gap for adult patient care.
- Large language models (LLMs) show promise in various applications, prompting investigation into their medical utility.
Purpose of the Study:
- To evaluate the ability of Llama-2-70b-chat (70b) and GPT-3.5 (GPT) to generate age-appropriate medical vignettes, dialogues, and management plans for genetic conditions in both child and adult hypothetical patients.
- To assess the correctness and completeness of LLM-generated content for 282 genetic conditions.
Main Methods:
- LLMs (70b and GPT) were prompted to create medical vignettes, patient-geneticist dialogues, and management plans for hypothetical pediatric and adult patients across 282 genetic conditions.
- Clinicians graded the generated content for correctness and completeness, focusing on age-appropriateness and clinical plausibility.
- A sub-analysis was performed on metabolic conditions, including those presenting neonatally.
Main Results:
- LLMs demonstrated age-appropriate responses in generating medical vignettes and dialogues for both child and adult patients, as indicated by clinician scoring.
- Sub-analysis confirmed age-appropriate LLM responses for metabolic conditions, even those with neonatal crisis presentations.
- Both 70b and GPT models achieved lower correctness and completeness scores for generating plausible clinical management plans, ranging from 50-90% depending on the model and specific metrics.
Conclusions:
- LLMs show potential in generating age-specific medical information for genetic conditions, aiding in bridging pediatric-adult knowledge gaps.
- Current LLMs have limitations in producing clinically accurate and complete management plans, highlighting areas for future development and validation in healthcare applications.
More Related Videos
Related Concept Videos
Language and Cognition
287
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
287
Human Genetics
470
Human genetics provides a profound framework for understanding the interplay between genetic predispositions and human psychology. At the heart of this discipline lies the study of how genes influence physical traits, behaviors, and susceptibility to diseases. Each person carries a unique genetic code that subtly or significantly shapes their psychological and behavioral landscape.
The complex relationship between genetics and psychology is observable through common biological components such...
The complex relationship between genetics and psychology is observable through common biological components such...
470

