Me-LLaMA: Foundation Large Language Models for Medical Applications
Qianqian Xie1, Qingyu Chen1, Aokun Chen2
1Section of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT, USA.
Research Square
|June 3, 2024
Summary
Me-LLaMA, a new family of medical large language models (LLMs), demonstrates superior performance on medical tasks compared to existing open-source models. These models address limitations in clinical settings by utilizing specialized medical data and mitigating catastrophic forgetting.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Natural Language Processing
Background:
- Large language models (LLMs) show promise in medicine but lack specialized medical training.
- Existing open-source medical LLMs have limitations in clinical application.
- The challenge lies in adapting general LLMs for nuanced medical data.
Purpose of the Study:
- Introduce Me-LLaMA, a novel family of medical large language models (LLMs).
- Develop foundation and chat-enhanced models using LLaMA2 architecture and extensive medical datasets.
- Evaluate Me-LLaMA's performance against existing models and benchmarks.
Main Methods:
- Utilized continual pre-training on a 129B token medical dataset.
- Employed instruction tuning with a 214k sample dataset.
- Developed and used the Medical Instruction Benchmark Evaluation (MIBE) across six medical tasks and 12 datasets.
Main Results:
- Me-LLaMA models outperformed existing open-source medical LLMs in zero-shot, few-shot, and supervised learning.
- Task-specific instruction-tuned Me-LLaMA surpassed ChatGPT on 7/8 and GPT-4 on 5/8 MIBE datasets.
- Me-LLaMA demonstrated superior mitigation of the catastrophic forgetting problem.
Conclusions:
- Me-LLaMA represents a significant advancement in open-source medical LLMs, integrating biomedical and clinical data.
- The models offer enhanced performance on both general and medical tasks, making them suitable for medical AI applications.
- Me-LLaMA provides a robust foundation for future medical AI research and deployment.
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Language Development
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
Introduction to Language of Pathophysiology l
Pathophysiology investigates how biological mechanisms—typically starting at the cellular level—disrupt normal bodily functions. It bridges anatomy and physiology to explain the progression of disease. With this foundation, it is important to understand the following key terms used to describe disease processes: Diagnosis:The process of identifying a disease using clinical evaluation, including signs (objective evidence like rashes), symptoms (subjective experiences like pain), laboratory test...
Introduction to Language of Pathophysiology ll
This lesson explores key terms that describe how diseases progress, their outcomes, and their distribution in populations.Diagnostic tests identify diseases and monitor treatment. These include blood and urine tests, biopsies, imaging (X-ray, MRI), and detection of infectious agents.Remission is a reduction or disappearance of symptoms.Exacerbation refers to the worsening of symptoms, such as increased wheezing during an asthma attack.A precipitating factor triggers an acute episode, while a...


