Related Experiment Video
Updated: Jul 2, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
555
Taiyi: a bilingual fine-tuned large language model for diverse biomedical tasks.
Ling Luo1, Jinzhong Ning1, Yingwen Zhao1
1School of Computer Science and Technology, Dalian University of Technology, Dalian 116024, China.
Journal of the American Medical Informatics Association : JAMIA
|February 29, 2024
Summary
Taiyi, a bilingual large language model (LLM), demonstrates superior performance on diverse biomedical natural language processing (NLP) tasks. This fine-tuned LLM shows potential for bilingual biomedical multitasking, outperforming general LLMs.
Area of Science:
- Biomedical Natural Language Processing (NLP)
- Artificial Intelligence in Healthcare
- Computational Linguistics
Background:
- Existing biomedical large language models (LLMs) primarily focus on monolingual question answering and conversation.
- There is a need to evaluate LLM performance across diverse biomedical NLP tasks and languages.
Purpose of the Study:
- To develop and evaluate Taiyi, a bilingual fine-tuned LLM for a wide range of biomedical NLP tasks.
- To assess the effectiveness of supervised fine-tuning strategies on diverse biomedical datasets.
Main Methods:
- Curated 140 biomedical text mining datasets (102 English, 38 Chinese) across 10+ task types.
- Converted corpora into instruction data for supervised fine-tuning of a general LLM.
- Employed a 2-stage fine-tuning strategy to optimize performance.
Main Results:
- Taiyi achieved superior performance on 13 test sets, including named entity recognition, relation extraction, text classification, and question answering.
- Demonstrated considerable potential for bilingual biomedical multitasking in a case study.
- Outperformed general LLMs on various biomedical NLP benchmarks.
Conclusions:
- High-quality biomedical corpora and effective fine-tuning strategies significantly enhance LLM performance in the biomedical domain.
- Taiyi exhibits bilingual multitasking capabilities via supervised fine-tuning.
- Generative LLMs still face challenges in non-generation tasks like information extraction, where discriminative models remain superior.
Related Concept Videos
Improving Translational Accuracy
10.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.4K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K

