Related Experiment Video
Updated: Jan 22, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
VetTag: improving automated veterinary diagnosis coding via large-scale language modeling
Yuhui Zhang1, Allen Nie2, Ashley Zehnder2
11Department of Computer Science and Technology, Tsinghua University, Beijing, China.
This study introduces a novel algorithm for automatically predicting all 4577 veterinary diagnosis codes from free-text clinical notes. This advancement overcomes a major barrier in veterinary medicine, enabling better public health and translational research.
Area of Science:
- Veterinary Medicine
- Natural Language Processing
- Machine Learning
Background:
- Veterinary records are predominantly unstructured free text, lacking standardized diagnosis coding.
- This data format hinders the use of veterinary records for public health and translational research.
- Previous machine learning efforts were limited to predicting only 42 broad diagnosis categories.
Purpose of the Study:
- To develop a large-scale algorithm for automatically predicting all 4577 standard veterinary diagnosis codes from free-text clinical notes.
- To improve the utility of veterinary records for public health and translational research.
- To explore the application of advanced machine learning techniques in veterinary clinical data.
Main Methods:
- Developed a novel algorithm based on an adapted Transformer architecture.
- Trained the algorithm on a large dataset including over 100,000 expert-labeled and over one million unlabeled veterinary notes.
- Utilized large-scale language modeling through pretraining and an auxiliary objective during supervised learning.
- Employed hierarchical training to address data imbalances for fine-grained diagnoses.
Main Results:
- Successfully developed an algorithm to predict all 4577 standard veterinary diagnosis codes from free text.
- Demonstrated significant performance improvements through large-scale language modeling on unlabeled data.
- Evaluated model performance in challenging cross-hospital settings with substantial domain shift.
- Showcased the effectiveness of hierarchical training for rare or fine-grained diagnoses.
Conclusions:
- The developed algorithm effectively addresses the challenge of systematic coding in veterinary medicine.
- The study highlights the power of unsupervised learning and advanced NLP techniques for clinical data.
- This work facilitates leveraging veterinary records for broader public health and research applications.
More Related Videos
10:15Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia
Published on: July 2, 2013
06:16Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
Published on: June 6, 2020
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
lncRNA - Long Non-coding RNAs
lncRNA - Long Non-coding RNAs
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition