Related Experiment Video
Updated: Jul 25, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
628
Localizing in-domain adaptation of transformer-based biomedical language models
Tommaso Mario Buonocore1, Claudio Crema2, Alberto Redolfi2
1Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Pavia, 27100, Italy.
Journal of Biomedical Informatics
|June 29, 2023
Summary
Creating biomedical language models for under-resourced languages like Italian is challenging. This study found that while data quantity is crucial, combining high-quality data can significantly improve model performance, even with limited resources.
Area of Science:
- Computational linguistics
- Biomedical informatics
- Natural language processing
Background:
- Digital healthcare generates vast textual data, valuable for improving patient care.
- Fine-tuning language models on domain-specific resources enhances performance.
- Resources for adapting models to less-resourced languages like Italian are scarce.
Purpose of the Study:
- To investigate accessible methods for creating Italian biomedical language models.
- To compare the impact of data quantity versus quality in model adaptation.
- To address the gap in in-domain adaptation for non-English medical resources.
Main Methods:
- Two approaches were explored: neural machine translation of English resources (quantity-focused) and using a native Italian corpus (quality-focused).
- Models were fine-tuned using these distinct datasets.
- Performance was evaluated to determine the effectiveness of each approach.
Main Results:
- Data quantity proved to be a more significant constraint than data quality for biomedical adaptation.
- Concatenating high-quality data enhanced model performance, even with smaller corpora.
- The developed models offer potential for Italian medical research and institutions.
Conclusions:
- Accessible methods can be developed for creating biomedical language models in less-resourced languages.
- Balancing data quantity and quality is key for effective domain adaptation.
- The findings provide insights for generalizing biomedical language model creation to other languages and domains.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K
Transformers
1.1K
A device that transforms voltages from one value to another using induction is called a transformer. A transformer consists of two separate coils, or windings, wrapped around the same soft iron core. However, they are electrically insulated from each other.
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
1.1K

