Related Experiment Video
Updated: May 12, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
2.7K
Oversampling effect in pretraining for bidirectional encoder representations from transformers (BERT) to localize
Shoya Wada1, Toshihiro Takeda1, Katsuki Okada1
1Department of Medical Informatics, Osaka University Graduate School of Medicine, Japan.
Artificial Intelligence in Medicine
|May 10, 2024
Summary
Oversampling domain-specific corpora and pretraining Bidirectional Encoder Representations from Transformers (BERT) models in a balanced manner significantly improves performance on specialized language tasks. This method enhances information extraction in domains with limited high-quality data.
Area of Science:
- Natural Language Processing
- Machine Learning
- Bioinformatics
Background:
- Large-scale neural language models enhance transfer learning in NLP.
- Transformer-based models like BERT improve information extraction but struggle in data-scarce domains.
- Training domain-specific BERT models is challenging due to limited high-quality, large-scale databases.
Purpose of the Study:
- To address the challenge of training high-performance BERT models in domains with limited data.
- To investigate the effectiveness of oversampling domain-specific corpora for pretraining.
- To develop and evaluate a novel pretraining method for domain-specific BERT models.
Main Methods:
- Simultaneous pretraining of models using oversampled knowledge from distinct domains.
- Development of English biomedical BERT from a small corpus.
- Creation of Japanese medical BERT from a small corpus.
- Enhanced biomedical BERT pretraining using balanced PubMed abstracts.
Main Results:
- English BERT achieved practical performance on the Biomedical Language Understanding Evaluation (BLUE) benchmark.
- Proposed method outperformed conventional methods for biomedical corpora of similar sizes.
- Japanese medical BERT surpassed conventional models on most medical tasks.
- Enhanced biomedical BERT showed superior clinical and biomedical scores on the BLUE benchmark.
Conclusions:
- Balanced pretraining with oversampled, task-appropriate corpora enables high-performance BERT model construction.
- The proposed method offers a viable solution for developing effective domain-specific language models.
- This approach is particularly beneficial for specialized fields with data limitations.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...

