Related Experiment Video
Updated: Jul 11, 2025

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
241
BioBERTurk: Exploring Turkish Biomedical Language Model Development Strategies in Low-Resource Setting
Hazal Türkmen1, Oğuz Dikenelli1, Cenk Eraslan2
1Department of Computer Engineering, Ege University, 35100 İzmir, Turkey.
Journal of Healthcare Informatics Research
|November 6, 2023
Summary
This study introduces BioBERTurk, pretrained language models for Turkish biomedical Natural Language Processing (NLP). The best model, further pretrained on a biomedical corpus, significantly outperformed others in classifying radiology reports.
Area of Science:
- Biomedical Natural Language Processing (NLP)
- Low-resource language NLP
Background:
- Pretrained language models excel in English biomedical NLP.
- Limited research exists for low-resource languages, necessitating effective models for these settings.
Purpose of the Study:
- Introduce the BioBERTurk family of four pretrained models for Turkish biomedicine.
- Evaluate model performance on classifying Turkish head CT radiology reports.
Main Methods:
- Developed four BioBERTurk models, including variations pretrained on biomedical corpora and radiology theses.
- Compared BioBERTurk models against Turkish BERT (BERTurk), multilingual BERT (mBERT), and LSTM+attention baselines.
- Evaluated models on classifying 'impressions' and 'findings' sections of radiology reports.
Main Results:
- The model pretrained on a general biomedical corpus significantly outperformed BERTurk, mBERT, and baselines.
- A model continually pretrained using only radiology theses showed slight improvement on the 'impressions' dataset.
- Combining radiology and biomedicine corpora with BERTurk's corpus and pretraining from scratch resulted in the poorest performance.
Conclusions:
- Continued pretraining of BERTurk with a biomedical corpus is effective for Turkish biomedical NLP.
- Task-specific data, like radiology theses, can enhance model performance for specific sub-domains.
- Model architecture and corpus composition significantly impact performance in low-resource biomedical NLP.

