Related Experiment Video
Updated: Jul 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
German Medical NER with BERT and LLMs: The Impact of Training Data Size
Suteera Seeha1, Sihan Wu1, Justin Hofenbitzer1
1Institute of AI and Informatics in Medicine (AIIM), TUM University Hospital, Technical University of Munich, Munich, Germany.
Abstract:
Named Entity Recognition (NER) in the medical domain often presents significant challenges due to the complexity and specificity of medical terminology, especially in lower-resource settings where annotated data is scarce. This study explores the performance of an exemplary large language model and a BERT-based model in the context of NER for German medical texts. We focus on the impact of different data sizes for training and their performance to simulate lower-resource conditions. Both models are evaluated on two German annotated corpora. Our results reveal that both models perform rather similar on both datasets, with LLaMA3.1 performing slightly better with less training material.