Related Experiment Video
Updated: Jul 11, 2026

Recording Mouse Ultrasonic Vocalizations to Evaluate Social Communication
Published on: June 5, 2016
Beyond Tokens: Fair Evaluation of French Large Language Models for Clinical Named Entity Recognition
Jamil Zaghir1,2, Mina Bjelogrlic1,2, Jean-Philippe Goldman1,2
1Division of Medical Information Sciences, Geneva University Hospitals, Geneva, Switzerland.
Abstract:
Named Entity Recognition (NER) models based on Transformers have gained prominence for their impressive performance in various languages and domains. This work delves into the often-overlooked aspect of entity-level metrics and exposes significant discrepancies between token and entity-level evaluations. The study utilizes a corpus of synthetic French oncological reports annotated with entities representing oncological morphologies. Four different French BERT-based models are fine-tuned for token classification, and their performance is rigorously assessed at both token and entity-level. In addition to fine-tuning, we evaluate ChatGPT's ability to perform NER through prompt engineering techniques. The findings reveal a notable disparity in model effectiveness when transitioning from token to entity-level metrics, highlighting the importance of comprehensive evaluation methodologies in NER tasks. Furthermore, in comparison to BERT, ChatGPT remains limited when it comes to detecting advanced entities in French.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Components of Language
Language and Cognition

