Related Experiment Video
Updated: Mar 27, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Evaluating Encoder and Decoder Models for Extended Clinical Concept Recognition in Japanese Clinical Texts: A
Yuya Tsukiji1, Satoshi Kataoka1, Masafumi Itokazu1
1Center for Disease Biology and Integrative Medicine, Graduate School of Medicine, The University of Tokyo, 7-3-1 Hongo, Bunkyo-kuClinical Research Center A646, The University of Tokyo Hospital, Tokyo, JP.
Encoder models excel at extended clinical concept recognition (E-CCR), outperforming decoders for extracting long medical phrases. Domain-specific pretraining showed limited benefits, with encoder token classification achieving top performance.
Area of Science:
- Natural Language Processing
- Medical Informatics
- Computational Linguistics
Background:
- Digitized medical documents offer vast data, but extracting complex clinical concepts remains challenging.
- Conventional named entity recognition (NER) struggles with long phrases vital for applications like diagnostic support.
- Extended Clinical Concept Recognition (E-CCR) is essential for capturing these complex medical expressions.
Purpose of the Study:
- To identify optimal strategies for E-CCR model selection.
- To compare encoder vs. decoder models and general-purpose vs. domain-specific pretraining.
- To analyze model effectiveness based on target length and propose a novel evaluation metric.
Main Methods:
- Evaluated 17 encoder and decoder models on the J-CaseMap database (approx. 20,000 Japanese case reports).
- Utilized a novel "weighted soft matching score" to penalize fragmentation and weight by target length.
- Assessed performance variations concerning target length and pretraining strategies.
Main Results:
- Encoder models outperformed decoder models in the E-CCR task.
- JMedDeBERTa(s), an encoder model, achieved the highest mean performance (F1 = 0.758).
- Domain-specific pretraining showed limited benefits; token classification outperformed instruction tuning for long expressions.
Conclusions:
- Encoder-based token classification is highly effective for E-CCR, offering potential advantages in resource-constrained settings.
- Model performance was robust against fragmentation, indicating reliable extraction of long phrases.
- Findings suggest generalizability to Japanese medical text information extraction, warranting further cross-lingual and cross-document type investigation.
More Related Videos
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
Natural and Artificial Concepts
Nursing Clinical Information System
A Nursing Clinical Information System (NCIS) is a specialized type of healthcare information system tailored to meet the unique needs of nursing practice. It incorporates the principles of nursing informatics to streamline information management and improve the quality of care delivery.
Critical attributes of NCIS include:
Stereotype Content Model

