Related Experiment Video
Updated: Sep 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Infusing clinical knowledge into language models by subword optimisation and embedding initialisation
Abul Hasan1, Jinge Wu1, Quang Ngoc Nguyen1
1University College London, Institute of Health Informatics, 222 Euston Rd., London, NW1 2DA, UK.
Objective:
This study introduces a novel tokenisation methodology, K-Tokeniser, to infuse clinical knowledge into language models for clinical text processing.
Methods:
Technically, at initialisation stage, K-Tokeniser populates global representations of tokens based on semantic types of domain concepts (such as drugs or diseases) from either a domain ontology like Unified Medical Language System or the training data of the task related corpus. At training or inference stage, sentence level localised context will be utilised for choosing the optimal global token representation to realise the semantic-based tokenisation. To avoid pretraining using the new tokeniser, an embedding initialisation approach is proposed to generate representations for new tokens.
Results:
Using three transformer-based language models, a comprehensive set of experiments are conducted on four real-world datasets for evaluating K-Tokeniser in a wide range of clinical text analytics tasks including clinical concept and relation extraction, automated clinical coding, clinical phenotype identification, and clinical research article classification. Overall, our models demonstrate consistent improvements over their counterparts in all tasks. In particular, substantial improvements are observed in the automated clinical coding task with 13% increase on Micro F1 score. Furthermore, K-Tokeniser also shows significant capacities in facilitating quicker convergence of language models.
Conclusion:
Models built using K-Tokeniser have shown faster convergence. Specifically,the language models would only require 50% of the training data to achieve the best performance of the baseline tokeniser using all training data in the concept extraction task and less than 20% of the data for the automated coding task. It is worth mentioning that all these improvements require no pre-training process, making the approach generalisable. Code availability: Our full implementation is openly available at https://github.com/abulhasanbbk/K-Tokenizer.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
12:49Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
Related Concept Videos
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Components of Language
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Language and Cognition
Clinical Trials
There are four phases in a clinical trial. A phase one...