Related Experiment Video
Updated: Feb 7, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Evaluating semantic relations in neural word embeddings with biomedical and general domain knowledge bases
Zhiwei Chen1, Zhe He2, Xiuwen Liu1
1Department of Computer Science, Florida State University, Tallahassee, FL, USA.
Background:
In the past few years, neural word embeddings have been widely used in text mining. However, the vector representations of word embeddings mostly act as a black box in downstream applications using them, thereby limiting their interpretability. Even though word embeddings are able to capture semantic regularities in free text documents, it is not clear how different kinds of semantic relations are represented by word embeddings and how semantically-related terms can be retrieved from word embeddings.
Methods:
To improve the transparency of word embeddings and the interpretability of the applications using them, in this study, we propose a novel approach for evaluating the semantic relations in word embeddings using external knowledge bases: Wikipedia, WordNet and Unified Medical Language System (UMLS). We trained multiple word embeddings using health-related articles in Wikipedia and then evaluated their performance in the analogy and semantic relation term retrieval tasks. We also assessed if the evaluation results depend on the domain of the textual corpora by comparing the embeddings of health-related Wikipedia articles with those of general Wikipedia articles.
Results:
Regarding the retrieval of semantic relations, we were able to retrieve diverse semantic relations in the nearest neighbors of a given word. Meanwhile, the two popular word embedding approaches, Word2vec and GloVe, obtained comparable results on both the analogy retrieval task and the semantic relation retrieval task, while dependency-based word embeddings had much worse performance in both tasks. We also found that the word embeddings trained with health-related Wikipedia articles obtained better performance in the health-related relation retrieval tasks than those trained with general Wikipedia articles.
Conclusion:
It is evident from this study that word embeddings can group terms with diverse semantic relations together. The domain of the training corpus does have impact on the semantic relations represented by word embeddings. We thus recommend using domain-specific corpus to train word embeddings for domain-specific text mining tasks.
More Related Videos
Related Concept Videos
Acid and Bases: Ka, pKa, and Relative Strengths
Relative Strengths of Conjugate Acid-Base Pairs
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Membrane Domains
Protein Domains
The membrane comprises a group of distinct proteins responsible for carrying out a cell's specific function. For example, the plasma membrane of the human sperm, or a single germ cell, contains a unique set of proteins in the...
Three Developmental Domains
Physical Development
Physical processes, also known as maturation, encompass the biological changes that occur across an individual's life. These changes begin with genetic inheritance and continue through various stages, including growth in height and weight,...
Three-Domain System of Life

