Related Experiment Video
Updated: Jul 9, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Quantifying domain-specific relevance of computational biology Wikipedia articles using TF-IDF and cosine similarity
Arya Pradeep Menon1, Trenton Davis2, Megha Hegde1
1School of Computing and Mathematics, Faculty of Engineering, Computing, and Environment, Kingston University, London, KT1 2EE, United Kingdom.
Motivation:
Wikipedia is one of the world's most visited websites and serves as the principal open educational resource for computational biology. However, identifying which articles are most relevant to distinct sub-disciplines of computational biology remains largely subjective.
Results:
This study collected short descriptions for 22 Communities of Special Interest (COSI) groups maintained by the International Society for Computational Biology and downloaded 1536 computational biology articles from English Wikipedia. Following standard text preprocessing, COSI descriptions and Wikipedia articles were embedded in a common TF-IDF vector space. Semantic relatedness was quantified using cosine similarity, yielding a real-valued relevance matrix that maps each COSI to the most pertinent computational biology articles. The resulting scores, typically low in absolute value, captured nuanced differences: general-interest pages such as 'Computational biology' and 'Bioinformatics' ranked highest, whereas niche pages showed high relevance only for specific COSIs. Unsupervised analysis using principal component analysis, k-nearest neighbours, and Leiden community detection revealed clusters of articles corresponding to the particular COSIs and highlighted inter-COSI relationships. This automated pipeline reduces bias compared with manual tagging and enables more precise curation of domain-specific educational resources.
Availability And Implementation:
The relevance matrix developed in this study is available in the Zenodo repository (doi: 10.5281/zenodo.18311878).
Related Concept Videos
Protein-protein Interfaces
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
The Significance of Membrane Transport
Transporters facilitate either an active or passive movement of solutes. They can allow a single-molecule transport down its...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Globular and Fibrous Proteins
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
Cis-regulatory Sequences
