Related Experiment Video
Updated: Feb 4, 2026

A Semantic Priming Event-related Potential ERP Task to Study Lexico-semantic and Visuo-semantic Processing in Autism Spectrum Disorder
Published on: April 12, 2018
Identifying High-Priority Proteins Across the Human Diseasome Using Semantic Similarity
Edward Lau1, Vidya Venkatraman2, Cody T Thomas3
1Stanford Cardiovascular Institute , Stanford University , Stanford , California 94305 , United States.
This study introduces a new computational method to identify and rank important proteins for biological topics and diseases. The weighted copublication distance (WCD) metric prioritizes proteins using literature analysis, aiding research and hypothesis generation.
Area of Science:
- Biomedical Research
- Computational Biology
- Bioinformatics
Background:
- Identifying key genes and proteins for biological processes and diseases is crucial but often relies on subjective expert evaluation and manual literature review.
- Systematic, objective methods for evaluating protein-disease associations are limited, hindering efficient research.
- Automating the prioritization of proteins for specific research topics is needed.
Purpose of the Study:
- To develop and validate a computational method for objectively identifying and ranking proteins associated with specific research topics and diseases.
- To introduce the weighted copublication distance (WCD) metric for quantifying protein-topic relationships.
- To create a comprehensive human "diseasome" of protein-disease associations.
Main Methods:
- Developed a method to calculate the semantic similarity between proteins and query terms in biomedical literature.
- Incorporated adjustments for the impact and immediacy of associated research articles.
- Calculated the weighted copublication distance (WCD) metric to rank protein-topic relationships.
Main Results:
- The WCD metric demonstrated strong performance in predicting benchmark protein lists across multiple biological processes.
- Prioritized "popular proteins" were extracted for various cell types, anatomical regions, and over 20,000 human disease terms.
- A human "diseasome" of protein-disease associations was generated, enabling reverse queries and functional annotation.
Conclusions:
- The WCD metric provides a robust and objective approach for prioritizing proteins relevant to research topics and diseases.
- The generated "diseasome" facilitates data analysis, hypothesis generation, and the annotation of experimental protein lists.
- This bibliometric approach enhances the understanding of protein functions and their roles in human diseases.
More Related Videos
Related Concept Videos
Causes of Similarity-Dissimilarity Effect
Factors Influencing Attraction III: Similarity
Role of Proteins in the Human Body
Protein Families
Identifying Statistically Significant Differences: The F-Test
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...

