Related Experiment Video
Updated: Sep 2, 2025

Sample Preparation to Bioinformatics Analysis of DNA Methylation: Association Strategy for Obesity and Related Trait Studies
Published on: May 6, 2022
A dataset of mentorship in bioscience with semantic and demographic estimations
Qing Ke1, Lizhen Liang2, Ying Ding3
1School of Data Science, City University of Hong Kong, Kowloon, Hong Kong. q.ke@cityu.edu.hk.
Abstract:
Mentorship in science is crucial for topic choice, career decisions, and the success of mentees and mentors. Typically, researchers who study mentorship use article co-authorship and doctoral dissertation datasets. However, available datasets of this type focus on narrow selections of fields and miss out on early career and non-publication-related interactions. Here, we describe Mentorship, a crowdsourced dataset of 743176 mentorship relationships among 738989 scientists primarily in biosciences that avoids these shortcomings. Our dataset enriches the Academic Family Tree project by adding publication data from the Microsoft Academic Graph and "semantic" representations of research using deep learning content analysis. Because gender and race have become critical dimensions when analyzing mentorship and disparities in science, we also provide estimations of these factors. We perform extensive validations of the profile-publication matching, semantic content, and demographic inferences, which mostly cover neuroscience and biomedical sciences. We anticipate this dataset will spur the study of mentorship in science and deepen our understanding of its role in scientists' career outcomes.
More Related Videos
Related Concept Videos
Biostatistics: Overview
Discrete variables are...
Overview of Biostatistics in Health Sciences
Statistical Methods for Analyzing Epidemiological Data
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Longitudinal Studies

