Related Experiment Videos
MeSH term explosion and author rank improve expert recommendations
Danielle H Lee1, Titus Schleyer
1School of Information Sciences;
AMIA ... Annual Symposium Proceedings. AMIA Symposium
|February 25, 2011
Summary
Finding research collaborators is easier with a new algorithm. This study shows that using Medical Subject Headings (MeSH) terms and author rank improves expert recommendations for scientists.
Area of Science:
- Biomedical research
- Information science
- Computer science
Background:
- Information overload hinders scientific productivity and collaboration.
- Existing expertise location methods are often unsuitable for biomedical research.
- Identifying suitable research collaborators is a significant challenge for scientists.
Purpose of the Study:
- To develop and evaluate a vector space model-based algorithm for calculating researcher similarity.
- To assess the effectiveness of different input features for expert recommendations in biomedical research.
Main Methods:
- Developed a vector space model algorithm using MeSH terms, exploded MeSH terms, and author rank.
- Evaluated the algorithm on a dataset of 17,525 authors and 22,542 publications.
- Compared four variations of the algorithm based on input data.
Main Results:
- The algorithm using exploded MeSH terms and author rank demonstrated the highest accuracy in predicting coauthors.
- MeSH terms combined with author rank also showed high accuracy.
- On average, the algorithms correctly predicted 2.5 of the top 5/10 coauthors.
Conclusions:
- Combining MeSH terms with author rank significantly enhances the accuracy of researcher similarity calculations.
- The developed algorithm offers a promising approach for expert recommendations in biomedical research.
- Metadata such as author rank can improve the effectiveness of MeSH term-based matching for collaboration discovery.
Related Concept Videos
Ranks
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
Outliers and Influential Points
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the vertical...
Weighted Mean
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Geometric Mean
The mean is a measure of the central tendency of a data set. In some data sets, the data is inherently multiplicative, and the arithmetic mean is not useful. For example, the human population multiplies with time, and so does the credit amount of financial investment, as the interest compounds over successive time intervals.
In cases of multiplicative data, the geometric mean is used for statistical analysis. First, the product of all the elements is taken. Then, if there are n elements in the...
In cases of multiplicative data, the geometric mean is used for statistical analysis. First, the product of all the elements is taken. Then, if there are n elements in the...
Predicting Molecular Geometry
VSEPR Theory for Determination of Electron Pair Geometries
Maxwell-Boltzmann Distribution: Problem Solving
Individual molecules in a gas move in random directions, but a gas containing numerous molecules has a predictable distribution of molecular speeds, which is known as the Maxwell-Boltzmann distribution, f(v).
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by