Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Systematic Sampling Method01:17

Systematic Sampling Method

11.5K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
Systematic sampling is one of the simplest methods...
11.5K
Classification of Systems-II01:31

Classification of Systems-II

264
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
264
Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

940
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
940
Biostatistics: Overview01:20

Biostatistics: Overview

419
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
419
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

2.8K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.8K
Classification of Systems-I01:26

Classification of Systems-I

370
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
370

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

High-Frequency Physiological Measures Predict Post-Admission Surgical Intervention After Severe Traumatic Brain Injury.

Journal of neurotrauma·2025
Same author

Association of past-year mental and physical health conditions with intentional or unintentional drug overdoses.

Journal of substance use and addiction treatment·2025
Same author

Increasing COVID-19 Testing and Vaccination Uptake in the Take Care Texas Community-Based Randomized Trial: Adaptive Geospatial Analysis.

JMIR formative research·2025
Same author

Correction: Social connectedness as a determinant of mental health: A scoping review.

PloS one·2024
Same author

Factors associated with elevated SARS-CoV-2 immune response in children and adolescents.

Frontiers in pediatrics·2024
Same author

Baseline characteristics of SARS-CoV-2 vaccine non-responders in a large population-based sample.

PloS one·2024

Related Experiment Video

Updated: Oct 22, 2025

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases
05:02

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases

Published on: October 24, 2019

32.5K

Recommender system of scholarly papers using public datasets.

Jie Zhu1, Braja G Patra1, Ashraf Yaseen1

  • 1University of Texas Health Science Center at Houston Houston, TX, USA.

AMIA Joint Summits on Translational Science Proceedings. AMIA Joint Summits on Translational Science
|August 30, 2021
PubMed
Summary

Scholarly recommender systems improve dataset findability. Term-frequency methods like BM25 and TF-IDF are most effective for recommending PubMed papers relevant to Gene Expression Omnibus datasets.

More Related Videos

Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study
07:50

Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study

Published on: April 18, 2025

507
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
07:41

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases

Published on: May 17, 2019

9.2K

Related Experiment Videos

Last Updated: Oct 22, 2025

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases
05:02

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases

Published on: October 24, 2019

32.5K
Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study
07:50

Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study

Published on: April 18, 2025

507
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
07:41

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases

Published on: May 17, 2019

9.2K

Area of Science:

  • Bioinformatics
  • Information Science
  • Computational Biology

Background:

  • The increasing volume of public biological datasets necessitates improved methods for discovery and reuse.
  • Scholarly recommender systems are crucial for navigating and utilizing large-scale data resources.

Purpose of the Study:

  • To develop and evaluate a scholarly recommendation system connecting research papers with public datasets.
  • To enhance the findability and reusability of public datasets, specifically those in the Gene Expression Omnibus (GEO).

Main Methods:

  • Implementation of a scholarly recommendation system linking PubMed articles to GEO datasets.
  • Comparison of various text representation techniques, including term-frequency (TF-IDF, BM25) and Natural Language Processing (NLP) embedding models (doc2vec, ELMo, BERT).

Main Results:

  • Term-frequency based methods (BM25 and TF-IDF) demonstrated superior performance in relevance prediction.
  • Traditional methods outperformed advanced NLP embedding models for this specific recommendation task.

Conclusions:

  • Term-frequency based approaches offer a highly effective and efficient solution for scholarly dataset recommendations.
  • The developed system aids researchers in identifying relevant literature, thereby increasing dataset utility and accelerating scientific discovery.