Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

2.3K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.3K
Goodness-of-Fit Test01:16

Goodness-of-Fit Test

4.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
4.3K
Multiple Regression01:25

Multiple Regression

3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
High-Resolution Mass Spectrometry (HRMS)01:15

High-Resolution Mass Spectrometry (HRMS)

1.6K
The resolution of a mass spectrometer depends on the efficiency of separating ions with different ion masses. The mass of an atom is approximated to the sum of the masses of protons and neutrons inside, considering the masses of protons and neutrons as equal. However, the masses of the proton (1.6726 × 10−24 g) and neutron (1.6749 × 10−24 g) are not truly equal. There is a minor error in the expression of atomic masses relative to the simplest atom of hydrogen. For...
1.6K
Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

14.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
14.5K
Decision Making: P-value Method01:09

Decision Making: P-value Method

5.8K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A Comprehensive HLA-DR4 MHC Class II Tetramer Platform for the Detection and Functional Validation of Post-Translational Modification Neoantigens.

International journal of molecular sciences·2026
Same author

Development and validation of a predictive model for chronic pain after thoracoscopic pulmonary resection.

Frontiers in public health·2026
Same author

Comparison of Liposomal Bupivacaine for Single-Injection Versus Dual-Injection Thoracic Paravertebral Block on Analgesic Effect After Dual-Microport Thoracoscopic Lobectomy: A Prospective Randomized Controlled Trial.

Pain research & management·2026
Same author

Ecotype-Specific Drilosphere Microbiome Reprogramming Influencing Microplastic Impacts on Soil Carbon-Nitrogen Characteristics and Earthworm Health.

Environmental science & technology·2026
Same author

Association of body mass index and visceral fat with heart rate variability in university students: a BMI-stratified analysis.

Frontiers in physiology·2026
Same author

In-situ ceramic nanoparticle assembly within wood microstructure for strong, tough, and resilient ceramic wood.

Nature communications·2026

Related Experiment Video

Updated: Sep 28, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

703

Multi-attribute scientific documents retrieval and ranking model based on GBDT and LR.

Xuedong Tian1,2,3, Jiameng Wang1,2,3, Yu Wen1

  • 1School of Cyber Security and Computer, Hebei University, Baoding 071002, China.

Mathematical Biosciences and Engineering : MBE
|March 28, 2022
PubMed
Summary

This study introduces a novel model for scientific document retrieval, integrating mathematical expressions and text. The proposed method enhances search accuracy and ranking using gradient boosting decision tree (GBDT) and logistic regression (LR) models.

Keywords:
GBDTLRhesitation fuzzy setsmathematical expressionmulti-attribute featurescientific document retrieval and ranking

More Related Videos

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.7K
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

633

Related Experiment Videos

Last Updated: Sep 28, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

703
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.7K
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

633

Area of Science:

  • Information Science
  • Computer Science
  • Computational Linguistics

Background:

  • Scientific documents heavily rely on mathematical expressions and semantic text, posing retrieval challenges.
  • Existing retrieval methods struggle to effectively integrate both mathematical and textual features.
  • Accurate retrieval and ranking of scientific documents are crucial for research and knowledge discovery.

Purpose of the Study:

  • To propose a multi-attribute model for scientific document retrieval and ranking.
  • To effectively integrate mathematical expressions and textual features for improved search.
  • To enhance the performance of scientific document retrieval systems.

Main Methods:

  • A multi-attribute model was developed using gradient boosting decision tree (GBDT) and logistic regression (LR).
  • Five key attributes were calculated: mathematical expression symbols, sub-forms, context, document keywords, and expression frequency.
  • GBDT was employed for feature discretization and reorganization, followed by input into the LR model for final ranking.

Main Results:

  • The model achieved an average MAP@20 of 81.92% for scientific document recall.
  • The average nDCG@20 for scientific document ranking reached 86.05%.
  • Experimental results on the NTCIR dataset demonstrate the model's effectiveness.

Conclusions:

  • The proposed GBDT and LR-based model significantly improves scientific document retrieval and ranking.
  • Integrating mathematical and textual features is key to overcoming current retrieval limitations.
  • The model shows strong performance, offering a valuable advancement in scientific information access.