Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Survival Tree01:19

Survival Tree

132
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
132
Regression Toward the Mean01:52

Regression Toward the Mean

6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Calibration Curves: Linear Least Squares01:20

Calibration Curves: Linear Least Squares

1.6K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
1.6K
Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
2.6K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

1.8K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.8K
Kaplan-Meier Approach01:24

Kaplan-Meier Approach

229
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
229

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Distinct molecular subgroups in pediatric and young-onset meningiomas require age-adapted risk stratification.

Nature communications·2026
Same author

AI-based selection of tumor regions for genomic profiling in neuropathology.

Neuro-oncology advances·2026
Same author

Cystobactamid off-target profiling reveals favorable safety, superoxide reduction, and SCARB1 inhibition in eukaryotes.

npj drug discovery·2026
Same author

Centrifugal microfluidic automation of the protein aggregation capture workflow for robust mass spectrometry-based proteomics.

Lab on a chip·2026
Same author

Hetairos is a histology-based artificial intelligence model for predicting central nervous system tumor methylation subtypes.

Nature cancer·2026
Same author

Natural Product-Derived Ianthelliformisamines Inhibit Protein Translation and Block Bacterial Flagellum Assembly.

ACS chemical biology·2026

Related Experiment Video

Updated: Aug 22, 2025

Stepwise Dosing Protocol for Increased Throughput in Label-Free Impedance-Based GPCR Assays
06:13

Stepwise Dosing Protocol for Increased Throughput in Label-Free Impedance-Based GPCR Assays

Published on: February 21, 2020

6.6K

Tuning gradient boosting for imbalanced bioassay modelling with custom loss functions.

Davide Boldini1, Lukas Friedrich2, Daniel Kuhn2

  • 1Center for Functional Protein Assemblies, Technical University of Munich (TUM), Ernst-Otto-Fischer-Straße 8, 85784, Garching, Germany.

Journal of Cheminformatics
|November 11, 2022
PubMed
Summary

Novel loss functions significantly improve Gradient Boosting for imbalanced bioassay datasets in drug discovery, achieving state-of-the-art results faster than traditional methods.

Keywords:
Gradient boostingImbalanced classificationVirtual screening

More Related Videos

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.9K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K

Related Experiment Videos

Last Updated: Aug 22, 2025

Stepwise Dosing Protocol for Increased Throughput in Label-Free Impedance-Based GPCR Assays
06:13

Stepwise Dosing Protocol for Increased Throughput in Label-Free Impedance-Based GPCR Assays

Published on: February 21, 2020

6.6K
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.9K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K

Area of Science:

  • Computational chemistry
  • Machine learning in drug discovery
  • Bioinformatics

Background:

  • Bioassay datasets are crucial for drug discovery but often exhibit severe class imbalance.
  • This imbalance hinders the performance of machine learning models in identifying active compounds.

Purpose of the Study:

  • To investigate alternative loss functions for imbalanced classification within Gradient Boosting models.
  • To benchmark these novel approaches against conventional methods on diverse bioassay data.

Main Methods:

  • Exploration of computer vision-inspired loss functions for imbalanced classification.
  • Application and benchmarking of these loss functions within a Gradient Boosting framework.
  • Evaluation across six public and proprietary bioassay datasets, totaling 42 tasks and 2 million compounds.

Main Results:

  • Statistically significant performance improvements were observed on five out of six datasets compared to standard cross-entropy loss.
  • Gradient Boosting with optimized loss functions matched or surpassed existing classifiers and neural networks.
  • Training convergence speed increased up to 8 times faster.

Conclusions:

  • Tuning loss functions is an effective and efficient strategy for Gradient Boosting on imbalanced bioassay data.
  • This approach achieves state-of-the-art performance without sacrificing interpretability or scalability.
  • Offers a practical solution for improving drug discovery pipelines.