Related Experiment Video
Updated: Jun 16, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A refined reweighing technique for nondiscriminatory classification.
Yuefeng Liang1, Cho-Jui Hsieh2, Thomas C M Lee1
1Department of Statistics, University of California at Davis, CA, United States of America.
This study introduces a refined reweighing technique for machine learning to reduce discrimination. By considering sensitive and insensitive attributes, it improves fairness with minimal accuracy loss and enhanced scalability.
Area of Science:
- Machine Learning
- Artificial Intelligence
- Data Science
Background:
- Machine learning systems can perpetuate and exacerbate socioeconomic disparities.
- Existing discrimination-aware classification methods often focus solely on sensitive attributes.
Purpose of the Study:
- To propose a novel data pre-processing technique for reducing discrimination in machine learning.
- To refine instance weighting by incorporating both sensitive and insensitive attributes.
Main Methods:
- Developed a data pre-processing technique assigning weights to training instances.
- Formulated weight assignment as a linear programming problem.
- Incorporated both sensitive and insensitive attributes for weight refinement.
Main Results:
- Achieved significant discrimination reduction with a minimal impact on classification accuracy.
- Demonstrated superior scalability compared to existing pre-processing methods.
- Enabled explicit monitoring of the trade-off between fairness and accuracy.
Conclusions:
- The refined reweighing method effectively reduces discrimination without altering input data or labels.
- This approach offers a scalable and transparent solution for enhancing fairness in machine learning models.
- The method provides users with control over the fairness-accuracy balance.
More Related Videos
Related Concept Videos
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Classifying Matter by Composition
According to its composition, the matter can be classified into two broad categories — pure substances and mixtures.
A pure substance is a form of matter that has a constant composition throughout with uniform properties. For example, any sample of sucrose has the same composition and same physical properties, such as melting point, color, and sweetness, regardless of the source from which it is isolated.
A mixture is composed of two or...
Precipitation Gravimetry
In determining nickel by gravimetric analysis, a precipitant of ethanolic dimethylglyoxime is added to a hot nickel salt solution. This is quickly followed by the dropwise addition of dilute ammonia solution until precipitation occurs. A...
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Classification of Titrimetric Analysis Based on Reaction Types
Titrations between an acid and a base lead to neutralization reactions that form...

