Related Experiment Video
Updated: Aug 10, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Equal accuracy for Andrew and Abubakar-detecting and mitigating bias in name-ethnicity classification algorithms
Lena Hafner1, Theodor Peter Peifer2, Franziska Sofia Hafner3
1Department of Politics and International Studies, University of Cambridge, Cambridge, UK.
Abstract:
Uncovering the world's ethnic inequalities is hampered by a lack of ethnicity-annotated datasets. Name-ethnicity classifiers (NECs) can help, as they are able to infer people's ethnicities from their names. However, since the latest generation of NECs rely on machine learning and artificial intelligence (AI), they may suffer from the same racist and sexist biases found in many AIs. Therefore, this paper offers an algorithmic fairness audit of three NECs. It finds that the UK-Census-trained EthnicityEstimator displays large accuracy biases with regards to ethnicity, but relatively less among gender and age groups. In contrast, the Twitter-trained NamePrism and the Wikipedia-trained Ethnicolr are more balanced among ethnicity, but less among gender and age. We relate these biases to global power structures manifested in naming conventions and NECs' input distribution of names. To improve on the uncovered biases, we program a novel NEC, N2E, using fairness-aware AI techniques. We make N2E freely available at www.name-to-ethnicity.com.
Supplementary Information:
The online version contains supplementary material available at 10.1007/s00146-022-01619-4.
Related Concept Videos
Stereotypes, Prejudice, and Discrimination
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Confirmation Biases
Surveys
Ethnic Identity within a Larger Culture
Bias in Epidemiological Studies

