Related Experiment Video
Updated: Apr 5, 2026

Author Spotlight: Exploring Sex-Specific Glial Signatures and Therapeutic Leads for Alzheimer's Disease
Published on: May 20, 2024
Performance of name-to-gender inference: comparison between Gender API, NamSor, and genderize.io in a multicultural
Paul Sebo1, Amrollah Shamsi2, Ting Wang3
1University Institute for Primary Care (IuMFE), University of Geneva, 1211, Geneva, Switzerland. paul.seboe@unige.ch.
None:
Inferring gender from first names is widely used in medical research when gender data are unavailable, but tool performance may vary across cultural and linguistic contexts. We compared the real-world performance of three name-to-gender inference tools (Gender API, NamSor, and genderize.io) in a large multicultural dataset with self-reported gender labels. We assembled publicly available results from seven major 2025 marathons: New York, Berlin, Paris, Shanghai, Tokyo, Dubai, and Abu Dhabi (n = 11,999, women = 6,000, men = 5,999). First names, with and without country of nationality, were submitted to each tool. We computed confusion matrices and performance metrics: overall error (errorCoded: misclassifications + nonclassifications), misclassifications among classified names (errorCodedWithoutNA), and nonclassifications (naCoded). A priori, a tool was defined as accurate when errorCoded was < 10%. Pairwise tool comparisons used McNemar's test. We also assessed confidence-threshold subsetting (≥ 60%- ≥ 90%) and region-specific performance. All three tools met the prespecified accuracy criterion overall. NamSor outperformed both comparators (p < 0.001). Without country information, errorCoded was 0.0481 (NamSor), 0.0867 (Gender API), and 0.0798 (genderize.io); NamSor produced no unclassified names. Performance was substantially lower for Asian nationalities and improved after excluding Asia/China. Increasing confidence thresholds reduced misclassifications but increased nonclassifications, raising overall error. Adding country improved performance for Gender API (p < 0.001) and genderize.io (p = 0.03), but not for NamSor (p = 0.41). In Conclusion, although all tools were accurate overall, performance depended strongly on geographic composition. Researchers should report expected misclassification rates based on external validation studies, justify confidence thresholds, and rely on tools validated in culturally comparable populations.
Related Concept Videos
Socioemotional Experience and Gender Development
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Test for Homogeneity
Improving Translational Accuracy
Improving Translational Accuracy
Causes of Similarity-Dissimilarity Effect
