BiasBuster:一种神经方法,使用偏差位置数据准确估计人口统计数据
Sepanta Zeighami1, Cyrus Shahabi2
1UC Berkeley.
移动位置数据有偏见,导致人口统计数据不准确. 一个神经网络BiasBuster纠正了这种偏见,提高了所有人群的准确性,特别是代表性不足的群体.
科学领域:
- 数据科学数据科学数据科学
- 计算社会科学 计算社会科学
- 机器学习 机器学习
背景情况:
- 移动设备位置数据被广泛用于城市流动性,商业洞察力和公共卫生政策.
- 这些数据集经常受到人口偏见的影响,某些社区的代表过多或不足.
- 从有偏见的数据中获得的综合统计数据导致了不准确的人口表示,并不成比例地影响了边缘化群体.
研究的目的:
- 为了应对从偏见的移动位置数据中生成准确的人口统计数据的挑战.
- 为了评估传统的统计失调方法的有效性.
- 引入和验证一种用于偏差校正的新型神经网络方法.
主要方法:
- 提出了BiasBuster,一种神经网络模型,利用位置特征和人口统计数据之间的相关性.
- 使用真实世界的位置数据进行了广泛的实验.
- 比较BiasBuster的性能与传统的统计失误分析技术.
主要成果:
- 统计失调方法往往无法显著提高准确性.
- 在估计人口统计数据方面,BiasBuster表现出了显著的改进.
- 准确度提高到一般的两倍,而代表性不足的人口则增加了三倍.
结论:
- 偏差位置数据对政策制定和研究构成重大风险.
- BiasBuster提供了一个强大的解决方案,可以从移动位置数据中获得更准确的人口统计数据.
- 这些发现突显了机器学习在减轻数据偏差和获得公平见解方面的潜力.
更多相关视频
08:45Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
Published on: October 24, 2012
08:19Stereological Estimation of Dopaminergic Neuron Number in the Mouse Substantia Nigra Using the Optical Fractionator and Standard Microscopy Equipment
Published on: September 1, 2017
相关概念视频
Distributions to Estimate Population Parameter
Estimating Population Standard Deviation
Estimating Population Mean with Unknown Standard Deviation
William S. Gosset (1876–1937) of the...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used;...
