Reproducibility of social science research using aggregate statistics with noise infused for differential privacy.
Privacy-preserving analytics such as differential privacy are designed to allow the analysis of sensitive datasets while protecting individuals' privacy. Their deployment, however, has been controversial. Critics maintain that statistical noise injected to preserve privacy can degrade the quality and feasibility of social science research. We select a benchmark of empirical findings from 93 published social science studies involving regression analyses over aggregate statistics. We evaluate whet
Privacy-preserving analytics such as differential privacy are designed to allow the analysis of sensitive datasets while protecting individuals' privacy. Their deployment, however, has been controversial. Critics maintain that statistical noise injected to preserve privacy can degrade the quality and feasibility of social science research. We select a benchmark of empirical findings from 93 published social science studies involving regression analyses over aggregate statistics. We evaluate whether their findings replicate on privacy noise-infused data. Under privacy budgets typical in industry, around 91% of simulated findings still support the original claims at significance level [Formula: see text]. Claims based on weaker original effect sizes are more likely to be nullified or sometimes reversed. We compare distortions caused by privacy noise to those due to measurement errors and other kinds of nonsampling errors common in social statistics, and we find that the marginal impacts of privacy protection are smaller. Moreover, discrepancies due to privacy noise are often much smaller than discrepancies observed in traditional replication and robustness studies.
