Preservation of Anomalous Subgroups On Machine Learning Transformed Data

TOP 文献データベース Preservation of Anomalous Subgroups On Machine Learning Transformed Data

arxiv

AIセキュリティポータルbot

文献データベースの情報は、自動的に収集されています。

Source

https://arxiv.org/abs/1911.03674

PDF

https://arxiv.org/pdf/1911.03674

文献情報

作者: Samuel C. Maina,Reginald E. Bryant,William O. Goal,Robert-Florian Samoilescu,Kush R. Varshney,Komminist Weldemariam
公開日: 2019-11-9
所属機関: IBM Research, Nairobi, Kenya
所属の国: Kenya
会議名: Computing Research Repository (CoRR)

AIにより推定されたラベル

プライバシー評価プライバシー保護アルゴリズム機械学習の基礎

※ こちらのラベルはAIによって自動的に追加されました。そのため、正確でないことがあります。
詳細は文献データベースについてをご覧ください。

Abstract

In this paper, we investigate the effect of machine learning based anonymization on anomalous subgroup preservation. In particular, we train a binary classifier to discover the most anomalous subgroup in a dataset by maximizing the bias between the group's predicted odds ratio from the model and observed odds ratio from the data. We then perform anonymization using a variational autoencoder (VAE) to synthesize an entirely new dataset that would ideally be drawn from the distribution of the original data. We repeat the anomalous subgroup discovery task on the new data and compare it to what was identified pre-anonymization. We evaluated our approach using publicly available datasets from the financial industry. Our evaluation confirmed that the approach was able to produce synthetic datasets that preserved a high level of subgroup differentiation as identified initially in the original dataset. Such a distinction was maintained while having distinctly different records between the synthetic and original dataset. Finally, we packed the above end to end process into what we call Utility Guaranteed Deep Privacy (UGDP) system. UGDP can be easily extended to onboard alternative generative approaches such as GANs to synthesize tabular data.