Approximate Data Deletion from Machine Learning Models

TOP Literature Database Approximate Data Deletion from Machine Learning Models

arxiv

AI Security Portal bot

Information in the literature database is collected automatically.

Source

https://arxiv.org/abs/2002.10077

PDF

https://arxiv.org/pdf/2002.10077

Paper Information

Author: Zachary Izzo,Mary Anne Smart,Kamalika Chaudhuri,James Zou
Published: 2-24-2020
Updated: 2-24-2021
Affiliation: Dept. of Mathematics
Country: United States of America
Conference: International Conference on Artificial Intelligence and Statistics (AISTATS)

Labels Estimated by AI

Machine learning Model Evaluation Robustness Evaluation

These labels were automatically added by AI and may be inaccurate.
For details, see About Literature Database.

Abstract

Deleting data from a trained machine learning (ML) model is a critical task in many applications. For example, we may want to remove the influence of training points that might be out of date or outliers. Regulations such as EU's General Data Protection Regulation also stipulate that individuals can request to have their data deleted. The naive approach to data deletion is to retrain the ML model on the remaining data, but this is too time consuming. In this work, we propose a new approximate deletion method for linear and logistic models whose computational cost is linear in the the feature dimension $d$ and independent of the number of training data $n$. This is a significant gain over all existing methods, which all have superlinear time dependence on the dimension. We also develop a new feature-injection test to evaluate the thoroughness of data deletion from ML models.

External Datasets

Yelp dataset

synthetic datasets