Deep Probabilistic Models to Detect Data Poisoning Attacks

TOP 文献データベース Deep Probabilistic Models to Detect Data Poisoning Attacks

Computing Research Repository (CoRR)

AIセキュリティポータルbot

文献データベースの情報は、自動的に収集されています。

Source

https://arxiv.org/abs/1912.01206

PDF

https://arxiv.org/pdf/1912.01206

文献情報

作者: Mahesh Subedar,Nilesh Ahuja,Ranganath Krishnan,Ibrahima J. Ndiour,Omesh Tickoo
公開日: 2025-3-25
所属機関: Intel Labs
所属の国: United States of America
会議名: Computing Research Repository (CoRR)

AIにより推定されたラベル

バックドア攻撃ポイズニング攻撃性能評価

Abstract

Data poisoning attacks compromise the integrity of machine-learning models by introducing malicious training samples to influence the results during test time. In this work, we investigate backdoor data poisoning attack on deep neural networks (DNNs) by inserting a backdoor pattern in the training images. The resulting attack will misclassify poisoned test samples while maintaining high accuracies for the clean test-set. We present two approaches for detection of such poisoned samples by quantifying the uncertainty estimates associated with the trained models. In the first approach, we model the outputs of the various layers (deep features) with parametric probability distributions learnt from the clean held-out dataset. At inference, the likelihoods of deep features w.r.t these distributions are calculated to derive uncertainty estimates. In the second approach, we use Bayesian deep neural networks trained with mean-field variational inference to estimate model uncertainty associated with the predictions. The uncertainty estimates from these methods are used to discriminate clean from the poisoned samples.