Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation

TOP 文献データベース Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation

arxiv

AIセキュリティポータルbot

文献データベースの情報は、自動的に収集されています。

Source

https://arxiv.org/abs/1909.01492

PDF

https://arxiv.org/pdf/1909.01492

文献情報

作者: Po-Sen Huang,Robert Stanforth,Johannes Welbl,Chris Dyer,Dani Yogatama,Sven Gowal,Krishnamurthy Dvijotham,Pushmeet Kohli
公開日: 2019-9-4
更新日: 2019-12-21
所属機関: DeepMind
所属の国: United Kingdom
会議名: EMNLP/IJCNLP

AIにより推定されたラベル

敵対的サンプルの脆弱性敵対的サンプル学習の改善

※ こちらのラベルはAIによって自動的に追加されました。そのため、正確でないことがあります。
詳細は文献データベースについてをご覧ください。

Abstract

Neural networks are part of many contemporary NLP systems, yet their empirical successes come at the price of vulnerability to adversarial attacks. Previous work has used adversarial training and data augmentation to partially mitigate such brittleness, but these are unlikely to find worst-case adversaries due to the complexity of the search space arising from discrete text perturbations. In this work, we approach the problem from the opposite direction: to formally verify a system's robustness against a predefined class of adversarial attacks. We study text classification under synonym replacements or character flip perturbations. We propose modeling these input perturbations as a simplex and then using Interval Bound Propagation -- a formal model verification method. We modify the conventional log-likelihood training objective to train models that can be efficiently verified, which would otherwise come with exponential search complexity. The resulting models show only little difference in terms of nominal accuracy, but have much improved verified accuracy under perturbations and come with an efficiently computable formal guarantee on worst case adversaries.

外部データセット

Stanford Sentiment Treebank (SST)

AG News corpus