Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks | AIセキュリティポータル

EN

JA

EN

TOP 文献データベース Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks

arxiv

Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks

AIセキュリティポータルbot

文献データベースの情報は、自動的に収集されています。

Source

https://arxiv.org/abs/2408.11587

PDF

https://arxiv.org/pdf/2408.11587

文献情報

作者: Ziqiang Li;Yueqi Zeng;Pengfei Xia;Lei Liu;Zhangjie Fu;Bin Li
公開日: 2024-8-21
所属機関: Nanjing University of Information Science and Technology
所属の国: China
会議名

AIにより推定されたラベル

バックドア攻撃ポイズニング

※ こちらのラベルはAIによって自動的に追加されました。そのため、正確でないことがあります。
詳細は文献データベースについてをご覧ください。

Abstract

With the burgeoning advancements in the field of natural language processing (NLP), the demand for training data has increased significantly. To save costs, it has become common for users and businesses to outsource the labor-intensive task of data collection to third-party entities. Unfortunately, recent research has unveiled the inherent risk associated with this practice, particularly in exposing NLP systems to potential backdoor attacks. Specifically, these attacks enable malicious control over the behavior of a trained model by poisoning a small portion of the training data. Unlike backdoor attacks in computer vision, textual backdoor attacks impose stringent requirements for attack stealthiness. However, existing attack methods meet significant trade-off between effectiveness and stealthiness, largely due to the high information entropy inherent in textual data. In this paper, we introduce the Efficient and Stealthy Textual backdoor attack method, EST-Bad, leveraging Large Language Models (LLMs). Our EST-Bad encompasses three core strategies: optimizing the inherent flaw of models as the trigger, stealthily injecting triggers with LLMs, and meticulously selecting the most impactful samples for backdoor injection. Through the integration of these techniques, EST-Bad demonstrates an efficient achievement of competitive attack performance while maintaining superior stealthiness compared to prior methods across various text classifier datasets.

外部データセット

SST-2

AG News

HSOL

参考文献

Science China Technological Sciences

Pre-trained models for natural language processing: A survey

Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, Xuanjing Huang

Published: 2020

OpenAI Technical Report

Language models are few-shot learners

T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, D. Amodei

Published: 2020

IEEE Transactions on Emerging Topics in Computational Intelligence

A new perspective on stabilizing gans training: Direct adversarial training

Z. Li, P. Xia, R. Tao, H. Niu, B. Li

Published: 2022

ACM Computing Surveys

A systematic survey of regularization and normalization in gans

Z. Li, M. Usman, R. Tao, P. Xia, C. Wang, H. Chen, B. Li

Published: 2023

Proceedings of the IEEE

A comprehensive survey on transfer learning

Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, Qing He

Published: 2021

Proceedings of NAACL-HLT

Bert: Pre-training of deep bidirectional transformers for language understanding

Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova

Published: 2019

Chinese Computational Linguistics: 18th China National Conference

How to fine-tune bert for text classification?

C. Sun, X. Qiu, Y. Xu, X. Huang

Published: 2019

Universal adversarial triggers for attacking and analyzing nlp

E. Wallace, S. Feng, N. Kandpal, M. Gardner, S. Singh

Published: 2019

Computing Research Repository (CoRR)

Weight Poisoning Attacks on Pre-trained Models

Keita Kurita, Paul Michel, Graham Neubig

Published: 2020.4.15

Recently, NLP has seen a surge in the usage of large pre-trained models. Users download weights of models pre-trained on large datasets, then fine-tune the weights on a task of their choice. This raises the question of whether downloading untrusted pre-trained weights can pose a security threat. In this paper, we show that it is possible to construct ``weight poisoning'' attacks where pre-trained weights are injected with vulnerabilities that expose ``backdoors'' after fine-tuning, enabling the attacker to manipulate the model prediction simply by injecting an arbitrary keyword. We show that by applying a regularization method, which we call RIPPLe, and an initialization procedure, which we call Embedding Surgery, such attacks are possible even with limited knowledge of the dataset and fine-tuning procedure. Our experiments on sentiment classification, toxicity detection, and spam detection show that this attack is widely applicable and poses a serious threat. Finally, we outline practical defenses against such attacks. Code to reproduce our experiments is available at https://github.com/neulab/RIPPLe.

敵対的学習バックドア攻撃ポイズニング

Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing

Mind the style of text! adversarial and backdoor attacks based on text style transfer

Fanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li, Zhiyuan Liu, Maosong Sun

Published: 2021