Detecting Language Model Attacks with Perplexity

TOP 文献データベース Detecting Language Model Attacks with Perplexity

arxiv

AIセキュリティポータルbot

文献データベースの情報は、自動的に収集されています。

Source

https://arxiv.org/abs/2308.14132

PDF

https://arxiv.org/pdf/2308.14132

文献情報

作者: Gabriel Alon;Michael Kamfonas
公開日: 2023-8-28
更新日: 2023-11-7
所属機関: University of Michigan
所属の国: United States of America
会議名: Computing Research Repository (CoRR)

AIにより推定されたラベル

プロンプトインジェクション悪意のあるプロンプト LLMセキュリティ

※ こちらのラベルはAIによって自動的に追加されました。そのため、正確でないことがあります。
詳細は文献データベースについてをご覧ください。

Abstract

A novel hack involving Large Language Models (LLMs) has emerged, exploiting adversarial suffixes to deceive models into generating perilous responses. Such jailbreaks can trick LLMs into providing intricate instructions to a malicious user for creating explosives, orchestrating a bank heist, or facilitating the creation of offensive content. By evaluating the perplexity of queries with adversarial suffixes using an open-source LLM (GPT-2), we found that they have exceedingly high perplexity values. As we explored a broad range of regular (non-adversarial) prompt varieties, we concluded that false positives are a significant challenge for plain perplexity filtering. A Light-GBM trained on perplexity and token length resolved the false positives and correctly detected most adversarial attacks in the test set.

外部データセット

machine-generated adversarial prompts

human-designed adversarial prompts

Puffin

DocRED

SuperGLUE

SQuAD-v2

Platypus

Tapir

instructional code-search-net-python