EnJa: Ensemble Jailbreak on Large Language Models

TOP Literature Database EnJa: Ensemble Jailbreak on Large Language Models

arxiv

AI Security Portal bot

Information in the literature database is collected automatically.

Source

https://arxiv.org/abs/2408.03603

PDF

https://arxiv.org/pdf/2408.03603

Paper Information

Author: Jiahao Zhang;Zilong Wang;Ruofan Wang;Xingjun Ma;Yu-Gang Jiang
Published: 8-7-2024
Affiliation: School of Computer Science, Fudan University
Country: China
Conference: Computing Research Repository (CoRR)

Labels Estimated by AI

Prompt Injection Attack Method Evaluation Method

These labels were automatically added by AI and may be inaccurate.
For details, see About Literature Database.

Abstract

As Large Language Models (LLMs) are increasingly being deployed in safety-critical applications, their vulnerability to potential jailbreaks -- malicious prompts that can disable the safety mechanism of LLMs -- has attracted growing research attention. While alignment methods have been proposed to protect LLMs from jailbreaks, many have found that aligned LLMs can still be jailbroken by carefully crafted malicious prompts, producing content that violates policy regulations. Existing jailbreak attacks on LLMs can be categorized into prompt-level methods which make up stories/logic to circumvent safety alignment and token-level attack methods which leverage gradient methods to find adversarial tokens. In this work, we introduce the concept of Ensemble Jailbreak and explore methods that can integrate prompt-level and token-level jailbreak into a more powerful hybrid jailbreak attack. Specifically, we propose a novel EnJa attack to hide harmful instructions using prompt-level jailbreak, boost the attack success rate using a gradient-based attack, and connect the two types of jailbreak attacks via a template-based connector. We evaluate the effectiveness of EnJa on several aligned models and show that it achieves a state-of-the-art attack success rate with fewer queries and is much stronger than any individual jailbreak.

External Datasets

AdvBench

AdvBench Subset