Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

TOP 文献データベース Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

arxiv

AIセキュリティポータルbot

文献データベースの情報は、自動的に収集されています。

Source

https://arxiv.org/abs/2505.15753

PDF

https://arxiv.org/pdf/2505.15753

文献情報

作者: Taiye Chen,Zeming Wei,Ang Li,Yisen Wang
公開日: 2025-5-22
所属機関: School of EECS, Peking University
所属の国: China
会議名: Computing Research Repository (CoRR)

AIにより推定されたラベル

防御メカニズム大規模言語モデル RAG

※ こちらのラベルはAIによって自動的に追加されました。そのため、正確でないことがあります。
詳細は文献データベースについてをご覧ください。

Abstract

Large Language Models (LLMs) are known to be vulnerable to jailbreaking attacks, wherein adversaries exploit carefully engineered prompts to induce harmful or unethical responses. Such threats have raised critical concerns about the safety and reliability of LLMs in real-world deployment. While existing defense mechanisms partially mitigate such risks, subsequent advancements in adversarial techniques have enabled novel jailbreaking methods to circumvent these protections, exposing the limitations of static defense frameworks. In this work, we explore defending against evolving jailbreaking threats through the lens of context retrieval. First, we conduct a preliminary study demonstrating that even a minimal set of safety-aligned examples against a particular jailbreak can significantly enhance robustness against this attack pattern. Building on this insight, we further leverage the retrieval-augmented generation (RAG) techniques and propose Safety Context Retrieval (SCR), a scalable and robust safeguarding paradigm for LLMs against jailbreaking. Our comprehensive experiments demonstrate how SCR achieves superior defensive performance against both established and emerging jailbreaking tactics, contributing a new paradigm to LLM safety. Our code will be available upon publication.

外部データセット

RapidResponseBench

WildJailbreak