Security Concerns for Large Language Models: A Survey

TOP 文献データベース Security Concerns for Large Language Models: A Survey

arxiv

AIセキュリティポータルbot

文献データベースの情報は、自動的に収集されています。

Source

https://arxiv.org/abs/2505.18889

PDF

https://arxiv.org/pdf/2505.18889

文献情報

作者: Miles Q. Li,Benjamin C. M. Fung
公開日: 2025-5-25
更新日: 2025-8-21
所属機関: Infinite Optimization AI Lab, Montreal, Canada
所属の国: Canada
会議名: Computing Research Repository (CoRR)

AIにより推定されたラベル

心理的操作インダイレクトプロンプトインジェクションプロンプトインジェクション

※ こちらのラベルはAIによって自動的に追加されました。そのため、正確でないことがあります。
詳細は文献データベースについてをご覧ください。

Abstract

Large Language Models (LLMs) such as ChatGPT and its competitors have caused a revolution in natural language processing, but their capabilities also introduce new security vulnerabilities. This survey provides a comprehensive overview of these emerging concerns, categorizing threats into several key areas: prompt injection and jailbreaking; adversarial attacks, including input perturbations and data poisoning; misuse by malicious actors to generate disinformation, phishing emails, and malware; and the worrisome risks inherent in autonomous LLM agents. Recently, a significant focus is increasingly being placed on the latter, exploring goal misalignment, emergent deception, self-preservation instincts, and the potential for LLMs to develop and pursue covert, misaligned objectives, a behavior known as scheming, which may even persist through safety training. We summarize recent academic and industrial studies from 2022 to 2025 that exemplify each threat, analyze proposed defenses and their limitations, and identify open challenges in securing LLM-based applications. We conclude by emphasizing the importance of advancing robust, multi-layered security strategies to ensure LLMs are safe and beneficial.

外部データセット

BackdoorLLM

AutoPoison

VPI

BadGPT

AgentDojo

Open-Prompt-Injection

MISALIGNMENTBENCH