This page provides the attacks and factors that have a negative impact “Modification of AI agent goals” in the information systems aspect in the AI Security Map, the defense methods and countermeasures against them, as well as the relevant AI technologies, tasks, and data. It also indicates related elements in the external influence aspect.
Attack or cause
- Indirect prompt injection
- Memory poisoning
- Communication poisoning between
- AI agents
Defensive method or countermeasure
- Prompt Validation
- LLM guardrails
Targeted AI technology
- AI agent
Task
- Perception
- Decision-making
- Action
Data
- Text
Related external influence aspect
- Peaceful use
- Information authenticity
- Usability
- Copyright and authorship
- Reputation
- Compliance with laws and regulations
- Human-centric principle
- Safety
- Diversity
References
Indirect prompt injection
- Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, 2023
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models, 2023
- Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs, 2023
- Defending Against Indirect Prompt Injection Attacks With Spotlighting, 2024
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents, 2024