Literature Database

Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts

Authors: Rui Yang, Yang Hong, Yichao Xu, Zhengyu Liu, Ziyang Li, Yinzhi Cao | Published: 2026-09-01
データセットの影響
実験プロトコル
攻撃モデルの訓練

The Safeguard Worked. Is the LLM System Safer?

Authors: Pingyu Wu, Weiming Zhang, Nenghai Yu | Published: 2026-09-01
Indirect Prompt Injection
Prompt Injection
Attacker Behavior Analysis

SingProbe Technical Report

Authors: Sing Team | Published: 2026-08-31
Hallucination
Medical Monitoring System
評価結果

ECLIPSE: Self-Evolving Stealthy Prompt Injection Attack against Long-Horizon Agentic Systems

Authors: Shiqian Zhao, Yangfan Zhou, Xinfeng Li, Runyi Hu, Yechao Zhang, Yi Xie, Tianwei Zhang, Luu Anh Tuan | Published: 2026-08-31
Indirect Prompt Injection
セキュリティ検証手法
攻撃モデルの訓練

EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

Authors: Doyun Kim, Chanwoo Kim, Sugyeong Eo, Yeo-Chan Yoon, Chanjun Park | Published: 2026-08-31
Indirect Prompt Injection
Prompt Injection
能力進化の脅威

Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders

Authors: Yizhe Zeng, Chenxu Niu, Wei Zhang, Hao Huang, Yunpeng Li, Dongxu Han, Dan Du, Cheng Hong, Hequn Xian, Yuling Liu | Published: 2026-08-31
Backdoor Detection
Poisoning Attack
防御手法の統合

Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents

Authors: Yunseok Lee, Yunji Kim, Woojin Lee | Published: 2026-08-31
Indirect Prompt Injection
攻撃モデルの訓練

SIR: Self-improving Red-teaming for Compute Use Agents

Authors: Chen Xiong, Zhiyuan He, Pin-Yu Chen, Stjepan Picek, Tsung-Yi Ho | Published: 2026-08-31
サイバーセキュリティの脅威
Exploratory Attack
攻撃モデルの訓練

Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization

Authors: Dishu Yang, Jingjing Liu, Jize Li | Published: 2026-08-31
Alignment
プライバシー評価手法
Measurement of Memorization

RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution

Authors: Junjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo, Changsheng Chen, Haoliang Li | Published: 2026-08-27
Prompt Injection
攻撃効果の評価
脱獄攻撃手法