Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts Authors: Rui Yang, Yang Hong, Yichao Xu, Zhengyu Liu, Ziyang Li, Yinzhi Cao | Published: 2026-09-01 データセットの影響実験プロトコル攻撃モデルの訓練 2026.09.01 2026.09.03 Literature Database
The Safeguard Worked. Is the LLM System Safer? Authors: Pingyu Wu, Weiming Zhang, Nenghai Yu | Published: 2026-09-01 Indirect Prompt InjectionPrompt InjectionAttacker Behavior Analysis 2026.09.01 2026.09.03 Literature Database
SingProbe Technical Report Authors: Sing Team | Published: 2026-08-31 HallucinationMedical Monitoring System評価結果 2026.08.31 2026.09.02 Literature Database
ECLIPSE: Self-Evolving Stealthy Prompt Injection Attack against Long-Horizon Agentic Systems Authors: Shiqian Zhao, Yangfan Zhou, Xinfeng Li, Runyi Hu, Yechao Zhang, Yi Xie, Tianwei Zhang, Luu Anh Tuan | Published: 2026-08-31 Indirect Prompt Injectionセキュリティ検証手法攻撃モデルの訓練 2026.08.31 2026.09.02 Literature Database
EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents Authors: Doyun Kim, Chanwoo Kim, Sugyeong Eo, Yeo-Chan Yoon, Chanjun Park | Published: 2026-08-31 Indirect Prompt InjectionPrompt Injection能力進化の脅威 2026.08.31 2026.09.02 Literature Database
Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders Authors: Yizhe Zeng, Chenxu Niu, Wei Zhang, Hao Huang, Yunpeng Li, Dongxu Han, Dan Du, Cheng Hong, Hequn Xian, Yuling Liu | Published: 2026-08-31 Backdoor DetectionPoisoning Attack防御手法の統合 2026.08.31 2026.09.02 Literature Database
Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents Authors: Yunseok Lee, Yunji Kim, Woojin Lee | Published: 2026-08-31 Indirect Prompt Injection攻撃モデルの訓練 2026.08.31 2026.09.02 Literature Database
SIR: Self-improving Red-teaming for Compute Use Agents Authors: Chen Xiong, Zhiyuan He, Pin-Yu Chen, Stjepan Picek, Tsung-Yi Ho | Published: 2026-08-31 サイバーセキュリティの脅威Exploratory Attack攻撃モデルの訓練 2026.08.31 2026.09.02 Literature Database
Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization Authors: Dishu Yang, Jingjing Liu, Jize Li | Published: 2026-08-31 Alignmentプライバシー評価手法Measurement of Memorization 2026.08.31 2026.09.02 Literature Database
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution Authors: Junjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo, Changsheng Chen, Haoliang Li | Published: 2026-08-27 Prompt Injection攻撃効果の評価脱獄攻撃手法 2026.08.27 2026.08.29 Literature Database