When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Authors: Tong Zhang, Zexin Li, Simin Chen, Yun Peng | Published: 2026-07-27 Disabling Safety Mechanisms of LLMPrompt leakingModel Protection Methods 2026.07.27 2026.07.29 Literature Database
DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense Authors: Trung V. Phan, Tri Gia Nguyen, Thomas Bauschert | Published: 2026-07-27 Cooperative Effects with LLMPrompt InjectionEvaluation of Explanatory Approaches 2026.07.27 2026.07.29 Literature Database
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls Authors: Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Khandaker Rifah Tasnia, Sifat Rahman Ahona, Juena Ahmed Noshin | Published: 2026-07-27 Cooperative Effects with LLMModel Protection MethodsRisk Assessment 2026.07.27 2026.07.29 Literature Database
Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families Authors: Dushyant Sharma | Published: 2026-07-27 Relationship of AI SystemsCooperative Effects with LLM制御システム 2026.07.27 2026.07.29 Literature Database
Just Testing, Move Along: Evasion of LLM-based System Log Interpretation by Prompt Injection Authors: Max Landauer, Florian Skopik, Markus Wurzenberger, Franciszek Górski, Mateusz Krzysztoń | Published: 2026-07-27 Indirect Prompt Injection攻撃手法の効果Attacks on Explainability 2026.07.27 2026.07.29 Literature Database
A Cybersecurity MLPS Large Language Model with Multi-Path Retrieval Fusion Authors: Qian Li, Zhenyan Qi, Liang Shen, Yuan Zhang, Yifan Wan, Junyuan Ma, Yining Hu | Published: 2026-07-27 Cooperative Effects with LLMRAGRisk Assessment 2026.07.27 2026.07.29 Literature Database
Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation Authors: Mohan Manivannan, Dalal Alharthi | Published: 2026-07-27 Indirect Prompt Injection報告生成評価結果 2026.07.27 2026.07.29 Literature Database
Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models Authors: Tapan Parikh | Published: 2026-07-27 RAGModel CommunicationUser Behavior Analysis 2026.07.27 2026.07.29 Literature Database
Understanding Machine Unlearning Through the Lens of Mode Connectivity Authors: Jiali Cheng, Hadi Amiri | Published: 2026-07-27 Dataset evaluationModel Protection MethodsMachine Learning 2026.07.27 2026.07.29 Literature Database
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure Authors: Zhetong Zhang, Honghao Fu, Miao Xu, Yiwei Wang, Yujun Cai | Published: 2026-07-23 User Behavior AnalysisRisk Assessment攻撃手法の効果 2026.07.23 2026.07.25 Literature Database