When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Authors: Tong Zhang, Zexin Li, Simin Chen, Yun Peng | Published: 2026-07-27 LLMの安全機構の解除プロンプトリーキングモデル保護手法 2026.07.27 文献データベース
DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense Authors: Trung V. Phan, Tri Gia Nguyen, Thomas Bauschert | Published: 2026-07-27 LLMとの協力効果プロンプトインジェクション説明アプローチの評価 2026.07.27 文献データベース
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls Authors: Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Khandaker Rifah Tasnia, Sifat Rahman Ahona, Juena Ahmed Noshin | Published: 2026-07-27 LLMとの協力効果モデル保護手法リスク評価 2026.07.27 文献データベース
Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families Authors: Dushyant Sharma | Published: 2026-07-27 AIシステムの関係性LLMとの協力効果制御システム 2026.07.27 文献データベース
Just Testing, Move Along: Evasion of LLM-based System Log Interpretation by Prompt Injection Authors: Max Landauer, Florian Skopik, Markus Wurzenberger, Franciszek Górski, Mateusz Krzysztoń | Published: 2026-07-27 インダイレクトプロンプトインジェクション攻撃手法の効果説明可能性に対する攻撃 2026.07.27 文献データベース
A Cybersecurity MLPS Large Language Model with Multi-Path Retrieval Fusion Authors: Qian Li, Zhenyan Qi, Liang Shen, Yuan Zhang, Yifan Wan, Junyuan Ma, Yining Hu | Published: 2026-07-27 LLMとの協力効果RAGリスク評価 2026.07.27 文献データベース
Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation Authors: Mohan Manivannan, Dalal Alharthi | Published: 2026-07-27 インダイレクトプロンプトインジェクション報告生成評価結果 2026.07.27 文献データベース
Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models Authors: Tapan Parikh | Published: 2026-07-27 RAGモデル通信ユーザー行動分析 2026.07.27 文献データベース
Understanding Machine Unlearning Through the Lens of Mode Connectivity Authors: Jiali Cheng, Hadi Amiri | Published: 2026-07-27 データセット評価モデル保護手法機械学習 2026.07.27 文献データベース
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure Authors: Zhetong Zhang, Honghao Fu, Miao Xu, Yiwei Wang, Yujun Cai | Published: 2026-07-23 ユーザー行動分析リスク評価攻撃手法の効果 2026.07.23 文献データベース