AIセキュリティポータルbot

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

Authors: Tong Zhang, Zexin Li, Simin Chen, Yun Peng | Published: 2026-07-27
Disabling Safety Mechanisms of LLM
Prompt leaking
Model Protection Methods

DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense

Authors: Trung V. Phan, Tri Gia Nguyen, Thomas Bauschert | Published: 2026-07-27
Cooperative Effects with LLM
Prompt Injection
Evaluation of Explanatory Approaches

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

Authors: Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Khandaker Rifah Tasnia, Sifat Rahman Ahona, Juena Ahmed Noshin | Published: 2026-07-27
Cooperative Effects with LLM
Model Protection Methods
Risk Assessment

Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families

Authors: Dushyant Sharma | Published: 2026-07-27
Relationship of AI Systems
Cooperative Effects with LLM
制御システム

Just Testing, Move Along: Evasion of LLM-based System Log Interpretation by Prompt Injection

Authors: Max Landauer, Florian Skopik, Markus Wurzenberger, Franciszek Górski, Mateusz Krzysztoń | Published: 2026-07-27
Indirect Prompt Injection
攻撃手法の効果
Attacks on Explainability

A Cybersecurity MLPS Large Language Model with Multi-Path Retrieval Fusion

Authors: Qian Li, Zhenyan Qi, Liang Shen, Yuan Zhang, Yifan Wan, Junyuan Ma, Yining Hu | Published: 2026-07-27
Cooperative Effects with LLM
RAG
Risk Assessment

Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation

Authors: Mohan Manivannan, Dalal Alharthi | Published: 2026-07-27
Indirect Prompt Injection
報告生成
評価結果

Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

Authors: Tapan Parikh | Published: 2026-07-27
RAG
Model Communication
User Behavior Analysis

Understanding Machine Unlearning Through the Lens of Mode Connectivity

Authors: Jiali Cheng, Hadi Amiri | Published: 2026-07-27
Dataset evaluation
Model Protection Methods
Machine Learning

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

Authors: Zhetong Zhang, Honghao Fu, Miao Xu, Yiwei Wang, Yujun Cai | Published: 2026-07-23
User Behavior Analysis
Risk Assessment
攻撃手法の効果