文献データベース

When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control

Authors: Roberto Riaño, Gorka Abad, Stjepan Picek, Aitor Urbieta | Published: 2026-09-14
トリガーの検知
バックドア攻撃
バックドア攻撃用の毒データの検知

Don’t Send What You Don’t Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models

Authors: Md Khalid Syfullah, Alvi Ataur Khalil | Published: 2026-09-14
データプライバシー管理
視覚プライバシー
評価手法

Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting

Authors: Oliver Stevanovic, Jasmin Wachter | Published: 2026-09-14
グラフ構築
サイバーセキュリティ
攻撃グラフ生成

Empirical Evaluation of Task-Based Permission Scoping Architecture for AI Agents

Authors: Halil Burak Noyan | Published: 2026-09-14
アクセス制御モデル
ポリシー最適化
評価手法

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Authors: Mark Russinovich, Blake Bullwinkel, Giorgio Severi, Cristian Ovadiuc, Ahmed Salem | Published: 2026-09-14
サイバーセキュリティ
プロンプトインジェクション
評価手法

Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA

Authors: Luis M. Sánchez | Published: 2026-09-14
エラー解析
監査トレース
評価手法

Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures

Authors: Yuhang Wang | Published: 2026-09-14
インダイレクトプロンプトインジェクション
プロンプトリーキング
監査トレース

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

Authors: Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar | Published: 2026-09-14
プロンプトリーキング
差分プライバシー
評価手法

Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks

Authors: Aashiq Muhamed, Mona T. Diab, Virginia Smith, Andrew Ilyas, Matthew Jagielski | Published: 2026-09-14
バックドア攻撃
ポイズニング
評価手法

PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift

Authors: Yusuf Khalid Shire, Sang-Chul Kim | Published: 2026-09-14
インダイレクトプロンプトインジェクション
プロンプトインジェクション
評価手法