文献データベース

Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

Authors: Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah, Yazan Boshmaf | Published: 2026-09-08
プロンプトインジェクション
プロンプトリーキング
評価結果

Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning

Authors: Mohammed-Yassine Habibi, Klea Ziu, Martin Takáč, Makoto Yamada | Published: 2026-09-08
モデルの堅牢性
敵対的学習
自己監視モデル

Windows Malware Detector as a Compound AI System: Trade-Offs in Accuracy, Efficiency, and Adversarial Robustness

Authors: Andrea Ponte, Luca Demetrio, Luca Oneto, Battista Biggio, Fabio Roli | Published: 2026-09-08
バックドアモデルの検知
マルウェア分類のためのデータセット
敵対的サンプルの検知

Do Input-Level Defenses Transfer to Observation-Level Attacks on VideoLLMs?

Authors: Bangshuo Zhu, Wei Song, Yuxin Cao, Yuezhong Wu, Zhiquan Liu, Yuekang Li, Jingling Xue | Published: 2026-09-08
モデルの堅牢性
攻撃手法の効果
評価結果

Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems

Authors: Yi Ting Shen, Kentaroh Toyoda, Alex Leung | Published: 2026-09-08
攻撃手法の効果
監査手法
評価結果

ACEA: An Adversarial Co-Evolution Arena for Head-to-Head Red-Team and Blue-Team LLM Testing

Authors: Yi Ting Shen, Kentaroh Toyoda, Alex Leung | Published: 2026-09-08
攻撃シナリオ分析
攻撃手法の効果
評価結果

Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts

Authors: Yongxi Zhou, Wenbo Ye, Yuanzhe Liu, Zihan Dong, Junwei Yao | Published: 2026-09-08
エラー解析
監査手法
評価結果

Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models

Authors: Yu-Hang Wu, Yu-Jie Xiong, Henghua Zhang, Bairui Zhang, Jia-Chen Zhang, Shaohua Li | Published: 2026-09-08
モデルの堅牢性
攻撃手法の効果
監査手法

LLM-Based Penetration Testing in the Presence of Honeypots

Authors: Xinhong Xie, Piyush Nagasubramaniam, Neeraj Karamchandani, Sencun Zhu | Published: 2026-09-08
ハニーポット技術
プロンプトインジェクション
評価結果

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

Authors: Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild | Published: 2026-09-03
サイバーセキュリティ
ユーザー認証システム
監査手法