AIセキュリティポータルbot

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

Authors: Hanwool Lee, Dasol Choi, Bokyeong Kim, Seung Geun Kim, Haon Park | Published: 2026-06-18
Indirect Prompt Injection
Robustness
Large Language Model

Quantization as a Malicious Task: Removing Quantization-Conditioned Backdoors via Task Arithmetic

Authors: Kaihsun Yang, Min-Yan Tsai, Chia-Mu Yu | Published: 2026-06-18
Backdoor Attack Mitigation
Model Extraction Attack
Adversarial Learning

Accelerating Trust Convergence in IIoT: A ML Approach for Dynamic Network Conditions

Authors: Aymen Bouferroum, Valeria Loscri, Abderrahim Benslimane | Published: 2026-06-18
ネットワーク条件モデリング
信頼管理手法
Self-Learning Method

Artificial Intelligence as Game Changer in Cybersecurity: What We Learned in 2025-2026, and how this is relevant for Africa

Authors: Mikael Alemu Gorsky | Published: 2026-06-18
Relationship of AI Systems
LLM Application
Generative AI in Financial Services

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

Authors: Kaiyue Yang, Yuyan Bu, Jingwei Yi, Yuchi Wang, Biyu Zhou, Juntao Dai, Songlin Hu, Yaodong Yang | Published: 2026-06-18
アクセス制御モデル
Large Language Model
Operational Scenario

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling

Authors: Haotian Xu, Zeyang Zhang, Linbao Li, Huadi Zheng, Yu Li, Cheng Zhuo | Published: 2026-06-18
Prompt Injection
Large Language Model
安全性向上手法

CodeSentinel: A Three-Layer Defense Against Indirect Prompt Injection in Code Contexts

Authors: Po-Han Cheng, Chia-Mu Yu, Ying-Dar Lin, Yu-Sung Wu, Wei-Bin Lee | Published: 2026-06-17
Prompt validation
Adversarial attack
脆弱性検出手法

Generalised Eigenvalue Geometry of Semantic Adversarial Attacks

Authors: Martin Anthony, Kaveh Salehzadeh Nobari | Published: 2026-06-17
Model Robustness
Robustness Analysis
Self-Supervised Learning

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection

Authors: Jinhan Li, Kexian Tang, Yihan Xu, Zhuorui Ye, Kaifeng Lyu | Published: 2026-06-17
Model Robustness
安全性反映トレーニング
Adversarial Learning

OpenAnt: LLM-Powered Vulnerability Discovery Through Code Decomposition, Adversarial Verification, and Dynamic Testing

Authors: Nahum Korda, Gadi Evron | Published: 2026-06-17
Prompt leaking
脆弱性検出手法
評価基準