Systematic Categorization, Construction and Evaluation of New Attacks against Multi-modal Mobile GUI Agents

TOP Literature Database Systematic Categorization, Construction and Evaluation of New Attacks against Multi-modal Mobile GUI Agents

arxiv

AI Security Portal bot

Information in the literature database is collected automatically.

Source

https://arxiv.org/abs/2407.09295

PDF

https://arxiv.org/pdf/2407.09295

Paper Information

Author: Yulong Yang,Xinshan Yang,Shuaidong Li,Chenhao Lin,Zhengyu Zhao,Chao Shen,Tianwei Zhang
Published: 7-12-2024
Updated: 3-16-2025
Affiliation: Xi’an Jiaotong University
Country: China
Conference

Labels Estimated by AI

Indirect Prompt Injection Attack Method Vulnerability Attack Method

These labels were automatically added by AI and may be inaccurate.
For details, see About Literature Database.

Abstract

The integration of Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) into mobile GUI agents has significantly enhanced user efficiency and experience. However, this advancement also introduces potential security vulnerabilities that have yet to be thoroughly explored. In this paper, we present a systematic security investigation of multi-modal mobile GUI agents, addressing this critical gap in the existing literature. Our contributions are twofold: (1) we propose a novel threat modeling methodology, leading to the discovery and feasibility analysis of 34 previously unreported attacks, and (2) we design an attack framework to systematically construct and evaluate these threats. Through a combination of real-world case studies and extensive dataset-driven experiments, we validate the severity and practicality of those attacks, highlighting the pressing need for robust security measures in mobile GUI systems.

External Datasets

Android-in-the-Wild (AITW)