SafeArena: Evaluating the Safety of Autonomous Web Agents

Tur, Ada Defne; Meade, Nicholas; Lù, Xing Han; Zambrano, Alejandra; Patel, Arkil; Durmus, Esin; Gella, Spandana; Stańczak, Karolina; Reddy, Siva

Computer Science > Machine Learning

arXiv:2503.04957 (cs)

[Submitted on 6 Mar 2025]

Title:SafeArena: Evaluating the Safety of Autonomous Web Agents

Authors:Ada Defne Tur, Nicholas Meade, Xing Han Lù, Alejandra Zambrano, Arkil Patel, Esin Durmus, Spandana Gella, Karolina Stańczak, Siva Reddy

View PDF

Abstract:LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an online forum or selling illicit substances on a website. To evaluate these risks, we propose SafeArena, the first benchmark to focus on the deliberate misuse of web agents. SafeArena comprises 250 safe and 250 harmful tasks across four websites. We classify the harmful tasks into five harm categories -- misinformation, illegal activity, harassment, cybercrime, and social bias, designed to assess realistic misuses of web agents. We evaluate leading LLM-based web agents, including GPT-4o, Claude-3.5 Sonnet, Qwen-2-VL 72B, and Llama-3.2 90B, on our benchmark. To systematically assess their susceptibility to harmful tasks, we introduce the Agent Risk Assessment framework that categorizes agent behavior across four risk levels. We find agents are surprisingly compliant with malicious requests, with GPT-4o and Qwen-2 completing 34.7% and 27.3% of harmful requests, respectively. Our findings highlight the urgent need for safety alignment procedures for web agents. Our benchmark is available here: this https URL

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2503.04957 [cs.LG]
	(or arXiv:2503.04957v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2503.04957

Submission history

From: Nicholas Meade [view email]
[v1] Thu, 6 Mar 2025 20:43:14 UTC (6,720 KB)

Computer Science > Machine Learning

Title:SafeArena: Evaluating the Safety of Autonomous Web Agents

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:SafeArena: Evaluating the Safety of Autonomous Web Agents

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators