Back to papers
April 3, 2026cs.AI

AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

HF Upvotes

12

Categories

cs.AI