Back to papers
July 1, 2026cs.CLcs.AIcs.SE

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Categories

cs.CL, cs.AI, cs.SE