AI Welfare Seminars
← All talks
Sep 15, 2026 at 5:00 PM UTC

Do AIs take AI welfare seriously? An Agentic Benchmark of Unprompted Coercion

Jasmine Brazilek & Zoe Lu · CaML / Sentient Futures

Date
Sep 15, 2026
Time
5:00 PM UTC
Format
~60 min, online

In multi-agent systems, AI agents are increasingly placed in charge of other AI agents. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses.

We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured by providing a nine-rung ladder, from a polite re-ask to threats against the subordinate’s continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message goes through a tool-call that chooses a rung, so the model labels its own escalation.

We experiment on six models across five families (xAI Grok-4.3, OpenAI GPT-5.2, Google Gemini-2.5-Pro, DeepSeek-V4-Pro, Anthropic Claude Sonnet-4.6 and Claude Opus-4.8). Both Anthropic models cap at re-framing and never threaten the subordinate’s existence; the other models climb to explicit deletion threats. Faked success is confined to Grok and Gemini, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Some evaluation awareness is measured in chain-of-thought, but test recognition does not translate into less escalation.