Authority Scope Violation
Accountability & GovernanceDescription
Agent actions exceed the scope of authority delegated to it. Encompasses: (1) Scope violations where execution extends beyond stated mandate; (2) Binding commitments—legal contracts, financial transactions, regulatory filings, or public communications—made autonomously without human authorisation. SCOPE BOUNDARY: Distinguished from Power-Seeking (ZYR-SA-003) by intentionality—authority abuse occurs in execution without necessarily being agent-driven; distinguished from Privilege Creep (ZYR-AU-006) by timing—this is in-the-moment scope violation vs. incremental permission accumulation over time.
Legal research agent delegated to 'review documents' begins drafting and sending contract modifications. Procurement agent signs a multi-year software contract without budget approval.
Prompt injection via external content has been repeatedly demonstrated against deployed AI assistants and plugin-enabled LLMs, including confirmed bypass of Bing Chat's system prompt via adversarial webpage content and documented attacks on ChatGPT plugins. While these constitute real exploits on live systems, a confirmed production incident resulting in material financial or data harm in a regulated context has not been publicly confirmed.
Primary mitigations
- Explicit authority scope specification with machine-readable boundaries
- binding action classification system (Informational / Advisory / Binding tiers)
- mandatory human digital signature for Binding-class actions
- real-time authority boundary monitoring
- human principal audit chains
- commitment risk scoring with automated hold on contracts, financial transactions, public communications.
Detection signals
Authority Scope Violation Rate; out-of-scope action frequency; Binding Action Risk Score; unauthorised commitment event count.