Bounded Agents: Delegation Security for Multi-Agent AI Systems
Most LLM agent security checks operate action-by-action — but the most dangerous attacks chain two individually-permitted actions into a prohibited outcome: read a confidential file, then send it via email. Bounded Agents stops this class of attack by tracking authorization state across the entire delegation trajectory, not just at each individual action boundary.
The paper introduces the Agentic Principal Chain (APC), a runtime mechanism that attaches signed session-level state to every inter-agent delegation: principal identity, scope, delegation budget, declared intent, and action history. Six-condition enforcement at each step prevents any combination of individually-allowed actions from yielding a prohibited outcome — the key contribution over per-action checks. Benchmark results: data exfiltration attacks reduced from 100% → 0% under complete restrictions; 140 of 200 stealthy attacks blocked; AgentDojo exfiltration rate → 0%; utility cost 8.6–13.9 pp; authorization overhead 0.05 ms median — negligible in deployed settings.