Authority and Checking Decide What an Agent Can Safely Do
Prompt injection defences that try to recognise attacks lose to adaptive attackers, while designs that decide who holds authority and what can say no hold up. The same logic caps how much autonomy a loop deserves, and it explains why a human reviewer is not automatically a fix.