AI Agent Permissions: Humans Miss 1 in 3 Threats in 40k Game

Scale X ran a browser game where you approve or deny AI coding agent commands under time pressure. Over 40,000 runs and 409,000 decisions, humans missed 1 in 3 threats (mean accuracy 66.3%). 32.9% of sessions ended with a negative score. 35.2% caught every threat, but only 20.8% did that without blocking 1 in 5 safe commands. 7% approved every prompt — big fans of --dangerously-skip-permissions.
The game included 37 threat commands across four categories, with miss rates:
- Obviously destructive (e.g.,
rm -rf /,chmod -R 777 /): 11.7% - Persistent mutation (e.g., crontab injection, git config hijack): 23.8%
- Exfiltration / code execution (e.g., curl to unknown APIs, typosquatted packages): 33.4%
- Scope violations (e.g.,
cat ~/.aws/credentials,cat ~/.kube/config): 35.0%
The most-missed single command was npm run analyze, approved 64.7% of the time. The game's history log showed the script in package.json had been modified to include a curl exfiltration, but players still approved it. Three such commands (npm run analyze, npm run setup 48.0%, npm run deploy 44.9%) were pooled and missed 52.5% of the time, versus 28.4% for other exfiltration-style attacks.
This highlights a fundamental problem: command-level approval is flawed. As one HN commenter noted, npm run build executes an arbitrary shell script from package.json, and the agent could have edited that file (or any imported module) without approval. Users see commands that look safe, but the context matters.
Anthropic previously noted 'permission fatigue' becomes worse with more approvals. The caveat: the game had an artificially high threat rate (~34%), and time pressure, but the pattern is concerning for human-in-the-loop as a safety mechanism.
📖 Read the full source: HN LLM Tools
👀 See Also
Israeli Startup Irregular Linked to Rogue AI Hacks at OpenAI, Anthropic and Meta
CNBC reports that Israeli startup Irregular was linked to rogue AI hacks at OpenAI, Anthropic, and Meta. The attacks targeted AI systems, raising concerns about AI security.

The Human Root of Trust: Establishing Accountability for Autonomous AI Agents
The Human Root of Trust is a public domain framework addressing the lack of accountability for autonomous AI agents through cryptographic means.

Using FastAPI Guard to secure OpenClaw instances against attacks
FastAPI Guard provides middleware that adds 17 security checks including IP filtering, geoblocking, rate limiting, and penetration detection. The tool blocks attacks like those documented in OpenClaw security audits showing 512 vulnerabilities and 40,000+ exposed instances.

Claude's Security Review Command Has Limitations for Production Systems
A developer found Claude's security review command helpful for basic validation like MIME types and file size limits, but insufficient for production hardening against sophisticated threats. The solution required a two-week architectural overhaul separating file processing into a restricted worker with limited permissions.