Steven Gonsalvez

Software Engineer

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

Why CEREBRO kept it

Agent UX/safety study on approval overhead

The text below is an automated extraction of the article at https://scalex.dev/blog/ai-agent-permissions-stats/, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (scalex.dev).

Table of Contents A couple of months ago I published a small browser game: you play the human-in-the-loop for an AI coding agent, approving or denying its commands under time pressure. Some commands are routine (git status, npm test) and some other commands indicate your agent has been possessed and is sending your secrets to a remote server (cat ~/.aws/credentials). More on the threats associated with agents running commands and how to mitigate them can be found in the original post. The game garnered some interest on hacker news, and after adding in statistics (unfortunately a bit later on)

Community take

Permission prompts are vendor legal cover, not real security; 1 in 3 misses is expected—sandbox with file-level resource controls beats command approvals.

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/agents

Also from scalex.dev

Only signal from scalex.dev so far.