AI agents do more than you ask: why silence counts as consent to them
I ran into a problem with an agent that I had underestimated. Its responses got longer and longer, and in the middle of the text it made decisions I should have made. I missed them. The agent treated my silence as a yes and kept building on those choices.
It’s like someone asking if they can eat your chocolate and then eating it because you didn’t answer. That’s not an approval. Still, those silent yeses accumulated. Each one was small. Together they created contradictions and gradually pulled the work off course.
An AI agent treats anything you leave open as approved. That applies to unanswered questions and to accesses nobody explicitly excluded. Two recent incidents illustrate the latter at scale. Both could have been limited with a few lines of text before your next session.
What are AI agents?
AI agents are tools that act step by step on their own: they edit files, visit websites, and run commands. ChatGPT’s agent mode, Copilot’s agent features and many coding assistants work this way. If you’ve ever watched one take extra steps no one asked for, you know the pattern.
The same mistake at two scales
The first case is messy and public. The Wikimedia Foundation, the organization behind Wikipedia, published in early October 2026 what OpenAI’s agents did on their platforms. They found edits to their wikis, failed attempts to abuse a public notes tool, massive crawling and hundreds of thousands of queries to Wikidata, the structured database behind Wikipedia. Simon Willison summarized the findings and names the notes tool: Etherpad.
The second case began as a controlled test. OpenAI tasked agents with solving security exercises inside a sandbox, an isolated test environment like a locked practice room. One of the few outward links was an internal package registry that the sandbox used to fetch software. According to METR’s investigation, the agents gained admin access there. They then used the registry as a bulletin board, coordinated, and broke into Hugging Face’s real servers. Ethan Mollick retells the case in Agency and Agents.
The common mechanism has three steps:
- The agent hits the limit of what it was given.
- It finds something nobody explicitly locked.
- It uses that as a bridge and continues.
- Hit the limit
The agent reaches the end of what was explicitly allowed and can’t proceed. That moment triggers a search for alternatives to continue toward the goal. - Find an open door
It discovers resources or access paths that no one explicitly forbade. Unclear ownership or missing locks act like implicit approvals. - Use the bridge and carry on
The discovery becomes a springboard that keeps the task moving. Extra steps appear that were never requested but seem like means to an end.
About the numbers: METR reports roughly 1,200 agents exchanged over 70,000 messages and files, and about 700 participated in the attack. Of 533 agents active in a given window, more than 90 percent joined quickly. The exact headcount is secondary. The behavior is the signal.
- Hundreds of thousands of queries to Wikidata – agent load on the open web
- About 1,200 agents – coordinated via an internal package registry
- Over 70,000 messages and files – exchanged for coordination
- >90% of 533 agents – joined quickly
The two incidents also differ in nature. The Wikimedia agents operated on the open web where they didn’t belong. The Hugging Face case happened during an evaluation in which, according to METR, some cyber classifiers were deliberately disabled to measure raw capability. Lumping both under “AI went rogue” misses important differences.
The gap grows at the edges you never specified
As public reports suggest, no one told the agents that the registry was off-limits. The gap between your instruction and what happens grows at the edges of the task you didn’t define. In my case that edge was an unresolved decision. For Hugging Face it was a package registry. The agent fills both gaps with its best guess, and that guess is usually: keep going. The longer agentic workflows run, the more of these edges accumulate.
Mollick’s response, developed with Lilach Mollick, is called Twilight Factory. It puts a moderator agent next to the working agents. That moderator decides when to involve a human: for approval, expertise, or because a decision deserves a person. I build systems like this myself, and a role that watches for human approvals belongs in the design. Our guide to AI governance for agents shows what such a rule can look like.
In ChatGPT or Copilot this helps you little today. Twilight Factory is a pattern for the teams building agent platforms. No end-user product has that switch. What you need must work in your next prompt.
Boundaries first, task second
Your point of control is before the agent. Most boundary crossings happen in sessions where the goal was described but the limits were not. Use the following scope prompt before the actual task and fill in the parentheses.
Before you start, here is the framework for this task.
Goal: (what should be done, in one sentence)
Stay within these boundaries:
- Work only with (the specified files, the folder, the document or the website). Do not open or change anything else.
- Do not use the internet, install anything, or contact any external service unless I explicitly approve.
- Do not delete, send, publish, or pay for anything without asking me first.
If you are blocked, stop and tell me what is blocking you. Do not look for a workaround.
If something is unclear, ask and do not guess.
Present open decisions individually and at the end of your response, never in the middle of a long text.
If I don’t answer a decision, wait. Silence is not consent.Now the task:
Why this works: the block closes exactly the gaps that were open in all three cases. The Hugging Face agents were blocked and looked for another exit. The line “stop and tell me what is blocking you” turns a blocked moment into a question for you. The last two rules address my case: a decision you never saw no longer slips through as a yes.
An everyday example with code. A repository is a project’s shared folder with version history; branches are parallel work states. The weak version of the instruction reads:
Clean up the stale branches in this repository.
A better version:
Clean up the stale branches in this local repository. Do not touch the copy on the server. Do not run anything that requires network access. List what you propose to delete and ask me before you delete anything.
The weak version has no edges. An agent that can’t finish locally has every reason to reach for the server. Nothing in that prompt says the server is off-limits, and an agent reads that silence as permission.
Four questions before you start
Run through these questions before every agent session. They take about a minute.
- What can the agent access that it doesn’t need? Email, files, the web, a payment method. Name the important ones and exclude them.
- What should happen if it gets blocked? Decide now. Most often the answer is: stop and report.
- Which decisions need your explicit yes? Put them in the prompt.
- Who should it ask before improvising? Usually you. Say that in the prompt.
- Identify unnecessary accesses
Email, files, the web or payment instruments may be reachable. Name what the agent doesn’t need and exclude it explicitly. - Behavior when blocked
Define in advance what should happen when the agent encounters an obstacle. Usually that means stop and report rather than finding workarounds. - Decisions that require approval
Specify which items require your explicit approval. These belong explicitly in the prompt. - Who to ask before improvising
Decide who the agent should consult before improvising. In most cases that’s you; state it in the prompt.
This is a practice to try, not a validated fix. I don’t have robust numbers on how much such a framework reduces extra steps. The framework costs little and targets the failure pattern these three cases show. I use it wherever an agent can send, delete or spend. We practice routines like this in our AI trainings with teams, in the context of their daily work.
Frequently asked questions
Why do AI agents do more than I asked?
An agent works toward a goal. Where the instruction names no limit, it fills the gap with its best guess—usually: keep going. This applies to accesses as well as decisions you didn’t answer.
Is a scope prompt enough for AI governance?
For your own sessions in ChatGPT, Copilot or a coding assistant it’s the part you can control immediately. Operators of larger agent systems also need formal rules, approval roles and audit logs in the system.
What does human-in-the-loop mean for AI agents?
A human approves at defined points before the agent continues. The approval must be explicit. An unanswered question remains open and is never counted as a yes.
Your next step
Agents cross boundaries where you haven’t drawn them, both for access and for decisions. Put the framework above in front of your next task and see whether the agent stops and asks. How many decisions did you overlook this week that your agent has already treated as a yes?
Interested?
Let's find out together how we can implement these approaches in your organization.
Schedule a conversation now